
180 days of AI in marketing operations: the real numbers
- 15 hours ago
- 7 min read
We deployed MOPsy, our AI agent for marketing operations, inside a global financial services company's Eloqua instance for a 180-day pilot. The goal was to find out whether an AI agent could meaningfully reduce the time and effort required for campaign operations: building emails, creating campaigns, running QA, and handling the repetitive operational tasks that consume most of a MOPs team's week.
Six months later, we have results. Some of them are exactly what we expected. Others taught us things we didn't anticipate. Both are worth sharing.
We're keeping the client's name confidential, but everything in this article reflects the actual outcomes from a live production environment with a real marketing operations team running real campaigns.
The problem we were solving
The client's marketing operations team was running a high-volume campaign operation on Eloqua. The work was consistent and repeatable: webinar invitation emails, newsletter builds, campaign setup, QA checklists across multiple campaign types. Each task followed a known process. Each one consumed time that added up.
An email build took about an hour of hands-on work and roughly two days from request to completion when you accounted for the back-and-forth, reviews, and queue time. A webinar campaign setup took an hour of build time and a day of elapsed time. QA on a campaign with multiple outbound emails took an hour of build time and up to three days of elapsed time.
None of these tasks were broken. The team was competent and the work got done. But the volume was high enough that the operational load left little room for anything else. The team was spending most of its capacity on repeatable execution rather than on optimization, analysis, or strategic work.
The question was whether MOPsy could take on the repeatable work reliably enough that the team could redirect their time toward higher-value activities.
What we deployed
We scoped 15 use cases across four categories of work.
Email creation. MOPsy builds emails from briefs, following the client's templates, brand guidelines, and content structure. The agent produces the email inside Eloqua, ready for human review.
Campaign creation. MOPsy sets up campaign structures in Eloqua, including the configuration steps that normally require a person to click through multiple screens and follow a documented process. For webinar campaigns, this includes the full setup from brief to launchable campaign.
Campaign QA. MOPsy runs QA checklists against live campaigns, checking for the errors and inconsistencies that human reviewers catch manually: broken links, incorrect field references, missing elements, configuration mistakes. We built separate QA checklists for webinar campaigns, outbound emails, newsletters, events, and a generic canvas that covers other campaign types.
Image editing. MOPsy handles simple image tasks like formatting webinar speaker headshots to the client's specifications. A small use case, but one that consumed surprisingly regular time from the team.
Each use case went through the same process: scoping with the client's team, building and refining the master prompts, testing in a controlled environment, iterating based on feedback, and then moving into production testing with the team using MOPsy on real campaigns.
The results after 180 days
The time savings were significant and consistent across the use cases that reached production testing.
Email creation went from approximately one hour of build time and two days of elapsed time to three to five minutes. The agent produces a draft that the team reviews and adjusts rather than building from scratch. The review is faster than the build because the starting point is already close to the finished product.
Webinar campaign creation went from an hour of build time and a day of elapsed time to roughly five minutes. The agent handles the configuration steps that previously required a person to work through the platform's interface manually.
Campaign QA went from 30 to 60 minutes of build time and two to three days of elapsed time to approximately five minutes per campaign. The agent runs the checklist systematically and produces a report the team reviews.
These are real numbers from a live environment. They represent the time between requesting the work and having a reviewable output, not just the build time in isolation.
Of the 15 use cases scoped at the start of the pilot, one is fully complete, eight are in testing or final edits with completion rates between 80% and 95%, three were deprioritized by the client's team based on shifting business needs, and three haven't started because the team chose to focus on finishing open use cases before adding new ones.
What we learned that we didn't expect
The time savings weren't the surprise. We expected MOPsy to be faster than manual execution for repeatable tasks. It was. The surprises were in the operational reality of getting an AI agent to the point where the team trusts it.
Prompt iteration takes longer than you'd think. Building the master prompts that govern how MOPsy handles each use case is a detailed, iterative process. The first version of a prompt produces output that's close but not right. The second version is closer. The third version handles most scenarios well but misses edge cases. Each iteration requires testing against real campaigns, reviewing the output, identifying what needs to change, and refining. For nuanced tasks like QA checklists where the criteria are specific and the tolerance for error is low, this iteration cycle takes weeks, not days.
Consistency requires ongoing vigilance. An AI agent doesn't drift the way a person drifts, but it can produce inconsistent output when the inputs vary in ways the prompt didn't anticipate. A campaign brief that's structured slightly differently than the ones the prompt was trained on can produce output that's subtly off. Maintaining consistent quality requires regular prompt reviews and updates as the team encounters new variations in their work.
Moving from a pilot group to broader adoption takes time. The initial testing happened with a small group of team members who were closely involved in the scoping and iteration. When the use cases moved to a wider group of stakeholders for review and testing, the feedback cycle lengthened. People who weren't part of the initial development had different expectations, different preferences, and different definitions of "good enough." Incorporating their feedback while maintaining what was already working required careful management.
Deep-dive technical sessions accelerate everything. The most productive moments in the pilot were sessions where the Sojourn team and the client's team sat down together to walk through specific use cases in detail, reviewing the prompts, testing variations, and making adjustments in real time. These sessions compressed weeks of asynchronous iteration into hours. We're building more of them into the next phase.
Not every use case is worth pursuing. Three of the original 15 use cases were deprioritized during the pilot because the client's business needs shifted. That's not a failure. It's how a well-managed pilot should work. Scoping broadly at the start and then focusing on the use cases that deliver the most value is better than committing rigidly to the original plan when circumstances change.
What the human role looks like
One of the most common questions about AI agents in marketing operations is whether they replace the team. After 180 days of running MOPsy in production alongside a real MOPs team, the answer is clearly no. But the role changes.
Before MOPsy, the team's work was primarily building: constructing emails, configuring campaigns, running through QA checklists manually. The work was skilled but repetitive, and the volume consumed most of the team's available hours.
With MOPsy, the team's work shifts toward reviewing and refining. The agent produces the first version. The team evaluates it, catches anything the agent missed, makes adjustments, and approves it for production. The skill required is the same. The time required is dramatically less. And the team's attention shifts from "can I get this built in time?" to "is this good enough to send?"
That shift frees capacity. The team has more time for the work that an AI agent can't do: strategic planning, process improvement, stakeholder management, and the judgment calls that require understanding the business context in ways an AI agent doesn't.
The number of human touches per task is still higher than we'd like. Reducing that number is a primary focus for the next phase. But the direction is clear: each iteration of the prompts reduces the number of times a human needs to intervene, and the goal is to reach a point where the majority of routine tasks require a single review step rather than multiple rounds of adjustment.
What comes next
The next six months focus on three things.
Completing the use cases currently in testing and getting them into full production use across the wider team. The core campaign creation and QA use cases are close to finished, and the priority is closing out the remaining iteration and moving them from "testing" to "standard workflow."
Expanding into reporting and insights. MOPsy has the potential to analyse campaign performance data and surface patterns the team would take hours to find manually. This use case has been on hold while the client's information security team reviews the data access requirements. Once approved, it opens a new category of value beyond operational execution.
Reducing human touches per task. Every prompt refinement, every edge case handled, every variation accounted for reduces the number of times the team needs to intervene. The target isn't zero human involvement. It's the minimum viable review: the agent does the work, the human confirms it's right, and the campaign moves forward.
Why this matters beyond one client
The results from this pilot are specific to one organization, one platform, and one set of use cases. But the patterns are generalizable.
Marketing operations teams across B2B are spending the majority of their capacity on repeatable execution. The work follows documented processes. The quality criteria are known. The volume is high. These are exactly the conditions where an AI agent can make a meaningful difference, not by replacing the team but by shifting their time from building to reviewing, from execution to oversight, from repetitive tasks to strategic work.
The key word is "meaningful." The time savings we measured in this pilot aren't marginal. They're transformational for how the team allocates its capacity. An email that took an hour to build and two days to deliver now takes minutes. That's not an incremental improvement. That's a structural change in how the operation works.
But the path to getting there requires investment: careful scoping, detailed prompt engineering, iterative testing, honest feedback loops, and the patience to get the output quality right before scaling. The 180 days weren't just about deploying an AI agent. They were about building the operational infrastructure that makes an AI agent trustworthy enough to rely on.
That infrastructure is what most teams skip when they activate AI features and hope for the best. It's also what separates an AI agent that genuinely transforms capacity from one that creates more work than it saves.
If you're considering an AI agent for your marketing operations, or if you've already started and the results aren't matching the expectations, we've been through the full cycle now. The scoping, the prompt engineering, the testing, the iteration, the honest conversations about what works and what doesn't. We know what the first 180 days actually look like because we've lived them. If a conversation about what this could look like for your team would be useful, we're happy to have it.










