1. What a RevOps Engagement Cost Us in Early 2025
Like any agency worth its salt, we had stellar processes and workflows that ran like this: discovery, then a roadmap, then a multi-month build, then a go-live.
Discovery meant an audit, and an audit meant a senior person and a CRM. They opened it and started reading. Workflows, properties, automations, data quality, one screen at a time, taking notes, and forming a view as they went.
It took us two to three weeks to tell our clients what was actually wrong.
But our audits weren’t designed to be purposefully slow and heavy. We wanted to position ourselves for perfection when we execute. So the hours we put in weren’t the problem. It’s just that our senior employees were embroiled in a lot of unglamorous work.
- Creating inventories
- Reconciling data
- Documentation
- Reconstruction of what a previous team built and why
None of it is the interesting part. Nevertheless, the most expensive people were doing it.
So senior capacity was the business constraint, and we were spending it on work that didn't need seniority. It needed thoroughness.
In short, this is what it looked like:
2. We were Selling the Fix but were Living the Problem we Solved for Others
There is a paragraph on our website about what happens to a revenue team that never systematized anything.
“Your team has become the automation.”
A revenue team that is neck-deep in grunt work assembles reporting by hand, every month, from the same sources as last month.
The more concerning effect is that as the number of accounts grows, headcount also does, rendering investments in automations moot.
We have said versions of that to prospects for years. But in August 2025, our own RevOps delivery had the same symptom on the list.
Hindsight is a gratifying feeling when we decode stories that don’t happen to us. However, it isn’t great when you are part of the problem.
While we could see the wrongs clearly in someone else's system, we had stopped seeing them in our own as we scaled.
Audits Were Comprehensive, but Error-prone, Wearisome and Unscalable
A senior person opened the client's CRM and started reading. They took notes. They built a picture, and in a couple of weeks, they could say what was actually wrong.
However, reading screens required thoroughness and stamina, not the senior person’s judgment.
And with two hundred things to check in a mature CRM by the time someone reaches the last forty workflows in week three, they grew weary.
For example, a workflow called Lead Routing that routes nothing survives in production for months. Nobody re-reads a workflow whose name sounds right when there are a 199 other things in the queue.
So the most common failure in a client's system and the most common failure in our method of finding it were the same failure. A name that sounds correct, accepted without inspection, because no one can inspect everything by hand.
Acquired Judgments and Knowledge didn’t Pass to the Team
Every CRM has a gap between what its documentation promises and what it does, such as:
- Property types that behave differently than the docs suggest
- Workflow limits you find by exceeding them
- Patterns that look fine in month one and break in month three for reasons nobody wrote down.
That knowledge is expensive. In our operation, we bought that knowledge once and then stored it in a person.
In practice, it looked like this:
Every CRM has a gap between what its documentation promises and what it does. For instance, a property set up as a dropdown cannot be converted to a text field later without breaking every report built on it.
A workflow enrolment trigger re-fires on records that re-enter the criteria, so a "notify on new MQL" automation notifies three times for the same account. A calculated property does not backfill, so it reads as null on every record created before it existed.
Such problems are learned by shipping and watching it fail.
So, an engineer configuring their first lifecycle model on HubSpot would have rebuilt a mistake an engineer three desks away had already made and solved on another account.
Such knowledge didn’t diffuse across the team.
We had playbooks. But they were more of a checklist than something that helps our engineers apply or vary their judgment across differing contexts.
Client Context Lived in One Person's Memory
This one cost the most and was the hardest to see. Every revenue system carries context that makes it make sense. Things like:
- Why a qualified lead means what it means at this company specifically
- Which automation exists for a compliance reason and must not be touched, however wrong it looks
- Which apparent bug is deliberate
- What was tried two years ago, why it was abandoned
None of that is documented anywhere in the client's system. While the respective account leads had context, anyone new has to learn the ropes by getting it wrong atleast once.
Which means a handover wasn't a handover. It was a re-discovery project, run at the client's expense, in which they re-explained their own business to a stranger and paid for the rediscovery in slower work and worse recommendations for a quarter.
We were one resignation away from that on every account we ran.
Growth Meant Headcount
Add an engagement, add a proportional slice of a senior person. That was the arithmetic, and it was the same arithmetic.
Senior capacity was being consumed by inventories, reconciliation and documentation. The work that gates everything and rewards thoroughness rather than expertise. So the business ceiling was set by how many CRMs our best people could read per quarter.
We Were Not Short of AI Tools
Models were available to anyone with a credit card, and we were using them in the ordinary way, for the ordinary things: drafting, summarising, first passes on documentation.
None of it changed the shape of the work. It made individual steps faster and left every handoff between steps exactly where it was, so the constraint didn't move.
A senior person still read the CRM. The knowledge still lived in their head. The next engagement still cost the same slice of the same person.
3. What We Found When We Opened Other People's Systems
Fifteen days of reading a CRM produces one useful thing, eventually: a picture of what is actually wrong. We built forty of those pictures. Three patterns turn up in them often enough to be predictable.
None of the three is an AI problem. That is the point of the chapter.
The System is Judging Leads it Cannot See
This is the most damaging and the least visible, which is a bad combination.
A qualification rule reads fields. That is all it does. When those fields are populated, it evaluates them against whatever thresholds someone configured and returns a verdict.
When those fields are empty, it does not pause. It does not flag. It does not ask. It evaluates against nothing and returns a verdict anyway.
A Rule with no Data Behind it Does not Return "unknown." It returns "no."
There is no queue of uncertain records, no exception report, no signal anywhere in the system that a decision was made on an empty file.
A good-fit buyer gets discarded for reasons that have nothing to do with fit, and the only trace is a record that was scored low and stopped being anybody's problem.
We have measured the relationship. Across our engagements, the share of leads with no resolvable firmographics tracks the share the logic cannot judge at r = 0.87.
Every record enrichment that cannot resolve becomes a record the rule cannot judge, which in practice means a record the rule declines. The failure is not occasional. It scales with how thin your database is, and most databases are thinner than their owners believe.
This is why we will not put a scoring or qualification layer on a database we have not enriched first, and why we are suspicious of any proposal that starts with the agent.
The right question to ask anyone building this for you is what happens when the fields are empty. "The rule handles it" is not an answer.
Demand Generation Channels Expand, but Automations did not Include Them
A CRM gets built for the sources that existed when it was configured. Forms, a demo request, maybe an event list.
But as the product moves to a self-serve model or marketing channels expand, demand starts coming from elsewhere: product usage, support conversations, and partner channels.
However, building the routing logic to capture new demands looks like a multi-week build. So sales teams handle them in spreadsheets as a temporary measure.
What all of the Above Problems look like
More than half the database is duplicated. Roughly two-thirds of leads go to the wrong owner, or to nobody.
They are also the reason AI projects fail. MIT's NANDA study found roughly 95% of enterprise AI pilots produced no measurable P&L impact, with tools attached to workflows nobody had mapped as the diagnosis. That number gets quoted as a statement about AI. Read against what is actually in these systems, it looks more like a statement about foundations.
Point a model at a database that is 56% duplicated, and it will produce wrong answers faster, with more confidence, at greater volume than the manual process it replaced.
Point one at routing logic that misfires two-thirds of the time, and it will misfire two-thirds of the time at machine speed. The model is not the variable.
4. How We Rebuilt It
Internal capability first, then client delivery. Never the other way round.
Phase 1: Rebuild the Audit as an Inventory
We started with the first thing every engagement starts with.
An audit.
The first pass is now systematic. We inventory and cross-check every workflow, property, automation and data-quality issue before anyone forms a view. The senior review then starts from a complete picture instead of a blank page.
Our AI-native systems cut through ambiguities like the Lead Routing problem we discussed earlier. It has no opinion about which names sound plausible and no stamina to run out of in week three. It checks what a workflow does and reports what it does, and the gap between that and what the workflow is called falls straight out.
Automations that lie about their own function are now the most common finding we surface in week one. They were presumably always there. We were finding them in month three, when someone hit the consequence.
The work did not get shallower. The inventory stopped being done by a human reading screens, and judgment moved to the front of the engagement instead of arriving exhausted at the end of the reading.
Phase 2: Codify Platform Knowledge and Client Context
Two artifacts, both closer to infrastructure than documentation.
Skills per platform: We packaged instructions that encode how a specific CRM actually behaves. Its property-type quirks, its workflow limits, the gap between what the documentation promises and what the platform does.
When an engineer builds a lifecycle model or a routing rule, that accumulated knowledge is available at the point of work rather than living in whoever hit the problem last time.
The effect shows up in delivery speed. Our fourth build on a platform is materially faster and better than our first, which wasn't true before, because the first build taught the fourth nothing.
That difference is the entire economic argument for an AI-native agency, and it only works if the knowledge is captured deliberately. An agency that uses AI gets faster at individual tasks. An agency that codifies each build gets structurally cheaper at every subsequent one.
A knowledge base per client: Every revenue system carries context such as:
- Why a qualified lead means what it means here
- Which automation exists for a compliance reason and must not be touched
- Which apparent bug is deliberate.
- What was tried two years ago, abandoned, and should not be proposed again.
The reason to care is turnover, on both sides. If that context lives in one account manager's head, the client is one resignation away from paying to re-explain their own business.
It is also what lets an embedded engagement compound rather than plateau. Running hundreds of requests a year for a dozen stakeholders only works if nobody is relearning the business each month.
And it is the honest version of the proprietary-data argument. The models are a commodity. Everyone has the same ones. The client-specific corpus is not.

An AI-native agency is where each build makes the next one cheaper.
Break that loop at step two, and you have an agency that uses AI. Most stop there because step two has no client attached and no invoice at the end.
Phase 3: Reshape Delivery Around Shipping
The old shape was discovery, roadmap, multi-month build, then a go-live where everyone discovered at once what had been misunderstood in month one.
Now each engagement is a scoped sprint solving a single bottleneck end to end. Two to six weeks to a live workflow, then the next bottleneck. Short cycles surface the misunderstanding in week two, when it is still cheap to be wrong.
Alerting into client channels. A system that only surfaces in a dashboard gets checked when someone remembers. So the work pushes into the channels the client's team already lives in. High-priority signals the moment they happen, ownership assigned automatically, handoffs announced where the receiving team will see them.
The same channels double as the debugging surface while logic is verified, so the client watches the system being proved rather than being told it was. Nothing waits for a weekly call.
Templates the client can run. An operations system that only works while the agency is engaged is not a system. It is a dependency. What we build is documented, monitored, and runnable without us. Buyers have gotten much sharper at spotting the difference, and they are right to have.
5. The Rule That Set the Sequence
Never Sell a Capability Which We Didn’t Run Ourselves.
If we could not point at it working inside our own delivery, it did not go into a proposal.
That rule decided the order of everything in the previous chapter. It is also why the previous chapter is about our audits, our platform knowledge and our client context rather than about anything a client bought.
The Order Matters More Than the Tools
The industry mostly did this in reverse. Buy the tools, announce the capability, learn on the client's engagement. It is a rational sequence if you believe the tool is the thing being bought.
But a tool is not a capability. A capability is a tool plus everything that had to be true around it for the tool to produce a result twice.
None of that is discoverable in a demo. All of it is discoverable in about six weeks of running it against your own work, where the cost of finding out is your time.
Run it internally first, and you learn it all on your own account. Sell it first, and the client funds the discovery, usually without being told that is what they are paying for.
What it Cost Us
The rule is expensive, and it is only worth stating if we are willing to say where.
It meant declining work we could probably have done. Through the back half of 2025, clients wanted things we had a credible path to build but no internal proof of, and the rule said no. It meant the gap between having an idea and being able to sell it was measured in months.
It also meant the first version of every capability ran against our own delivery, which is a lower-stakes environment in exactly one respect: when it broke, the person absorbing the cost had agreed to absorb it.
The Version of This a Buyer Can Use
The test transfers cleanly to anyone evaluating an agency, including us.
Ask what they run internally on the capability they are selling you. Not what they have built for clients, which tells you they can build it once. What they use themselves, daily, where they eat the failure. An agency that sells an AI-assisted audit and audits its own systems by hand is telling you something specific about how much they trust it.

Caption: What we rebuilt on our side, and what it changes on yours.
6. What Changed in How We Work
Tool Selection Became Evidential
We benchmark rather than assume, and the benchmark decides, including when it decides against a partner.
The practical change is when we make the decision. Vendor selection used to happen at the proposal stage, based on what we had used before and what integrated cleanly. It now happens after a measured comparison against ground truth, per field, costed per thousand records. It is slower at the start of an engagement, and it removes an entire category of problem three months in, when someone notices a field is wrong, and there is no way to establish whether it was ever right.
Every Decision Has to Explain Itself
This is the adoption rule, and it is the one that decides whether anything else in this chapter survives contact with the client's team.
The failure mode nobody quotes statistics about is the system working and the team routing around it. It happens for an entirely rational reason. If a rep cannot see why a lead was scored the way it was, they cannot act on it, defend it in a pipeline review, or trust it next week. So they open a browser and check manually, which is the exact work the client paid to remove.
So qualification returns a verdict and a reason code from a defined set. Some codes explain a rejection, some explain what qualified it. A rep who sees "No, this account is already a partner" can move on. A rep who sees "No" opens a browser.
Enriched values carry a receipt too: which provider supplied which attribute, contact by contact. When a field looks wrong, it gets traced in seconds instead of debated in a meeting, and the debate is the expensive part.
A verdict a rep cannot interrogate is a verdict a rep ignores. Worth asking any agency whether their output explains itself, and if the answer is a confidence score, worth asking what a rep is supposed to do with 0.72.
Two Design Rules We Now Apply by Default
Nulls pass, but they do not fail: This is the enrichment ceiling turned into a build standard. If half a database is missing a key attribute, a rule that fails on unknown disqualifies half the database for reasons unrelated to fit.
So missing data routes to review rather than producing a silent rejection. It makes the problem visible instead of making it disappear, which is the entire difference.
Strip your own noise explicitly: Internal records, test records, and anything containing the client's own company name will otherwise sit in the rep queue indistinguishable from real prospects. It is unglamorous, nobody specifies it, and it is usually a meaningful percentage of the queue.
Neither rule is clever. We now do both without being asked because we have seen what their absence costs.
The Build Covers History, Not Just New Arrivals
Most builds only improve what arrives after go-live. That leaves the majority of a database untouched and creates a two-tier system where a record's quality depends on when it landed, which is a distinction no rep can see and no report accounts for.
It is usually cheap to avoid. Build the logic so it evaluates history as well as new arrivals, and every existing record gets scored the moment it deploys. No separate re-processing project, and a queue that is sortable on day one rather than after a quarter of accumulation.
The question to ask, of us or anyone: does this apply to the data I already have, or only to what comes next?
Nothing Waits for the Weekly Call
Alerting into the client's channels changed what a status update is.
High-priority signals arrive the moment they happen, ownership is assigned automatically, and handoffs are announced where the receiving team will see them. The same channels double as the debugging surface while logic is being verified, so the client watches the system being proved rather than being told afterward that it was.
The second-order effect is on the meeting. A weekly call that exists to transmit information is mostly dead time. A weekly call where everyone already has the information is a decision meeting, which is the only kind worth scheduling.
What Runs Without a Person

An operations system that only works while the agency is engaged is not a system; it is a dependency, and buyers have gotten much sharper at telling the difference.
The knowledge base per client sits on our side but exists for the same reason.
7. The Numbers
Across roughly 40 RevOps engagements.
Three of those are worth unpacking, because two are less impressive than they look and one is the only number here we would defend as meaningful.
97% is Not a Model Doing Something Clever
A 97% cut in lead response time sounds like the headline. It is the most mechanical number in the table.
A form fill used to wait. It waited for an enrichment batch, or for a rep to work down a queue, or for whoever was on rotation to notice. The lead itself was fine. The time was spent on a sequence of handoffs, each needing a person to start it.
Now enrichment resolves firmographics before a human looks at the record, the qualification layer turns those fields into a verdict with a stated reason, and assignment puts the lead in front of a named owner. No queue, no overnight batch, no rotation. What is left for a rep is the only part that ever needed a person.
There is no intelligence in that. There is an absence of waiting. Most of what gets sold as AI speed is this, and it is worth having, but it is worth being clear about what produced it.
The Duplicate and Mis-Routing Numbers are Not Achievements
A duplicate rate going from 56% to 3% is a large movement, and it mostly measures how bad the starting point was.
More than half a database being duplicated is not unusual. Two-thirds of leads misrouted is not an unusual finding. These are the ordinary conditions of a CRM that has been running for several years through two or three ops owners and at least one migration, and fixing them is careful work rather than clever work.
We report them because they are the precondition for everything else and because they are what an AI project fails on. But an agency quoting a 53-point improvement in duplicate rate is quoting the client's starting position as much as its own contribution.
68% is the One we Watch
Client expansion rate since August 2025.
Most clients arrive with a single bottleneck. They have one thing that is visibly broken, and the engagement is scoped to that. Expansion means they came back for the second bottleneck after seeing what happened to the first.
It is the only measure of whether the first system was worth keeping, and it is the one number an agency cannot manufacture, because it requires someone to voluntarily buy again. Pipeline lift can be attributed generously. Response time can be measured from a favorable baseline. Expansion is a decision made by someone with a budget, after the fact, with full knowledge of what working with us is actually like.
If we were evaluating an agency, this is the number we would ask for, and the shape of the answer matters more than the figure. Expansion into what, from how many clients, and how many of those relationships existed before the period being measured.
8. What it Produced
Five builds, each solving a different bottleneck. They are written up separately, so this is what each one was for and what it demonstrates rather than a retelling.
Atlan: Six commercial providers plus a custom scraper, About pages as ground truth, a waterfall configured per field rather than per vendor. If you want the configuration detail behind the argument that the data layer comes first, it is here.
Atlas HXM: Our longest-running embedded engagement, and the clearest evidence for the knowledge base argument in chapter 5. Hundreds of requests a year across a dozen stakeholders only works if nobody is relearning the business each month. This is what the per-client corpus is for, and this account is where it stopped being a theory.
Domotz: A CRM built for form fills, and demand arriving from product usage. Lifecycle stages computed from what a customer actually did rather than tagged by hand in a dropdown. The test of whether adding a new demand source costs a project or a configuration.
Bench: The unglamorous one. Migrations are where historical attribution quietly dies, and the damage surfaces a quarter later when someone tries to run a comparison and finds the before and the after are not comparable.

