Writing

AI · GTM engineering

Amplification or Relocation?

Your AI stack can either make you more productive or trap you in an Escher-like maze of plausible lies. People are still calling both outcomes ‘productivity’.

· 10 min read

Over the last few months, honestly ever since the release of Opus 4.5, I automated a good chunk of my execution work. Account research, list building, personalization, sequencing, first-draft content, big chunks of reporting. For about three weeks it felt like I had achieved enlightenment and freed myself from the repetitive doldrums of keeping a GTM motion alive.

Then I stepped back and audited where that recovered time actually went. Tuning prompts. Reviewing outputs. Diagnosing why a pipeline that worked Tuesday broke Thursday, and kept breaking several more times that month. Rebuilding a research workflow after a model update quietly changed its behavior. Although my job was ostensibly to lead our growth function, my day-to-day role became that of a babysitter for a large, tireless, senile, and occasionally delusional workforce.

Within a quarter, the job had changed completely, and in ways that most teams haven’t acknowledged.

GTM is especially vulnerable to this mistake, because so much of the work already had a plausibility problem. A bad account list can look good in a spreadsheet. A bad outbound motion can look good in a sequence tool. AI doesn’t really fix this problem (and fundamentally can’t), but it does enable a small team to produce plausibility at industrial scale.

Your stack did one of two things to you, and the difference decides whether you are compounding or quietly depreciating.

The automation trap

As a whole, there are two core ways that AI impacts your work.

Amplification means the system does work you understand, your time moves up the stack into judgment, and the system needs less oversight as it matures. The name for this property is supervision decay. A junior hire should need less review after six months. A good AI agent should too.

Relocation means the machine took your labor and moved it somewhere else along the chain. Frequently this moves you into: review, repair, exception handling, prompt maintenance, source checking, and political cleanup when the output reaches another team and turns out to be wrong. This trades execution labor for supervision labor at roughly one-to-one, sometimes worse.

The new labor feels technical and current, which disguises the problem: it is less transferable than the skill it replaced, and it is owned by whoever owns the tools.

The simplest way to tell them apart: where do your corrections go?

Amplifying systems metabolize your corrections. Every fix changes the system, so the next output needs less fixing. Relocating systems just consume them. Every fix dies in a stateless chat window, and you will perform the identical fix tomorrow. If your corrections do not feed back into the process, you have not automated the work, you’ve just subscribed to the Groundhog Day experience, but with more spreadsheets and less charm.

Early on, the two feel identical. They are indistinguishable in month one. Both produce the same feeling of productivity. They separate on a longer time constant, six months, a year, and by then you have already reorganized your job around it.

Fast feedback tells you whether the system technically works. Only slow feedback tells you what it did to you. Almost everyone evaluates on the fast loop and lives with the results of the slow one.

Three taxes

Every automation levies a series of taxes that frequently go unsaid.

The review tax. Output scales faster than attention. When a system 10x’s your throughput, your review capacity does not 10x with it, so review gets thinner while accountability stays exactly where it was: on you. The system’s error rate hasn’t changed, but your catch rate has, which means your effective error rate is climbing while your confidence stays flat.

The failure tax. Savings arrive gradually, an hour here, an hour there. Failures arrive as events. I once shipped a territory list to sales with the wrong countries in it. The system had been reliable for weeks, my review had thinned accordingly, and the error walked straight through. Sales trust that took months to build was gone in roughly the time it takes to read a confused Slack message. The machine can generate plausible-looking work at the cost of your credibility.

The maintenance tax. Every automated workflow is another thing that can drift, break, or silently degrade when a model updates, an API changes, or the input data shifts. Add workflows linearly and the maintenance surface grows with them. Your job quietly moves from outcomes to platform maintenance, and nobody budgeted for the platform team, because the platform team is you.

Vendors don’t even have to lie for the story to be misleading. They sell the fast loop because the fast loop is what people think they want. The slow loop is what they actually need. That is not a conspiracy, just an incentive structure, and incentive structures do not need to be malicious to be reliably wrong about your interests; nobody is going to sell you something you don’t say you want.

The worst part is that relocation does not feel like decline while it is happening. The work feels like you’re on the technological frontier. You are busy, you are “with it”, you are “working with AI.” The market even rewards it for now. Busy-and-with-it is exactly what the middle managers of the last era felt, right up until the org chart flattened.

How to know it happened to you

Relocation announces itself quietly. You know the moods of each model better than the moods of your pipeline. Your evenings are punctuated by the occasional workflow breakdown. You have opinions about model releases the way you used to have opinions about your ICP.

When something ships wrong, your first instinct is no longer “what did I miss” but “what did it do this time,” which is the exact sentence every manager of a struggling team has muttered since the beginning of management, except your team costs a few hundred dollars a month, never learns, and cannot be fired because you built your process around them. And when someone asks what you have actually gotten better at this year, there is a pause followed by a thousand-yard stare.

If you recognize more than two of these, run the cheapest experiment available: turn the stack off for a couple of weeks and see what comes back. If the absence feels like lost leverage, a capable person suddenly missing their tools, you are fine, build more. If it feels like relief, or you discover you are no longer sure how to do the underlying work by hand, the machine did not extend you. It replaced the part of you that did the work and kept the part that signs off on liability.

No claim without a receipt

Here’s an example that I am sure everyone trying to implement AI agents has lived through.

I built an account research agent to compile briefs for our outbound motion: who the account is, why now, what to say. The first version was ‘textbook plausible’ to be generous. The model inferred relevance from whatever it found: “they are likely scaling their platform team,” “security is probably a priority this quarter.” It read well. Then a rep built a call around one of those confident claims, the claim turned out to be fiction, and from that point on every brief got quietly re-researched by hand. Which meant I had not automated research at all. I had relocated it. The reps were now doing verification work on top of their actual jobs, reading every brief and wondering which sentences were real.

The redesign came down to one rule: no claim ships without a source.

Every line in a brief now traces to something deterministic. A job posting that names the tools they are evaluating. A sentence their VP actually said on a podcast, quoted and linked. Product signup, observed site activity, detected stack.

If a claim has no source, it does not go in the brief, no matter how plausible it sounds. The model did not get smarter. The output got checkable. Spot-checks went from constant to rare and review time per account collapsed, because verifying a claim with its evidence attached takes seconds and verifying a naked claim takes a few browser tabs and ten minutes of pain.

That is the other way out of relocation (the first is making your corrections permanent), and it is the one GTM mostly skips: you do not earn supervision decay by generating ‘better’ or more. You get there by making work cheap to verify.

Chosen relocation vs drift

To be clear, relocating labor is not always bad and can be a strategy. Moving from executor to architect is a real career path, arguably the defining one of this decade, and it is roughly what I did on purpose. But chosen relocation and accidental relocation are different things. One is career design. The other is drift with a software budget.

‘More’ does not mean ‘good’

Step back and the market-level version of this is uncomfortable: the first wave of AI in GTM has mostly relocated labor from execution to supervision. The next wave has to prove it can reduce supervision, not just increase output. Output is cheap now. Review is the real bottleneck for most teams.

Software development is currently living in this movie, as they’re about one wave ahead of us. I watched it from the inside: at Ona our product was built around exactly this thesis, that the constraint in AI-assisted engineering had moved from generating code to governing it.

The pattern was consistent. Engineering orgs rolled out coding agents, individual engineers got dramatically faster, pull requests flooded in, and then the uncomfortable discovery: cycle time didn’t move. Individually faster engineers did not make the organization faster, because review had become the bottleneck, and the gains were compounding with the person instead of the system.

The serious end of that market responded by building verification infrastructure: sandboxed environments, audit trails enforced at the execution layer rather than in prompts, agents that prove their work through tests and diffs and review gates, and that run outside of your laptop. This moves engineers to being operators rather than trapped into painstakingly reviewing agent output.

GTM is at the start of the same curve with none of the base layers. Code has a native verification layer that GTM never built: a diff shows exactly what changed, a test suite fails loudly, a pull request is a built-in review gate. There is no diff for a territory list. There is no test suite for a pipeline narrative. A bad campaign doesn’t fail a build; it ships, performs plausibly, and spends your credibility (and real dollars) on a delay.

The asymmetry between the two is not speed, it is detectability. Engineering failures are loud: builds break, tests fail, error rates spike, customers file tickets. GTM failures are silent. A bad list performs “fine,” a bad narrative gets nodded at in the board meeting, and the miss arrives a quarter or two later, buried in a pipeline number with a dozen explanations. Nobody files a bug report about a deal that was never sourced. The revenue doesn’t break; it just never exists.

As a whole that means teams or individuals that build their own verification layers will be the ones for whom these systems ever become actual leverage.

If your stack creates more output without reducing the judgment required to trust that output, you did not buy leverage. You bought a faster way to create work for yourself.

Which means the real evaluation criterion for anything you add to your stack is supervision decay: does this system become cheaper to trust over time? Does it show receipts, check its own state, and earn reduced oversight the way a good hire does? An AI workflow that needs permanent adult supervision is not an agent. It is an intern with a hand grenade.

More writing