Skip to content
Qofi
← all insights
EssayDec 2025 · 2 min read

the last mile: where AI meets the org chart

The demo works. The pilot works. Then the tool meets the organization, and the organization wins. Notes on the mile everyone underestimates.

There’s a stage in every AI deployment that no benchmark measures: the day the tool stops being a pilot and has to live inside the organization. Who runs it. Who checks it. Whose name goes on what it produces. Whose job description just changed without anyone saying so. We’ve come to call this the last mile, and in our project retrospectives it’s where more deployments die than at any technical stage.

The pattern is consistent enough to write down.

tools ship into workflows, not desks

The pilot succeeds because a pilot has an owner — one motivated person who wanted the tool and forgives its rough edges. Production has no such person. The draft report the agent produces at 6am is only useful if someone’s morning includes reading it, and “someone should read this” is an org-chart question, not a product question.

The deployments that stick are the ones where the workflow was redesigned around the tool, explicitly: the agent drafts, the associate reviews by 8, the VP signs by 10. Written down, with names. The ones that fail kept the old workflow intact and hoped the tool would find a seam to live in. Tools don’t find seams. People assign them.

Tools don’t find seams in the org chart. People assign them.

the review step is a job, so treat it like one

Every responsible deployment includes human review, and almost every one under-designs it. Review gets bolted on as a checkbox — “output verified” — performed by whoever had the least leverage to refuse. Then one of two things happens: the reviewer rubber-stamps, and the safeguard is theater; or the reviewer re-does the work, and the tool saved nothing.

Real review design answers three questions. What, specifically, is the reviewer checking — sources, logic, tone, all three? What does the reviewer see — a wall of output, or a diff against precedent with the uncertain parts flagged? And is review time budgeted in someone’s actual workload, or stolen from it? On one engagement, redesigning the review surface — showing what the agent was least confident about first — cut review time by more than half without changing the model at all.

incentives are load-bearing

A mid-level analyst whose status rests on being the person who builds the model has no reason to feed the agent that builds it faster — unless the definition of their job moves with the tool. This isn’t cynicism; it’s the system working as designed. People maintain what they’re measured on.

The organizations that navigate this well say the new expectation out loud: your value is now the judgment applied to the output, and here’s how that’s evaluated. The ones that navigate it badly announce “AI won’t replace anyone,” change no incentives, and then wonder why the tool is quietly unused. Silence reads as threat. Specifics read as a plan.

walk the mile before you build

The uncomfortable conclusion: the last mile can’t be deferred to deployment, because it’s mostly decided before the first line of code. Which workflow, whose review, what changes in whose job — these are scoping questions. We now spend the first weeks of every engagement on them, before any technical work, and it’s the highest-leverage time in the project. Nobody writes that phase into a statement of work. It is where the deployment is won or lost anyway.

← all insightsstart a conversation →