Skip to content
Shenzhen · The Greater Bay Area · Earth

The 90-Day AI Roadmap: Ship One Workflow and Label Everything Else a Guess

Most AI roadmaps are three-year capability maps priced as commitments. The version that survives a P&L is 90 days long, names one workflow, and pre-commits its stop condition.

8 min read1,796 words
AI StrategyEnterpriseNot yet translated.

Does a three-year AI roadmap survive its first budget review?

It rarely does, because it is a forecast wearing the clothes of a plan. I have sat in the room where a fourteen-month AI roadmap was approved in March and quietly deleted in November, and the failure was not technical — nobody disagreed about the model or the vendor. The failure was that the document committed to dates for work nobody had yet scoped, and the first honest estimate destroyed it.

A credible AI roadmap is 90 days long and contains one production workflow; everything past day 90 is a hypothesis to be triggered by evidence, not a phase to be scheduled.

The three-year capability map is popular because it is legible. It has swimlanes, it has a maturity curve, and it fits on one slide for a board that is asking a reasonable question: where does this go? But a plan is a sequence of commitments, and you cannot commit to something you cannot yet scope. What you can commit to, honestly, is 90 days and one workflow that real people use on real volume.

Why does the long roadmap feel safer than it is?

Because it prices uncertainty as certainty. A three-year map is effectively an option on a future capability, but it is written in the language of an obligation: hiring, platform selection, integration sequencing, change management. When the underlying technology moves — and the model tier that made a use case uneconomic last year is usually the one that is cheap this year — the map does not flex, it just becomes wrong in a document people have already been held to.

I will not put a precise number on how fast inference prices have fallen, because it depends on which model tier you measure and whether you are comparing list prices or negotiated rates. Every procurement conversation I have had in the last two years opened with a lower rate card than the one before it, and the same is true of the cost of hitting a fixed quality bar. Planning around one model's economics for 36 months means betting on a price that will almost certainly move in your favour.

There is a second cost, and it is the one I watch for. A long roadmap defers the only work that generates information. Discovery workshops and capability inventories produce documents; building one workflow through to production produces facts — how often the output is wrong, who has to review it, what the integration actually costs, whether the process owner will defend it in a Tuesday meeting. Every month spent mapping is a month not spent learning those.

What has to be true for one workflow to ship in 90 days?

Not every process qualifies. The constraint is not ambition, it is testability. In my experience a candidate workflow is only 90-day-shaped when six things hold at once, and when one of them fails the project does not fail loudly — it drifts into month seven with everyone still working on it.

The workflow runs on volume. Hundreds of instances a month, not dozens. Volume is what amortises the integration cost and gives you a statistically meaningful error rate. A workflow that fires twice a week can never tell you whether it works.

There is a deterministic success test. You need a way to score an output as right or wrong without a committee. If the only verdict available is "the team feels it is better," you have built a demo with a budget code.

A system of record already exists. The output has to land somewhere a human already looks — a ticket queue, an ERP field, a CRM record. Workflows that create a new destination create a new adoption problem.

A wrong output is cheap and reversible. Drafts that a human edits are fine. Anything that pays, ships, or posts without review is not a first workflow.

One named person owns the process today. Not a steering committee. A person whose week gets easier or harder depending on the outcome.

You can see the current cost of the manual process. If nobody can state what the process costs per instance today, you have no baseline, and a year from now nobody will be able to say whether the new one is cheaper.

If all six hold, 90 days is not aggressive. If three of them hold, 90 days is fantasy and the honest move is to fix the missing ones first, which is itself a 90-day project.

Which candidate workflow should you actually pick?

Score each candidate against the constraints above and pick the highest-volume one that passes every gate. This is the table I use in the first working session, and it tends to end the debate quickly because most candidates fail on the same two rows.

Candidate workflowMonthly volumeDeterministic test exists?System of record exists?Cost of a wrong outputVerdict for day 90
Invoice and PO exception triage1,000-8,000Yes, matches against PO line itemsYes, ERPLow, human approvesStrong first workflow
Inbound lead qualification and routing300-3,000Partly, routing is checkableYes, CRMLowStrong if routing rules are written down
Contract clause extraction50-400Yes, against a clause libraryPartialMedium, legal reviewViable, slower because of review volume
Internal policy question answeringUnboundedNo, correctness is contestedNoLow but invisiblePoor first workflow, no baseline to beat
Customer-facing email drafting200-2,000No, quality is subjectivePartialMedium to high, brand riskHold until a review layer exists
Autonomous procurement decisionsLowNoYesHigh, financialNot a 90-day project under any framing

The pattern is consistent: the workflows that fail the table are the ones with the highest imagined value. Internal question answering is the most requested project I am asked about and the worst possible first one, because nobody can define a correct answer, so the evaluation harness becomes a research programme and the 90 days disappear into it.

Where does the money actually go in 90 days?

Almost none of it goes to inference. That is the number that surprises finance teams, and it is worth putting in front of them early, because it relocates the argument to where the cost really is.

A worked example, with the arithmetic stated so you can redo it against current published API list prices rather than trusting my inputs: a claims-triage workflow handling 4,000 documents a month, averaging roughly 1,200 input and 400 output tokens per document, is about 4.8 million input and 1.6 million output tokens. At a frontier-tier list price on the order of a few dollars per million input tokens and roughly four times that for output, that is tens of dollars per month, not thousands. At a mid-tier model it is single-digit dollars. Rate cards move constantly, so treat those as placeholder rates and re-run the multiplication — the shape of the answer does not change.

Line itemRealistic 90-day shapeWhere the risk sits
Model inferenceTens of dollars to low hundreds per month at current list pricesAlmost none; recheck the rate card quarterly
Integration engineering4-6 weeks of one engineer against one system of recordUnderestimated when the system has no API
Evaluation harness1-2 weeks to build a scored test set of 200-500 real casesSkipped most often, and it is the reason nobody can prove the result
Human review layerOngoing, sized as a percentage of volume, not a fixed headcountThe real operating cost, and the one that decides whether you scale
Process owner time3-5 hours a week for 12 weeksThe most commonly unbooked line in the plan
External advisoryFixed-scope, not retainer, for the 90 daysUnbounded when the scope is a roadmap instead of a workflow

Two things follow. First, the case for a first workflow does not rest on labour arbitrage against inference cost, because inference was never where the money went; it rests on cycle time and on whether the review layer is a permanent tax or a temporary scaffold. Second, the review percentage is the number to put in the day-90 report. If 40 of 100 outputs need a human correction after eight weeks, you have a tool. If it is 4 of 100, you have a candidate for a process change, and that is a different conversation.

What do you do with day 91 and beyond?

You label it. The section after the 90 days in my plans is not a roadmap, it is a register of named bets, each with an explicit trigger and an explicit kill condition. No dates, because dates past day 90 are invented.

Three entries are usually enough: the next workflow, conditioned on the first one holding its review percentage for four consecutive weeks; the platform decision, conditioned on a second workflow needing the same integration surface; and the headcount or vendor decision, conditioned on volume crossing a stated threshold. Each is written so a reader can tell whether it has fired. If a line cannot be phrased as "we will do X when we observe Y," it is not a bet, it is a wish, and wishes do not belong in a document that gets approved.

This is also where the honesty pays off organisationally. When a transformation lead presents a three-year map, every quarter is a referendum on the whole programme. When she presents 90 days and one workflow, the day-90 review is a single decision with three pre-agreed outcomes — expand, hold, or stop — and the stop outcome has to be genuinely available, so the person who can say stop belongs in the day-45 review, not reading about it afterwards. A pilot that cannot be cancelled is a rollout with worse documentation.

What changes on Monday?

Take the roadmap you already have, find the first workflow listed in phase one, and ask the six questions above in a single meeting. If the answers hold, cut the document at day 90, delete everything after it, and rewrite what remains as bets with triggers; if they do not hold, you have learned that before spending the discovery budget rather than after. The uncomfortable part is that this makes the plan shorter than the organisation expects, and someone will ask what happens in year two. The answer is that you will know in 90 days, and knowing is worth more than a schedule. If you want the failure mode this prevents spelled out first, the mechanics are in why AI pilots stall after the demo — the same stall is what a 36-month roadmap institutionalises.

Keep reading

More in AI Strategy

Ready to build a system?[ Book a Call ]