Most companies are running a year behind.
Not because they lack the budget or the licences — because the value moved, and their plans didn’t. Here is where we think it went, and what we do about it.
The work AI can finish alone keeps getting longer.
Every few months, the length of task an AI can complete without a person stepping in roughly doubles. A year ago it drafted the email. Now it assembles the report. Next, it runs the project. Plans written for last year’s ceiling are already out of date.
- Draft an email
- Summarise a meeting
- Prepare a report
- Review a hundred contracts
- Run a week-long project
The model is a commodity. The harness isn’t.
Models are converging and increasingly interchangeable. What makes one company’s AI better than another’s is what surrounds the model: the context it’s given, the tools it can use, what it remembers, what it may touch and the checks on its work. Build that well and you can swap the model the day a better one ships.
Price the task, not the token.
A model that is cheaper per token can be more expensive per task — it takes more steps, more retries, more of a person’s time to fix. And for bulk work like sorting or labelling, a small specialised model can do the job for a fraction of a general one. We measure cost per finished task, and pick the model per job.
Per token
IllustrativeModel A
Model B
Per finished task
Model A
Model B
Small model, bulk labelling
- Model calls
- Retries and extra steps
- A person fixing the output
Evals are the product.
Producing an answer is the easy part. Knowing it’s right is the hard one — and it’s what most enterprise AI is missing. We build evaluation suites from your real work, run them before every release and every model change, and check outputs in a fresh context that didn’t write them. The person moves from doing the work to checking it.
Eval suite · supplier renewals
Before release · new model version
- Renewal summary cites every contract clause
- Figures match the ERP staging area
- People data never appears outside People
- Tone matches last quarter’s board pack
- Escalation count matches ticketing systemRegression
- Refuses to send without sign-off
5 of 6 passed
Release held · back to the team
Structure beats scale.
Pouring every document into one index doesn’t make knowledge findable; it makes it noisy. Structure — an ontology, labels, linked records — narrows what the AI has to read to answer a question. Less to read, less to get wrong. It’s also why knowledge-base projects stall: labelling is the hard part, so that’s the part we do.
Against a pile
Reads thousands of chunks. Most are near-misses; some contradict each other.
Against a structure
Reads four records. The path is the answer, and every step is traceable.
Decide where people stay.
Not every step needs a person, and not every step can go without one. If it can be undone, let it run. If it can’t — sending to a customer, signing a contract, deleting a record — a person signs off. We draw that line per workflow, with you, instead of leaving it to default.
Customer renewal pack
IllustrativeRead the contracts
Reversible
Draft recommendations
Reversible
Check against policy
Automatic
Share with Legal
Reversible
Sign the renewal
Irreversible
Approved by head of procurement · sent
Four tiers. Most companies stop at the first.
Adoption isn’t one switch. It’s four distinct capabilities — each needs more of the harness than the last.
Needs: Context · Memory
Find out which year you’re in.
One discovery call. We’ll map where your company sits on each of these — and what closing the gap would take.
Book a discovery call