Stop paying frontier prices for a task that doesn't need a frontier model.

Whether you pay an AI provider for it, pay a team to do it by hand, or run it as a service line for your clients, it is one repetitive decision made over and over: fields read off documents, tickets routed to the right queue, records classified or matched. It moves onto a model trained on the answers your own people have already given — and it is measured against what you do today before anything switches.

The qualifying test is simple: you already have a history of correct answers. That history is the training data, so there is nothing to annotate and nothing to specify from scratch. Six questions gets you a written read, free, within a working day.

It is measured, not asserted

Two public tasks, same frontier model on the other side of both, scored the same conservative way. Neither number is a projection.

Mistakes routing messages 54% fewer than the same frontier model — micro F1 0.9217 against Claude Opus 5's 0.8296, on 600 held-out messages
Cost per document 8×–62× less depending which model you run today — $1.02 against $8.32 to $63.18 per 1,000, at the providers' standard rates
Accuracy on contract fields 1.29× Claude Opus 5's, given the same worked examples — micro F1 0.512 against 0.396, difference +0.116 [+0.069, +0.161]
Re-tested on 5,000 resamples 5,000 of 5,000 came out ahead — so the result is not an artefact of which documents happened to be in the test set

Those are the plain-English versions, with the measurement beside each one. The full results carry the rest: every field, both baselines, the confidence intervals, the scoring rules, and the places these models still get it wrong — including a flaw we found in our own benchmark and published rather than quietly corrected.

How it works

01

Audit

An assessment of one endpoint: whether it can move to a more efficient model, the projected saving, and the risks. 2–3 days, and it behaves as a deposit rather than a fee.

02

Migrate

A model trained on your own history — system logs if you have them, your team's past records if you don't — and measured against how the work is done today, on examples it was never trained on. Nothing switches until it matches, and the harness that proves it ships to you.

03

Run

Coldstill hosts and operates it — hardware, monitoring, weekly quality checks and retraining. If you already call an AI provider, that is a one-line change to point somewhere else; if the work is done by hand today, this is the step where it stops being. Either way, nothing else on your side changes.

Fixed fees at every stage, no hourly billing and no infrastructure to buy. What it costs.

If you do this work on behalf of other people — a bookkeeping or accounting practice, an outsourcer, an agency with a service line — the arithmetic works hardest for you. One model, trained once, handles every client whose work has the same shape, so the volume that justifies it is the sum of theirs rather than any one of them. The saving is yours, because the cost being replaced is yours.

Get a free read on your task

Answer these and a written read comes back within a working day: your current spend with the arithmetic shown, whether the task can move to a model you own, the likely saving, and the risk most likely to threaten it. If the honest answer is that your current setup already does the job, that is what it says.

How is the work done today?

Numbers, not documents — nothing confidential is needed, and what you send goes to one mailbox. What happens to it.

Prefer email? hello@coldstill.org — the same answers in a message work just as well.