Pricing
The same two engagements apply whether the task runs through a frontier API, has just started on one, or is still done by hand. Every migration ends with a working model you own; what happens next is your choice, below.
The two engagements
Distillability audit
£600
- Behaves like a deposit, not a fee. The whole £600 comes off the migration if you start within 90 days, and comes back in full if the task turns out to be one Coldstill can't take on
- 2–3 days on a sample of your production data
- Spend breakdown, and a verdict on whether the task can move to a more efficient model
- Projected post-migration cost and payback period
- Risk register: where the migration is most likely to fail
- Recommended approach, with reasoning
- The eval harness and its output, so every number can be re-run on your side
Migration
£3,000
- 2–3 weeks, fixed fee, 50% deposit
- Parity is the acceptance condition: nothing switches until the eval shows the new model holds your quality on your data
- Trained model weights, built from your production logs or historical records, owned by you
- Serving config — used as-is by Coldstill's hosted service, or on infrastructure you run yourself
- The eval harness behind the report, so every number can be re-run on your side
- Handover doc: when to retrain as your data drifts, and how to read eval failures
Then Coldstill runs it
Hosted operation
from £500/mo
- Coldstill hosts and operates everything, GPU costs included — your only change is the endpoint your calls point at
- Capped at half your current or equivalent API bill where that half exceeds the £500/mo minimum; the exact fee is fixed in the audit — and as market prices fall, the cap falls with them
- Same quality, held: monitoring, a weekly drift eval, and retraining when it crosses the agreed threshold
- Cancel monthly and take the weights with you — they're yours from day one
- Your data is processed only to operate and evaluate your model, under a UK GDPR data-processing agreement — see how your data is handled
Or run it yourself
included
- For teams with their own ML infrastructure and someone to own it
- You already have the weights, the serving config and the eval harness — nothing further to buy
- Handover doc covers retraining and reading eval failures
- Questions answered by email for the first 30 days
What protects you at each step
Every stage of this is designed so you can stop, check, or walk away with something in hand. Nothing asks you to take the result on trust.
- Finding out costs nothing. The free read is a written verdict on one endpoint or one manual process, from its volume figures, within a working day. If your current setup already does the job, that is what it says.
- The audit is a deposit, not a spend. The whole £600 comes off the migration if you start within 90 days, and comes back in full if the task turns out to be one Coldstill can't take on.
- Nothing switches until it matches. Parity is the acceptance condition, measured on your own held-out data — and the eval harness ships to you, so it is your measurement, not our assurance.
- You own it from day one. Weights, serving config, eval harness and handover doc are yours as they are produced, so the work survives regardless of what happens to Coldstill.
- Hosting is monthly, never locked. Cancel whenever, take the weights, keep running.
- You don't have to send us your data at all. The model can be built on a synthetic corpus generated for your task, or the work can run inside your own environment under access you grant and revoke. Coldstill does the work either way — you are not handed a pipeline to operate. The three routes, and what leaves your side under each.
Start with the free read — six questions, and a written verdict on one task comes back within a working day. Email works just as well: [email protected].
What you're paying now
The fees above are what Coldstill charges. This is the other half of the arithmetic — what the task costs you today, from published API prices.
Straight answers
- "Frontier prices keep falling — waiting is cheaper." On average they fall; for a fixed workload they can also rise, on the vendor's schedule — Claude Sonnet 5's published list price rises 50% on 1 September 2026. The hosted fee is capped against your bill, so the ceiling falls with the market, and the audit prices against your optimised bill, with cheaper tiers and caching applied.
- "A consultant who asks for my numbers will always find a problem." The audit has three possible verdicts and one is "do not migrate" — a complete deliverable, written and delivered before any migration is scheduled. If the task turns out to be one Coldstill can't take on, the fee comes back.
- "What happens when the model I'm renting gets retired?" That is the failure mode we resolve. In 2026 both providers retired models their customers were using, on their own schedules, and OpenAI has announced that fine-tunes built on GPT-3.5 Turbo and GPT-4 stop running on 23 October 2026 — the dates, with sources. A model you own runs until you decide otherwise.
Who this is for
One repetitive judgement, made thousands of times a day, where you already have a history of the right answers. That is the whole qualification. Four things make a task fit it.
- Real volume behind one task. Above about £5,000 a month of API spend on a single endpoint, or about £1,000 a month of staff time if the work is still done by hand. For practices, outsourcers and agencies working across many clients, the combined figure counts.
- An objective right answer. Output that can be scored field by field against records you already hold, which is what lets parity be proven before anything switches.
- A history to learn from. Production logs, processed records, resolved tickets — the labels are usually a byproduct of doing the work, so there's no annotation project to fund first.
- A spec that holds still. A model trained on today's task keeps doing today's task: an advantage once the definition has settled, a drawback while it's still moving.
Extraction, classification, routing, tagging, matching and scoring all take this shape. The free read says which side of the line your task falls on, before you commit anything.