Pricing

The same two engagements apply whether the task runs through a frontier API, has just started on one, or is still done by hand. Every migration ends with a working model you own; what happens next is your choice, below.

The two engagements

Distillability audit

£600

  • Behaves like a deposit, not a fee. The whole £600 comes off the migration if you start within 90 days, and comes back in full if the task turns out to be one Coldstill can't take on
  • 2–3 days on a sample of your production data
  • Spend breakdown, and a verdict on whether the task can move to a more efficient model
  • Projected post-migration cost and payback period
  • Risk register: where the migration is most likely to fail
  • Recommended approach, with reasoning
  • The eval harness and its output, so every number can be re-run on your side

Migration

£3,000

  • 2–3 weeks, fixed fee, 50% deposit
  • Parity is the acceptance condition: nothing switches until the eval shows the new model holds your quality on your data
  • Trained model weights, built from your production logs or historical records, owned by you
  • Serving config — used as-is by Coldstill's hosted service, or on infrastructure you run yourself
  • The eval harness behind the report, so every number can be re-run on your side
  • Handover doc: when to retrain as your data drifts, and how to read eval failures

Then Coldstill runs it

Hosted operation

from £500/mo

  • Coldstill hosts and operates everything, GPU costs included — your only change is the endpoint your calls point at
  • Capped at half your current or equivalent API bill where that half exceeds the £500/mo minimum; the exact fee is fixed in the audit — and as market prices fall, the cap falls with them
  • Same quality, held: monitoring, a weekly drift eval, and retraining when it crosses the agreed threshold
  • Cancel monthly and take the weights with you — they're yours from day one
  • Your data is processed only to operate and evaluate your model, under a UK GDPR data-processing agreement — see how your data is handled

Or run it yourself

included

  • For teams with their own ML infrastructure and someone to own it
  • You already have the weights, the serving config and the eval harness — nothing further to buy
  • Handover doc covers retraining and reading eval failures
  • Questions answered by email for the first 30 days

What protects you at each step

Every stage of this is designed so you can stop, check, or walk away with something in hand. Nothing asks you to take the result on trust.

Start with the free read — six questions, and a written verdict on one task comes back within a working day. Email works just as well: [email protected].

What you're paying now

The fees above are what Coldstill charges. This is the other half of the arithmetic — what the task costs you today, from published API prices.

Spend estimate — public API pricing

Current spend, per month $7,500/mo
Hosted with Coldstill, at most (minimum £500/mo) $3,750/mo
Your saving, at least $45,000/yr

List prices, verified 2026-08-04, in USD, assuming 75% of tokens are input. Your real bill will differ — long-context calls, newer tokenizers, batch discounts and prompt caching all move it. The free read works from your actual token counts, not from this estimate.

Hosted operation is capped at half your current bill, GPU costs included, from £500/mo — and as market prices fall, that ceiling falls with them. Send six questions and the free read works from your actual token counts rather than this estimate.

Straight answers

Who this is for

One repetitive judgement, made thousands of times a day, where you already have a history of the right answers. That is the whole qualification. Four things make a task fit it.

Extraction, classification, routing, tagging, matching and scoring all take this shape. The free read says which side of the line your task falls on, before you commit anything.