Pricing
No markup on tokens. Ever.
We add 5.5% when you buy credits and nothing after that. Tokens are billed at each provider's list price with no markup, so what you spend on a model is what the model costs. There is no subscription, no seat fee and no minimum - credits do not expire, and the smallest top-up is $5.00.
| Feature | FreeNo card required | Pay as you goExtras unlock at $25.00 bought |
|---|---|---|
| Platform feeCharged on the top-up, never on a request. Tokens are billed at each provider's list price with no markup. | None | 5.5% when you add credits |
| Models | 2: GPT-5 mini and DeepSeek V4 Flash | All 20 |
| Providers | 2 of 8 | All 8 |
| Requests per day | 25 | No limit |
| Requests per minuteEvery dollar on the balance buys more headroom, so a busy account is not throttled for being busy. | 20 | 60, rising to 600 with your balance |
| Reply length | Up to 8,000 tokens | Whatever the model allows |
| Fallback across providers | Included | Included |
| Streaming, tool calling, the whole API | Included | Included |
| Request logs and usage breakdown | Included | Included |
| Playground | Included | Included |
| API keysRevoked keys do not count against the free limit. | 2 | No limit |
| PromptsStored conversation prefixes with {{variables}}, called from the model field as prompt/<name>. The wording changes without a deploy. | 3 | Up to 100 |
| GuardrailsSpend caps, model and provider access, prompt-injection and sensitive-data inspection, attached per key. Editing the one you have is never blocked. | 1 | No limit |
| Injection allow list | 5 phrases | 200 phrases |
| Cached answers heldAn identical request is answered from the last reply instead of a provider, so it costs nothing at all. Off unless you turn it on - keeping the answer means your text on our disk. Entries expire after an hour. | 100 | 5,000 |
| Answer near-identical requests tooAnswers a request that means the same as an earlier one - above 95% similarity, with the score in the reply so a near match never reads as an exact one. The one cache that costs us on every attempt, matched or not. | Not included | Included |
| Routing policiesName the models a request tries and the order to try them, or spread traffic across providers by weight. | Not included | Yours, plus 4 built in |
| Bring your own provider keysYour key pays the provider; the routing, logging and storage on this side stay ours to carry. | Not included | 5% of list price instead of the tokens |
| Store prompts and repliesOff unless you turn it on, at either tier. Turning it off is never blocked and deletes what was stored. | Not included | Included |
| Automatic top-upBuy credit when the balance runs low, so requests never start failing while you are asleep. | Not included | Included |
| Support | Email, best effort | |
| Start free | Add credit |
Adding credit is the whole of it - there is no plan to choose and nothing to provision. The next request after your payment lands is served on the full account, and if the balance reaches zero you drop back to the free allowance rather than getting errors. The extras in the right-hand column unlock once your purchases total $25.00, and stay unlocked at any balance after that.
Questions people ask before paying
Billing
- How are tokens billed?
- At each provider's own list price, with no markup. We add 5.5% when you buy credit and nothing after that - the fee is on the top-up, never on a request. Prices are held in the model catalogue with the date a person last checked them against the provider's own page, and shown on the Models screen.
- Do you mark up provider pricing?
- No. What you spend on a model is what the model costs. Two providers charge a second, dearer band - Google on very long prompts, DeepSeek at busy hours - and we pass those through exactly as they are charged to us rather than averaging them into one figure that would be wrong in both directions.
- Are failed or fallback attempts billed?
- No. Only the model that actually answered is charged. When a provider fails and the request falls through to the next one, the failed attempt produced no tokens and costs nothing - it is recorded in your logs so you can see it happened, and the response lists what was tried.
- Is there a minimum spend or a lock-in?
- Neither. The smallest top-up is $5.00, there is no subscription, no seat fee and no monthly commitment, and credits do not expire. Stop using it and nothing is charged.
- What unlocks the extra features?
- Total purchases of $25.00, counted across everything you have ever bought rather than what is left on the balance. Cross it once and routing policies, your own provider keys, prompt storage and the unlimited key and guardrail counts stay unlocked for good - spending the balance back down does not take them away. The model list, the daily count and the reply ceiling are the other way round: those follow the balance, because you cannot spend what you do not have.
- What happens when my balance reaches zero?
- Requests fall back to the free allowance rather than failing - 25 a day on 2 models. Once that is used up the API returns 402 until you top up or the day rolls over. Features you unlocked by paying stay unlocked.
- Can I stop it charging my card automatically?
- Automatic top-up is off unless you switch it on and tick a separate consent. When it is on it charges only when your balance falls below the threshold you set, at most ten times in a day as a hard stop so nothing can loop. You can switch it off in Settings at any time.
Limits
- Do you rate limit?
- Yes, per minute, and the ceiling rises with your balance rather than with a plan you choose. A key can also carry a lower limit of its own, which can only ever tighten the account's - a key that could raise its own ceiling would not be a guard rail.
- Can I separate staging from production?
- Make a key for each and attach a guardrail to the staging one. A guardrail carries a spend cap, which models and providers the key may reach, and whether prompts are inspected - so a leaked staging key costs part of an account rather than all of it.
- What does the free allowance actually cover?
- 2 models, 25 requests a day, and replies capped at 8,000 tokens. It is meant to be enough to read the quickstart, paste a key and see a real answer come back - not enough to build a product on, and it is bounded three ways at once so that stays true.
Routing
- What happens if a provider is down?
- The request carries on to the next model in the chain instead of failing. One request went out, one answer came back, and the response tells you what was tried on the way. We route across 8 providers and 20 models, so an outage at one is not an outage for you.
- Can I pin a specific model?
- Yes - name it, and it is used. An unrecognised id is refused rather than quietly substituted, so a typo never bills you for a different vendor. If you want a fixed order across several models, write a routing policy and send its name instead.
- Do you train on my prompts, or keep them?
- We never train on anything you send. Prompts and replies are not stored at all unless you switch that on, and switching it off deletes what was kept. What each provider does with a request is their own published position, recorded per provider with the date it was last read and a link to their page - including the one provider that does train on prompts, which is available if you name it and is never chosen for you.
Anything else
- Do you offer volume terms?
- If you are spending enough for the 5.5% to matter, write to sales@finalrouter.com and we will talk. Bringing your own provider keys is the other lever - you pay the provider directly and we charge 5% of list price for the routing.
- Can I get an invoice?
- Tick “I need an invoice” when buying credit and you get a numbered VAT invoice. Every purchase also has a PDF receipt in the billing portal.
- What happens to unspent credit if I leave?
- Deleting your account removes everything hanging off it - keys, logs, stored prompts, the lot. Unspent credit is not refunded, and the confirmation dialog says so before you type the phrase.
How we keep these numbers honest
We bill each provider's list price with nothing added, which means a stale price is not a rounding error - it is us charging you one number while paying another. Five habits keep that from happening quietly.
Every price carries a date
Each model's token prices are recorded with the day somebody read them off the provider's own page. The catalogue shows that date, so a price you are quoted is always a price with a checkable age.
You can price a job before running it
POST /v1/cost/estimate returns what a request would cost, free, on an empty balance - and it runs the same arithmetic the biller runs, so an estimate cannot drift away from the invoice.
A daily drift alarm, honestly scoped
A job compares our catalogue against a community-maintained price map every day and emails on divergence. It covers the models whose upstream entry we have confirmed by hand; the rest are reported as not watched rather than silently assumed correct.
Time-of-day pricing is passed through, not smoothed
Two DeepSeek models cost exactly double inside two UTC windows - 01:00-04:00 and 06:00-10:00 - because that is what the provider charges. We publish both rates and bill whichever applies, rather than quoting a blended number that is never the number.
The catalogue is the same one the router uses
GET /v1/models returns the prices, context windows, capabilities and data policies the gateway actually routes on. There is no separate marketing price list that could disagree with it.
The drift job never changes what you are charged. It reports, and a human confirms - a community list can carry a typo too, and billing from one blind would swap our risk for somebody else's.
One endpoint, and a bill you can check
Start on the free allowance without a card. Add credit when you need the rest of the catalogue.