AI system design modeling: a tokenomics sketch

Ask what an AI system will cost and you often get a single number back: dollars per token, or a line item on a software bill. It's a tidy answer, and it hides most of what decides whether the system is worth running. It says nothing about whether the setup will stay up, how fast it answers, what it cost to tune the model, or who is on the hook when something breaks.
The short version: first check that a setup can run reliably, stay available, be serviced and stay secure. Only then look at cost, and look at it along three separate lines (speed, the cost of training or tuning a model, and the cost of serving answers) that you keep apart and never add into one figure.
This sketch comes from BridgeTek Labs, the research side of BridgeTek, and it is the frame we use when a team asks whether to rent capacity or buy it. BridgeTek is a data science practice. We don't resell the hardware, cloud or licenses we recommend.
A word on the numbers, which we'll say once. The dollar figures and multipliers here are order-of-magnitude scenarios for teaching. None of them is a quote, a bill of materials (BOM, the itemized hardware list behind a purchase) or a measured result for a named buyer, and channel SKUs and buyer-specific costings stay outside this page. Where a number comes from a third party we haven't verified, or rests on assumptions we can't publish, we call it fenced: we show it because it helps the picture, and we hold it at arm's length.
Why one number misleads
"Tokenomics", the economics of paying for AI by the token, usually gets squeezed into a single dollars-per-token figure. That squeezes together three costs that behave very differently:
- Firepower: how fast a setup produces answers, measured in tokens per second (a token is a small chunk of text, a word or part of one).
- Training and adaptation: what it costs to train a small model or tune an existing one to your work. We measure this as rental-equivalent spend, meaning what the GPU hours would cost at rental rates, whether you rented them or not.
- Serving: what it costs to answer requests day to day, which depends heavily on whether you pay a closed model's API or run open-weight models on hardware you own.
A single number also skips the production bar. A cheap option that fails on reliability, availability, serviceability or security is no win at all.
There's a related trap. A volume number, such as tickets per month or requests per day, tells you how much work there is. It says nothing about the hardness mix, the share of easy and hard tasks inside that volume, and it isn't the same as a count of scored eval tasks (an eval is a test suite with a fixed set of tasks and a scoring rule). Public benchmark boards price "assignment cost" per eval task at a stated capability floor. Where the model runs, rented or owned, is a different unit again, and we deal with it last so it never lands on the same dollar axis as the task costs. This article works at the level of a sketch. A Token Economics Sprint replaces these public proxy boards with harness scores on your own work (a harness is the code that runs a model through a fixed set of tasks and scores the results).
Satisfy reliability, serviceability, availability and security first. Then talk $/token. Modeling structures the tradeoffs and surfaces the gap between where you should be and where you are; a measured harness on your own workload is what settles the decision.
Start with the production bar
We call the first gate the production bar, and it has four parts:
- Reliability: it gives correct, consistent results.
- Availability: it's up when people need it.
- Serviceability: someone can monitor, fix and update it.
- Security: it protects your data and your users.
In day-to-day operation, these show up as tracing (a record of what the system did and why), guardrails, human oversight and sandboxes that contain what an AI agent can touch.
Meeting the bar takes people. Owning inference on-premises means having an AI engineering, MLOps (the practice of running models in production) and reliability team. Deciding whether to buy or build is really a question of whether you can support what you build.
The scale ladder
With the production bar in place, we can talk about size. We describe setups on a five-rung scale ladder, from a phone to a fleet of racks. Each rung has a typical job and a typical way of owning it:
| Rung | Typical job | Usual ownership |
|---|---|---|
| Phone / edge | Local assistant; low firepower | Device or personal software subscription |
| Laptop / 1× GPU or Mac unified memory | Local assistant plus small fine-tunes | Workstation, lightly owned |
| 1–few GPUs / small server | Mid-serve: serving a mid-sized model to a team, plus adaptation runs | Owned inference starts here |
| 1 rack | Mid-to-high throughput serving | Owned inference, fuller stack |
| 10–20+ racks | Frontier-class self-hosting (teaching band) | Fuller stack; CapEx plus team at a ~$1–10M/yr scenario |
The ~$1–10M/year and 10–20 racks labels are order-of-magnitude teaching labels. No measured bill of materials or list price stands behind them. (CapEx, short for capital expenditure, is the money that buys hardware up front.)
Figure 1 draws the same ladder as a row of boxes, smallest on the left. Its message: mid-serve is the rung between a laptop and a rack. That highlighted rung, 1 to a few GPUs, is the one we care about most right now, because it's where the 2× 96GB PRO ask sits: the two 96 GB RTX PRO GPUs we've asked to borrow. The 90-day ask calibrates this one rung. It tells us nothing directly about multi-rack builds or the $1–10M band.
Three separate cost axes
Once a setup clears the production bar and has a place on the ladder, we look at cost along three axes, which we label A, B and C. We keep them apart because each answers a different question, and we never put them ahead of the production bar.
Axis A: firepower
Firepower is speed: tokens per second, or how much intelligence you get per unit of time. Its teaching job is to make one point clear to leadership: AI comes in several performance classes.
The contrasts are easy to find in public posts and demos. Specialty stacks post very high token rates (a public demo by Taalas, a specialty hardware maker, reported about 14,000 tokens a second), while consumer-class chat typically runs at about 50–150 tokens a second. For the mid-serve rung, Light Foundry reported about 243 tokens a second (1 Aug 2026) on two RTX PRO 6000 Blackwell cards serving DeepSeek-V4-Flash. That is a third-party figure, which we fence: we haven't measured it ourselves, and none of these numbers is a service-level agreement (SLA) anyone has promised. Keep firepower separate from cloud hosting costs and from vendor performance claims.
Figure 2 shows illustrative firepower for each rung of the ladder. Its message: firepower rises as you move from a phone to a rack. The vertical axis is logarithmic: each labelled step (1, 10, 100, 1,000, 10,000 tokens a second) is ten times the one below, so the jump from bar to bar is far bigger than it looks. The bars run from a few tokens a second on a phone to over 10,000 for a multi-rack fleet. The highlighted bar is the few-GPU rung, where the 90-day ask will test firepower claims.
Axis B: training and adaptation cost
Axis B asks what owned or rented compute buys you when you train a small model from scratch or adapt an existing one. We measure it in rental-equivalent dollars or GPU hours.
One public example: headlines about training runs for Puro-2B, a small model of about two billion parameters, put the rental-equivalent cost at around $4.4k–$6.9k on consumer-class GPUs. That comes from social media and hasn't been independently verified, so it's fenced. We give more weight to figures from paper abstracts when they're cited as such. None of this is a channel quote or a cost worksheet for a named buyer.
Axis C: serving cost
Axis C is the cost of serving tokens: paying for a closed model's API, or serving open-weight models yourself on owned or rented hardware. We use two scenario labels here:
- About 3×: the rough cost reduction class when a workload moves from cloud compute to owned on-premises hardware. This is the general infrastructure story.
- About 10×: the rough reduction class when you move from paying for closed API tokens to serving open-source models on-premises yourself.
These are scenario labels. We haven't measured them here and we attach no absolute dollars to them. They aren't a validated multiplier for any named buyer, and they are never a reason to skip the production bar.
Artificial Analysis offers an outside check on how fast API prices move. Its cost-per-task Pareto view (the set of models where you can't get a higher score without paying more), its inference-price history and its Index-over-time charts all show the cheapest API price per million tokens falling quickly within each intelligence band. That supports a simple point: prices for cloud MaaS (model-as-a-service, hosted models you pay for per use) move quickly. It doesn't turn API prices into steps on the phone-to-rack hardware ladder, which stays on its own axis.
What each rung costs to own
Figure 3 brings the ladder and cost together. Its message: cost jumps by orders of magnitude as you climb the ladder. Each bar is the illustrative CapEx plus opex per year for one rung (opex is operating expenditure, the running costs such as power and people). The vertical axis is logarithmic, with real dollar labels from $100 to $1 million: each step is ten times the one below. The bars climb from a few hundred dollars a year for a phone to millions for a multi-rack fleet, which is the ~$1–10M/yr frontier self-hosting teaching band. The highlighted bar is the 2× 96GB PRO ask on the few-GPU rung.
Choosing how much to own
With the production bar and the three axes in hand, the choice becomes a menu of ownership options. Each option has a different set of things that dominate the decision:
| Ownership option | What it means | What usually dominates |
|---|---|---|
| SaaS / MaaS | Hosted models and agents; you trust the vendor's operations | Production bar (through vendor SLAs) plus Axis C |
| Owned inference | Private serving on your own or a partner's hardware | Production bar (you run operations) plus Axes A and C |
| Fuller stack | Hardware plus adaptation, harness and runtime | Production bar plus Axes B and C, plus enterprise agent frameworks |
In practice we sum each option up on a short posture card. Generic sizing (accelerator count, node count, network and storage bandwidth) maps onto a handful of illustrative classes: edge or workstation (the phone-to-laptop unified-memory band), mid-serve (one to a few accelerator nodes, with measured tokens per second when the evidence exists), rack or on-premises (an OEM-neutral factory shape) and cloud, metered or committed (a hyperscaler-neutral rental posture). Each card states the capacity class, the ownership posture, the inputs that would flip the card to a different answer, and a plain note that it is a teaching card with no quote or procurement list behind it. Commercial quotes are handled separately from the teaching model.
The gap between should and actual
The most useful output of all this modeling is a gap. We compare two things:
- Should: a production-grade posture across all the metrics, with a sensible ownership option and ladder rung for the kind of workload.
- Actual: what we observe. Common examples are a move to the cloud made without a total cost of ownership (TCO) sheet, a vendor switch driven by internal politics, or laptop-class hardware asked to serve a whole fleet.
The gap between them is what the model surfaces. It is input for planning; a finished audit is separate work.
We check the gap on seven things: the four production metrics, then firepower, serve cost, and TCO (whether the buyer has a quantitative picture of what good looks like). In the stylized picture we teach with, the should posture sits above the actual one on all seven. The widest gap is usually on TCO, because many buyers have no quantitative bar for what good looks like, and the narrowest is on firepower.
An interactive version (choose a rung, number of concurrent users, target tokens per second and model size, and get a rough CapEx and opex band) is available separately. You don't need it to follow this article.
What to do next
How do you get from a sketch like this to a plan? We use a progressive estimate, much like a building contractor who starts with a rough sketch, moves to a representative estimate, lays out options and then draws up plans. Figure 4 shows the six steps, left to right. Its message: estimates get tighter as the work moves from idea to build. The highlighted step is the hardware estimate, where a rung on the scale ladder turns into an order-of-magnitude figure.
- Ideation: the pillars and outcomes you want, before any vendor list.
- Architecture: the harness, the runtime and the inference plane.
- Representative hardware: a rung on the scale ladder and an order-of-magnitude estimate.
- Options: on-premises or cloud, and where you sit on the ownership menu.
- Plans: a timeboxed statement of work.
- Build, only after agreement.
Three things hold at every step: production metrics come first, the should-versus-actual gap stays in view, and on-premises ownership needs an AI engineering and MLOps team. None of it is a binding quote.
This is what the Token Economics Sprint does for a named use case. It takes the steps above, replaces public proxy boards with harness scores on your own work, and builds a model of metered, rented and owned inference cost for your use case, side by side, which is handed over to your team. Production comes first, then the three cost axes, then the gap between where you should be and where you are.
Caveats
A few terms in this article are easy to stretch, so here is exactly what each one covers:
- A public-cloud TCO worksheet takes CapEx, power, utilization and GPU-hour inputs for a buyer's purchase order. It's a different thing from an Axis A tokens-per-second demo.
- A tokens-per-second leadership demo illustrates firepower. It is neither a serving-cost multiplier nor a SKU quote.
- Training rental dollars from small-run reports teach training economics. They are no channel bill of materials.
- The 3× and 10× labels are teaching scenarios for Axis C. They are neither measured truth nor dollars ready for an account.
- $1–10M and 10–20 racks are the frontier self-hosting teaching band. There is no measured bill of materials or quote behind them.
- Should versus actual is a modeling output and a way to frame the gap. It stops short of a finished customer audit.
- Dollars per token alone is one slice of the economics and is never enough on its own to settle a production decision.
- Monthly ticket volume is only a count until the hardness mix is defined. It isn't an eval-task assignment cost.
Third-party social posts, preprints and verbal teaching numbers (Taalas, the Puro runs, 3× and 10×, and $1–10M rack economics) have not been independently verified by BridgeTek. They are safe for general education. They don't go into account packs automatically, and no named buyer is the modeled customer. The set doesn't come from live AA v4.3.2, OpenRouter or InferenceX.
All four figures are illustrative teaching charts. Their labels are scenario labels and show no vendor's SKUs.
Sources
- Artificial Analysis data, dated 2026-09-01 / 2026-09-02
- Light Foundry, third-party report, 1 Aug 2026: about 243 tokens a second on two RTX PRO 6000 Blackwell cards serving DeepSeek-V4-Flash (fenced; third-party measurement)
- Taalas, public demo: about 14,000 tokens a second on specialty hardware (fenced; a social and demo figure that hasn't been independently verified)
- Puro-2B training-run headlines: rental-equivalent cost of about $4.4k–$6.9k on consumer-class GPUs (fenced; from social media and unverified)
About the author

Jason Larkin
Consultant, BridgeTek professional services
Jason Larkin models and predicts how complex systems behave, from a back-of-the-envelope estimate to a full simulation. He applies that to the cost and design of AI systems in his writing for BridgeTek Labs, and to quantum and other emerging computing in his consulting work.
