Cost, intelligence and walltime: one datapoint, many access paths

Picture three teams inside one company. The first pays $20 a month per person for a chat subscription. The second calls a model through an API key and gets a bill that rises and falls with use. The third has borrowed two GPUs for 90 days and runs an open model on its own hardware. Ask each team what its AI costs and you get three answers that can't be lined up side by side. The quality of answer differs too: the subscription team's model sits around 30 on the Intelligence Index (a fenced estimate), while the API team and the borrowed hardware both run a model that scores 51.8.
The short version: the same level of AI capability can be bought in very different ways, and the price tag is only one of three things that change. We track cost, capability and speed as three separate measurements, and we suggest you do the same before choosing how your team gets its AI.
This article comes from BridgeTek Labs, the research side of BridgeTek. BridgeTek is a data science practice: we structure compute economics for enterprise buyers. We don't resell the hardware, cloud or licenses we recommend.
A word on how to read the numbers, which we'll say once. Everything here is a teaching scenario built from one working dataset of 16 access options, collected in September 2026. It shows how the options compare with each other. It doesn't quote a price, promise a service level or list channel SKUs for your account. List prices come from the vendors' public help pages. Where a number is a rough band, a stand-in, a third-party report or a cost nobody has published, we call it fenced: we show it because it helps the picture, and we hold it at arm's length. Read a fenced number as a rough guide and don't quote it as fact.
Why the price tag misleads
Most teams pick an access path by habit: a $20-a-month copilot, an API key, or a short-term GPU loan. Each one feels like "getting AI". Underneath, three things move independently of each other.
The first is compute cost: what you pay, whether that's a subscription, a metered API bill, or the capital expenditure (CapEx, money spent up front on hardware) and running costs of machines you own or borrow. The second is capability: how good the model's answers are. The third is walltime: the time you actually spend waiting and the time you actually get, which covers response latency, speed in tokens per second (a token is a small chunk of text, a word or part of one), queues, rate limits, and how long you have access at all.
Once you look at all three, a few common assumptions start to wobble:
- Subscriptions hide the true cost per token. Heavy use can end up costing more than pay-as-you-go would, and usage caps and queues are rarely spelled out.
- Borrowed hardware only looks free. A loan window still carries electricity, operations, calibration time and the logistics of sending the hardware back, and public materials seldom put a dollar figure on any of it.
- Capability scores depend on the test. A model's score under one test setup doesn't automatically carry over to a different one.
- A 90-day window is still a limit. Throughput, how many requests can run at once, and queue behaviour during the window all still need measuring.
- Free-tier limits move. They change by region, by account and whenever the product changes, so precise quotes are unsafe.
Subscription is rented intelligence plus a hidden queue. Owned mid-serve is walltime you can calibrate plus ops you own. Access paths differ in capability as well as cost: Gemini free sits near 15 on the Index, Plus near 30, the 90-day loan and the DeepSeek V4 Flash API at 51.8, and Pro $200 at 60.9. Each is a different ownership and walltime posture, and none of them is simply the best point.
Three measurements for every option
So for every access option in the dataset we record three things and keep them apart.
Cost. For a subscription, this is the public monthly price. For an API, it's the variable bill, driven by the price per million tokens. For hardware you own or borrow, it's CapEx or loan terms plus running costs. We label these separately because they are different kinds of money, and the hardware dollars are often fenced.
Intelligence. Where a plan maps to a known model, we use the Artificial Analysis (AA) Intelligence Index. Artificial Analysis runs a public set of evals (standard test suites, each a fixed set of tasks with a scoring rule) and combines a model's results into one number on a 0–100 scale, where higher means more capable. In this dataset, scores run from roughly 10–20 for free consumer tiers to the low 60s for the strongest paid plans and frontier APIs. As a rough guide, a gap of a few points separates close peers, and the jump from around 30 to 60 separates everyday consumer chat from the frontier models in this set. Scores also shift when the test setup changes, so treat a small gap as a hint. When we can't map a consumer product to a scored model, we give a fenced band instead.
Walltime. Latency, tokens per second, queues and waits, rate limits, and the access window. Ninety days of borrowed hardware is a real limit on walltime, however fast the hardware runs.
For short, we write each option as a single datapoint with three parts: D = (C, I, W), for cost, intelligence and walltime. We keep the three apart on purpose. Blend them into a single "value" score and a cheap, slow option can look identical to an expensive, fast one.
One rule comes before all three. An option first has to meet the production bar: reliability, availability, serviceability and security. A low cost number means little if the option fails there.
The options side by side
Here is the working dataset in one table. The intelligence column shows the Index where one exists and a fenced band where it doesn't. The last rows are hardware options: our 90-day calibration anchor (more on that below) and two illustrative rack sizes.
| Access option | Cost | Intelligence | Walltime |
|---|---|---|---|
| Google AI / Gemini free | $0/mo | Fenced ~10–20 (Flash-Lite API proxy; plotted near 15) | Variable caps and queue; no SLA (service-level agreement) |
| ChatGPT free | $0/mo | Fenced ~10–30 (plotted near 20) | Queue, and downgrades under load |
| ChatGPT Go (where sold) | ~$8/mo | Measured 18, just under ChatGPT free | Lower caps than Plus |
| ChatGPT Plus / Claude Pro | $20/mo | Fenced ~20–40 (consumer proxy; plotted near 30) | Session limits; hidden queue at peak |
| ChatGPT Pro $100 | $100/mo | 60.9 | 5× Plus usage; vendor queue |
| Claude Max 5× | $100/mo | Fenced ~58 (consumer proxy) | Higher usage tier; vendor queue |
| ChatGPT Pro $200 | $200/mo | 60.9, the same model as the $100 tier | 20× Plus usage; vendor queue |
| Claude Max 20× | $200/mo | 63.1 (Opus 5, max) | 20× usage; abuse guardrails |
| API pay-as-you-go, mid band | Variable; the two mid-band APIs here cost ~$0.23 and ~$0.63 per million tokens | Measured 51.6–51.8 | Provider latency tiers; $/token visible |
| API: DeepSeek V4 Flash 0731 | ~$0.23 per million tokens | 51.8 | Provider latency |
| API pay-as-you-go, frontier (60 and up) | Variable; the frontier API here costs ~$7.70 per million tokens | Measured 62.1 | Enterprise contracts separate |
| Enterprise / Team seat | ~$25–30/seat/mo+ (illustrative) | Product-dependent | Single sign-on and admin; negotiated limits |
| 90-day loan: 2× RTX PRO 6000 (calibration anchor) | CapEx/loan; $ not disclosed; 90-day access | 51.8 (DeepSeek V4 Flash 0731, max, on this hardware class) | 90-day window; ~243 tok/s reported by Light Foundry, 1 Aug 2026, with up to four requests at once; our own lab calibration still to come |
| 1 rack owned (illustrative) | CapEx + power + ops ($ fenced) | Illustrative 40–55 serve class | Owned walltime at planned load |
| Multi-rack frontier (teaching band) | ~$1–10M/yr scenario, order of magnitude only (no bill of materials behind it) | Illustrative 55–65+ | Fleet throughput and engineering to service-level objectives |
Two rows surprise people. ChatGPT Pro $200 scores the same 60.9 as ChatGPT Pro $100: the extra $100 buys more usage of the same model, and the model itself is no smarter. The Claude rows use the dataset's own values, fenced ~58 for Max 5× and a measured 63.1 for Max 20× (Opus 5, max), and we don't average them with the ChatGPT rows.
Same capability, different bills
The clearest way to see the problem is to hold intelligence steady. Look at the options that land near an Index of 52.
The DeepSeek V4 Flash 0731 API scores 51.8 and costs about $0.23 per million tokens, with walltime set by the provider's latency. The same model on borrowed hardware (two RTX PRO 6000 cards) also scores 51.8. There, the walltime story is about 243 tokens a second as reported by Light Foundry (with up to four requests at once) for a 90-day window, and the cost is a loan or CapEx figure nobody has published. For comparison, Gemini free (a fenced ~15) and ChatGPT Plus (~30) sit well below that band, and ChatGPT Pro $200 sits above it at 60.9, the same model as the $100 tier.
Matching intelligence on owned hardware changes walltime: you gain throughput, you control how many requests run at once, and you can calibrate privately. It also moves cost from a subscription bill to CapEx or a loan, plus the operations to run it. Neither path is free.
Why a 90-day loan is in the dataset
The hardware row is our calibration anchor, the 2× 96GB PRO ask: two NVIDIA RTX PRO 6000 Blackwell cards with 96 GB each (192 GB in total), for 90 days. The ask itself is documented; the dollar cost of this kind of loan appears nowhere in public materials, so it stays fenced.
For intelligence, we map the anchor to DeepSeek V4 Flash 0731 (max), which scores 51.77 on the Artificial Analysis manifest dated 2026-09-01 (51.8 when rounded), inside the 50–60 band. For speed, the best evidence today is a third-party report from Light Foundry (1 Aug 2026): about 243 tokens a second on a single stream, with higher aggregate throughput when a few requests run together. We haven't measured it in our own lab yet, so we rate it moderate confidence.
What the 90 days buy is time to calibrate our progressive estimates on real hardware: serve the model, evaluate it and publish the results. The anchor stands for one mid-sized serving setup and nothing larger. It says nothing about a multi-rack build or the $1–10M/yr frontier band, it involves no hardware resale, and it doesn't replace anyone's subscription.
Capability by access path
Figure 1 starts the story with how capable each access path is. Its title carries the message: a 90-day loan scores about 9 points below the 200 USD plans.
Each bar is one access path, and its height is its Intelligence Index. The other bars are subscriptions, labelled by name and monthly price: Plus at $20, Pro and Max at $100, Pro and Max at $200. "Max" here is Claude Max, so Max 100 USD is Max 5× and Max 200 USD is Max 20×. The highlighted bar is the 90-day loan running DeepSeek V4 Flash 0731, at 51.8. That puts it in the 50–60 band: about 9 points below the two Pro plans at 60.9 (9.1 below Pro 200 USD) and 11.3 below Max 200 USD at 63.1. Max 100 USD, at 58, sits in the same 50–60 band as the loan. Plus, at 30, is a fenced consumer proxy, and so is Max 100 USD at 58.
The two Pro bars are the same height, 60.9. That's the point from the table above: the $200 plan buys more usage of the same model.
Responsiveness by access path
Capability tells you how good the answers are. The next question is how quickly you get them, and how often you're kept waiting. Figure 2 answers it with one message: owned mid-serve is the most responsive option on this ladder.
Each bar is an access path, and its height is a relative responsiveness score from 0 to 10, where higher is better. It's a teaching score for walltime, built from the same working dataset, that folds together queueing, rate limits and throughput. It isn't a benchmark, and the points are only meaningful relative to each other. The subscriptions sit lower because their users share the vendor's queue and work within usage caps: Plus at $20 scores 3, and Pro and Max at $200 both score 5. The highlighted bar is the 90-day loan, at 10. It scores highest because the hardware is dedicated to you, with no shared vendor queue, and because the one speed report we have for it, the ~243 tokens a second from Light Foundry, is third-party evidence that it serves quickly.
Cost at the same capability
Now hold capability steady and look at the bill. Figure 3 takes three paths that all score about the same and compares what each costs per month. Its message: at about the same capability, the monthly bill depends on the access path and on how much you use it.
The other two bars are APIs. The DeepSeek V4 Flash API, at Index 51.8, comes to about $25 a month, and the Gemini 3.6 API, at Index 51.6, to about $32 a month. Both monthly figures come from assumed usage, and the volumes differ: about 109M tokens a month for the DeepSeek API and about 51M for Gemini. So the two bars don't rank the APIs by price; an API bill scales with how much you use it. The highlighted bar is the 90-day loan, running the same DeepSeek model at 51.8, and it reads $0 a month. The chart draws it as a thin sliver so you can see it's there.
That $0 is the metered cost of using hardware someone has lent you. It leaves out the loan's CapEx, the power to run it and the people to operate it, none of which has been disclosed. So the highlighted bar shows what you pay per use. It doesn't show that the loan is the cheapest path overall.
What 90 days cost, day by day
A buyer weighing a 90-day hardware loan against subscriptions usually asks one more question: what does each option cost per day across that window? Figure 4 answers it.
Each horizontal bar is an option, and its length is the cost per day, taking a month as 30 days. A $20 subscription comes to about $0.67 a day, $100 to about $3.33, and $200 to about $6.67, whichever vendor sells it. The highlighted bar at the top is the hardware loan, at $0 a day on the meter.
The same caveat applies as in Figure 3. The meter reads zero because the loan's CapEx, power and operations sit outside it. Counting those unlike costs changes the picture, so the short highlighted bar doesn't make the loan the cheapest option.
What to do with this
When you next compare ways to give a team access to AI, write down three numbers for each option: what it costs, which capability band it reaches, and how much walltime you actually get. Check the production bar first. Then compare options inside the same capability band, where the differences in cost and walltime are real differences you can act on.
Our companion article, AI system design modeling: a tokenomics sketch, steps back from single options to the whole range of setups, from a phone to a multi-rack fleet. It sets out the production bar and the separate cost measures we use when a team is deciding whether to rent or own.
When you need these numbers for your own workload, that's the job of the Token Economics Sprint. It starts from public benchmarks like the ones here and moves to measured harness scores on your work (a harness is the code that runs a model through a fixed set of tasks and scores the results). You get a model of metered, rented and owned inference cost for your use case, side by side, handed over to your team.
Caveats
We label every number in this article in one of three ways:
| Label | What it covers here |
|---|---|
| Measured | Index values from the public Artificial Analysis manifest; public subscription list prices; the anchor hardware and its 90-day duration; third-party tokens-per-second reports (measured by others, outside BridgeTek's lab) |
| Illustrative | Rack and multi-rack dollar bands; enterprise seat ranges; API dollars-per-month equivalents from assumed usage |
| Fenced | Free-tier limits; the mapping of consumer products to an Index value; loan and rack dollars; internal lab numbers we haven't reproduced |
All four figures are teaching charts, built from one 16-row working dataset. It samples the market and doesn't survey all of it, and it doesn't sweep every plan and API live. The Index values come from an Artificial Analysis extract dated 2026-09-01, which predates the current v4.3.2 release; it is a snapshot and doesn't update. The hardware options have no public monthly price. Subscription prices come from public list pages as of September 2026. The API dollars-per-month equivalents come from assumed usage and are estimates of spend. The set doesn't come from live AA v4.3.2, OpenRouter or InferenceX.
Sources
- OpenAI Help: ChatGPT Plus
- OpenAI Help: Pro tiers
- Claude Help: Pro plan
- Claude Help: Max plan
- Google Gemini
- Artificial Analysis Intelligence Index and API pricing data, dated 2026-09-01
About the author

Jason Larkin
Consultant, BridgeTek professional services
Jason Larkin models and predicts how complex systems behave, from a back-of-the-envelope estimate to a full simulation. He applies that to the cost and design of AI systems in his writing for BridgeTek Labs, and to quantum and other emerging computing in his consulting work.
More from BridgeTek Labs
- AI system design modeling: a tokenomics sketch
Production metrics first, then firepower, adaptation cost, and serve economics, without folding unlike numbers into one dollars-per-token headline.
BridgeTek Labs · AI system design ·
