Three Dollars Became Four

Three dollars struck through, four dollars in the fox gradient – with the real hardware sums beneath: 8 × MI300X at $14.80 an hour published 25 July, 8 × MI355X at $21.20 three days later

My feed keeps handing me the same post. Some open model beats Fable, or you can run one for free on your own servers, and here is a trick for wiring it into Claude Code. The posts are cheerful and they are almost never costed. So one day in July I costed one, out loud, in public – and a few days later I had to post a correction underneath it. Everything below carries a date. It is all my own arithmetic, done at particular moments in June and July 2026, at the prices of that week, and a year from now most of it will be wrong. That is not a defect of the exercise. That is the exercise.

What “free” weighs

Start with the parameter count, because everything else follows from it. Kimi K3 is 2.8 trillion parameters. Running that requires around 1.6 TB of memory, leaving some spare room for cache and ops. Your computer, most of the time, has a hundred times less than that.

So you do not run it. You rent the machine that runs it.

I went and found one of the very inexpensive ones: an AMD MI300X cluster, eight GPUs at $1.85 each, $14.80 an hour. You do not take a cluster like that by the hour. You take it by the month, and the month is $10,656.

Now make it pay for itself. At K3’s list price of $15 per million output tokens, and with input and output normally totalling 10/90, you need to move something like 640 million output tokens a month just to get level. The cluster’s ceiling, running 24/7, is around 4 billion tokens – GPU, bandwidth and architecture each take their cut before you see a token.

Which means that at 100% utilisation you can drive the cost of running Kimi K3 down to about $3 per million output tokens.

That was the figure I published on 25 July.

Then the weights actually dropped

I had posted it 1–2 days before the real weights of K3 came out. Then Vultr deployed it, with a page describing the actual hardware, so I redid the sum in the comments under my own post.

The first difference is the card. MI355X, not MI300X: $2.65 a GPU, eight of them, $21.20 an hour. The second is throughput. They quote up to 8k tokens per second, but that is blended – you have to open the table and read the output column. Maximum output is 3K per second; realistic, at 128 parallel requests, is 1.5K per second. That is 5.4 million output tokens an hour, about 3.8 billion in a month. $21.20 buys those 5.4 million.

So the floor is roughly $4 per million output tokens, not $3.

Not much off, is what I wrote, and I still think that. But it was my number in front of other people and it had moved by a third, so it wanted saying. And a floor is only a floor: it assumes the cluster is busy every hour of every day, which does not happen.

That price is fine when you genuinely need that model and there are no options – cyber security, defence, biology. For day to day work a $200 Claude Max plan beats it by a margin. In theory even a $20 Pro plan would beat it, given optimised utilisation and a scheduler that knows your limits and works around them.


The other side of the sum

Two days before the Kimi post I had put the same point crudely. On a $200 Claude plan I had run up more than $3,000 of usage at list API prices, counting only Opus and Fable output tokens. I was not really trying, and I finished almost everything in my pipeline. The effective discount from list price on a $200 plan is 10x at least.

Which is why I get short with the wire-a-cheaper-model-into-your-editor genre. I am using Fable at $5 a million in Claude Code, with projects and memory set up properly, without any somersaults.

Then I sat down and did the full sum, and it came out lower still. On the $200 plan, in a month, I can use around 50 million tokens of each model – 100 million blended. That is $2 per million blended. Against the list prices, that works out at roughly $2.70 per million for Fable output, where list is $50, and $1.35 for Opus, where list is $25.

Under $3, in other words. Under the best theoretical case of a rented cluster running flat out.

The floor under the floor

What still surprises me is that free tiers exist at all.

Generating on the most capable Claude model needs a GPU setup that would be at least $150K – in the original comment I wrote $5K, and that was simply wrong. Amortise the real number over a year and you get around $17 an hour of pure hardware, which is the absolute minimum cost of anything at all happening. On Pro, at $20 a month, I was maxing out around 20 hours a week. The compute alone is worth around $1,500 a month.

So the future of this is one of two things. Either a less expensive way to run them is found, or they become a thing of a few. And those few would be able to do a great deal.

Gym passes

Prepaid plans are a bit like gym monthly passes. You have one, I have one, our friend Joe has one. But only I am going six days a week, every week, and you are still considering your programme and your schedule.

Now imagine there is a robot that will exercise for you. You get the benefit, and it is included in your pass. Once you and Joe work that out, you will be at the gym six days a week as well, and so will everyone around you.

That is what autonomous agents are. Another month or two and everyone will notice that you do not need to sit in front of the computer prompting. The agent checks what limit you have left, launches itself, does the work.

This is why I think prepaid plans are doomed. One day or another, depending on adoption. Use it while it lasts and finish your projects.

I can put a number on my own side of that. I spent over a year training myself to use tokens efficiently, squeezing every drop out of a prepaid Pro plan; most weeks I was capped on the weekly limit and I hit the 5-hour limit many, many times. Then, in early July, a week on Max 20x with Fable 5 back. I ran 3–5 projects at once including night ops, with scheduling and an auto-scheduling cron list. Not once did I reach the 5-hour limit. By the end of the working week I was at about 40% of the weekly limit, and it took launching a heavy hobby project – the Zettelkasten – to reach 87% before the Sunday reset. Two long-running projects were finished that week, set to launch the Monday after.

So the robot already exercises for me, and I still could not empty the pass. That gap is where the discount lives. It closes when everybody has the robot.

Big sharks

In June people started saying that AI usage had become extremely expensive; GitHub’s usage limits were the example going round. It was obvious this would happen. First-party providers – OpenAI with Codex, Google with Antigravity, Anthropic with Claude Code – are for now less expensive than Cursor or Copilot. An average dev on Cursor is what, $600–700 a month at best? Claude Code Max at $200 beats that by a margin, in usage available and in cost per token.

The intermediaries get pushed out, gently. First-party providers these days are big sharks eating smaller sharks, and whatever great idea you have for using their API, as soon as it is successful it will be implemented inside their own systems.

There is one cost the arithmetic does not show. To collect that discount you have to leave Cursor for Codex or Claude Code, then learn the instrument and stick with it. That is a long-term decision, not easily changed if the market moves. I asked it in someone else’s comments in June and I have not answered it since: what would, and what should, an average firm choose?


Half a white bread and a bottle of kefir

There was a footnote on a post I wrote at the end of June. Claude Opus is reportedly between 1 and 5 trillion parameters, and running it reportedly needs a rack with 8 × MI300X, which would land at $150K at least.

Set that beside the other thing I wrote that day. I had done a 15 km cycling run, and my watch told me I had spent 300 calories. Inside those 300 calories was every complex motor task the ride demanded, and also the next steps of an app I was building, planned while riding. Half a white bread and a bottle of kefir buys three runs like that. A truly formidable achievement.

The most comprehensive AI ever made was finished quite some time ago, and it has the most powerful hardware behind it. Only one book was ever written about its creation, and a good many people believe that book is not a true story.

No 150K setups. Invest in yourself.

For the average Joe – which is me, most days – the advice has not changed since June. Choose one instrument, either GPT or Claude, take the $20 paid plan, and learn to use it every day until you have mastered it. Do not read AI news; it is no news. The smallest plan is the biggest bang for the buck, for your own learning and for keeping some balance between the hours you spend with the machine and the hours you do not.

What that $20 buys is harder to price than tokens. It more than doubles my development capacity – I would say triples it. And it gives me access to things outside my comprehension at any given moment, so that instead of spending hours learning the one thing I need for some small bridge, I build the bridge.

Every figure above has a date attached, and the dates sit close together on purpose. Three dollars became four in three days because a company published a page describing real hardware. The plans will be repriced. The gym will fill up. Someone will have to do all of these sums again with different rentals and different models, and the arithmetic will not be difficult.

The hardware for all of it, in my case, is a laptop with the lid shut, screen off, sleep disabled, sitting on the power cord. I reach it over Tailscale, work in zellij, move around the files with lf. No rack of mine. It is working through the night on a plan I have already paid for.


Sources – my own LinkedIn posts, oldest first:

The corrected K3 figures, the GPU amortisation, the switch-cost question and the closed-lid laptop come from my own comments under those posts and others, June and July 2026.

$ exit 0 – thanks for reading

The fox will keep the drafts warm.