Route by the shape of the work, not by how important it feels. Pick the lane first, then the effort level. Real agency jobs in each lane.
Work with no judgement in it. Fast, cheap, and it never gets tired on row 400.
Reading a lot and handing back less. Close to Opus on quality, under half the price.
Anything a client pays for, or anything that runs in a loop where small errors compound.
Twice the price of Opus. Reach for it on purpose, on calls you'd take to a board.
Every model takes an effort level: low, medium, high, xhigh, max. It moves cost and quality more than switching model does.
When something feels expensive, drop the effort before you drop the model. You usually keep the quality and lose the cost.
The board above is the answer. This is the reasoning under it: why a long chat costs what it does, what the free plan actually gives you, six habits that cut usage on any plan, and an honest test for when to start paying.
Claude doesn't bill you per message. It bills per token — and every turn re-reads the entire conversation before it answers. Turn five re-reads four turns. Turn fifty re-reads forty-nine.
So a long chat doesn't cost twice what a short one costs. It costs many times more, and the curve gets steeper the longer you stay in it.
That single fact explains nearly everything below. Every rule here is a way of not paying to re-read your own conversation.
More than most people assume:
Free is enough for words in, words out — drafting, rewriting, summarising, thinking a decision through. It is not enough for building. If your work is writing, free will carry you further than you expect. If your work touches files, deployments or a live browser, it won't carry you at all.
Opus at low effort can cost less and read better than Sonnet at high. Effort moves the number more than the model name does. That's why every lane on the board carries an effort level with it.
The per-million prices up there are API rates. On a subscription you don't pay them directly, but they tell you what each choice is worth.
The tempting next step is an AI call that decides which AI to call. Skip it. It adds cost, adds delay, and adds a new thing that can be wrong.
Fixed rules on task shape get you most of the benefit at none of the overhead. Better still, bake the choice into the tool itself — set the model and effort once, on the agent that does that kind of work, so there's no decision left to make at the moment someone is busy.
You can apply it without reading the prompt. If you have to read the work to know where to send it, the rule isn't finished.
Use free for a week on your actual work, not on trying it out.
If you hit the limit regularly, that's the product telling you you've found a real use. Upgrade. If you don't hit it, don't upgrade yet — and don't let anyone tell you otherwise.
What the paid plan buys isn't simply more Claude. It's the connected setup: your tools, your files, your context carried from one session to the next.
I build routing, agents and workflows for agencies and service businesses — the kind that cut a team's AI bill and raise the quality of what comes out of it. I'll show you live rather than talk you through a deck.
Tell me your worst workflow. I'll build the fix on your own tools — in person across London or over a call — and show you what it saves before you spend anything. Keep it, and it's yours.
Book your build sessionTakes 2 minutes · no commitment · prefer email? andrei@andreiachim.com
Not ready to talk? Get the AI Admin Audit — 10 questions that find your most expensive workflow.