85% reduced costs, 95% quality
I built the same financial model three times in one session.
Once using Claude Fable 5, once in Claude Opus, once in Claude Sonnet. To see what changes when you use a ‘more intelligent’ model.
I was surprised that the quality didn't drop. Opus matched Fable, and Sonnet came close – both for less than half the usage.
My conclusion: Fable 5 is useless if you don't use Claude Code. So, the model everyone's rushing to try is built for a job most finance teams never do (and will end up costing you).
I'm not alone in finding this. UC Berkeley's LMSYS lab built RouteLLM. It sent easy prompts to a cheap model, hard ones to the expensive one. It cut costs over 85% whilst keeping the quality at 95%.
Most finance pros get a bad response, turning settings to max, or run everything in Cowork or Work (because it’s the latest tool). But, as pricing goes consumption-based, this becomes a big problem.
So, to save you a lot of future spend, here’re my 4 principles for keeping your AI usage under control.
Tool vs Model vs Effort vs Thinking (4 Principles)
Tool
There are now 2 main places to work in most AI tools. A ‘Chat’ interface (what you’re used to) and a ‘Work/Cowork’ interface. Here’s what they’re called in each tool.
Claude – Chat and Cowork
|
|
ChatGPT – Chat and Work
|
|
Microsoft Copilot – Copilot and Copilot Cowork
|
|
‘Chat’ is a back-and-forth interface: you ask something, it answers, you react, and you steer it one step at a time. It's built for collaboration, and you’re heavily in the loop.
‘Work / Cowork’ is a handoff interface: you set the instructions – give it a goal, a folder of files, or a set of connected apps – then step away. It works through the task on its own, sometimes for several minutes, and only comes back with a result, or a question/approval if it gets stuck. It’s built for execution, and you are in the loop much less.
Principle 1
Use Chat to plan and collaborate. Use Cowork / Work to handle repeat tasks.
If I am doing some ad-hoc analysis, where there is no defined steps. Then I will just chat to AI.
If I am doing a well defined repeat task (e.g creating a company branded month-end report from a data file) then I will use Cowork / Work.
NOTE – With Microsoft Copilot, you need a Copilot license, and the Cowork capability is ‘Metered’ meaning you pay for consumption. Like with electricity. Claude Fable is also a consumption based model, meaning it’s not included within your normal plan usage.
Model
I like to think of ‘model’ as a base level of intelligence.
Like with your finance team. You have junior analysts, AR/AP clerks, controllers, Directors & CFOs.
This is not too much different to AI.
If the problem is easy -> Give it to the analyst
If the problem is super hard and requires a lot of experience -> Only a CFO can do it.
Principle 2
Do your first build or workflow with a more intelligent model.
Then use a less intelligent model to run it once you’ve built it.
Most of your day to day tasks should use middle models (Claude Sonnet, ChatGPT Terra, Gemini Thinking) unless something is super simple, where you can go down another model – providing it keeps the same quality level.
Effort
Effort defines how hard the model works.
The simplest comparison is time taken. With your team, there is a big difference between a fast answer (what’s your gut feel on this – you’ve got 30 seconds) and a slow answer (take this away and come back to me tomorrow).
Principle 3
Use the default effort setting for each model (normally ‘Medium’) and experiment with lowering the effort for tasks that require less accuracy, and increasing the effort for tasks that require more.
Thinking
This is an option that is only available in Claude and now Gemini. You can access it from the ‘Effort’ menu in Claude, and the model selection menu in Gemini under ‘Thinking Level’.
|
|
Note – Google are still rolling out this capability, so you might not see it just yet.
When you turn ‘Thinking’ on it becomes better at breaking down tasks into logical steps.
Good use cases for this are things like complex data analysis, or producing documentation when you’ve got a lot of data and context.
Principle 4
Only turn on ‘Thinking’ for tasks where there is benefit in breaking them down step by step.
The One Thing to Remember
![]() |
90% of your work should be with mid-intelligence models at default effort levels.
With 10% of your work reserved for smarter models. But even then, you might not need to use them all the time once you’ve built your workflow.
And as I said at the beginning. Right now, I don’t think any finance team needs to use Fable 5.
So, share this with your team, and start getting your AI usage under control before consumption based billing becomes standard for every tool.
Best,
Your AI Finance Expert,
– Nicolas
P.S. – The best way to learn from me directly is to join my weekly masterclasses. 60-mins + Q&A where I’ll get all your AI in Finance questions answered. Join me here.
P.P.S – I ran real finance work through 4 Claude models so you don't have to. Here's what I found → I Ran Real Finance Work Through 4 Claude Models

.png)