Your Chatbot Flunks Money. A 10,000-Answer Test Explains Why
A UK fintech firm ran 121 personal finance questions through 18 AI models and got roughly 10,000 answers back. Mainstream models failed 57% of them, and that number hit 88% once the questions got harder. Here's why that matters more than the usual benchmark noise.
I'll say this plainly. You shouldn't trust a chatbot with your money, and the numbers out of Saturn make that a lot harder to argue with.
The UK fintech firm ran 121 personal finance questions through 18 free and paid AI models, which works out to roughly 10,000 individual answers. Mainstream models got 57% of them wrong. Not slightly off, not missing a caveat. Wrong. And when the questions got harder, the multi-step stuff like comparing mortgage options or working out tax treatment across several accounts, the failure rate climbed to 88%.
So the next time a friend tells you they asked ChatGPT whether to pay down the credit card or invest the difference, you've permission to wince.
What the Test Actually Found
The setup was clean. Same question set, same scoring, every model. That's the part I care about. Most AI benchmarks get gamed within weeks, and finance is a category where a model can sound completely confident while being flatly wrong.
The 57% headline is bad enough. The 88% number on multi-step queries is the one that should worry you.
Basic questions are fine. What's a Roth IRA, how does compounding work, why do fees matter. Models have read the entire internet on those, and they'll answer clearly. The trouble starts when the correct answer depends on your income, your state, your filing status, and the order you do things in. That's not trivia. That's personal finance.
And the timing matters. Consumer reliance on chatbots for money questions has grown sharply over the past couple of years, which means a lot of people are already acting on answers they never verified. I'd guess the real-world error rate runs worse than the lab number, because in a lab nobody panics and abandons the conversation halfway through.
The Case for the Other Side
Granted, the skeptics will say benchmarks are artificial, and they're partly right. A 121-question set is small. Scoring open-ended finance answers is subjective, and reasonable people can disagree about where the line sits between a wrong answer and an incomplete one. Some slice of that 57% might be technically correct answers with the wrong caveat attached.
To be fair, the models are also improving fast. A study like this has a shelf life of maybe six months before the numbers move. And there's a real argument that a chatbot getting 43% of personal finance questions right is still a net positive for the millions of people who'd otherwise ask nobody at all.
But here's the thing. In finance, a wrong answer isn't neutral. It costs money. Sometimes it costs a lot of it.
Where I Land
I'm not entirely convinced by the "it's just a benchmark" defense. The 88% failure rate on harder queries isn't a measurement artifact. It's structural. Models are good at retrieving facts and bad at reasoning through a chain of decisions where every step depends on the one before it. Personal finance is almost entirely that kind of problem, which is a bad matchup.
The question worth asking: if a tool fails nearly nine times out of ten on the questions people actually need help with, what's it for? Use it to learn vocabulary. Use it to draft questions for a real advisor. Don't hand it the decision.
What I'd watch next is whether any of those 18 models publishes a follow-up showing real gains on multi-step reasoning. If one does, that's a signal. Until then, the burden of proof sits with the chatbots, not the skeptics. Your money's a bad place to run the experiment.