# Start with the smallest model. Move up only when it fails.

By Stanford’s count, the price of a fixed level of AI performance fell more than 280-fold in two years. For sorting work, a small model trained on your own cases can beat a large general one.

Tyler Gibbs. 30 September 2026. 4 minute read.

A bank wants each flagged transaction sorted into one of three bins: fraud, abuse, or neither. The team’s first move is usually to send every alert to the largest AI model they can buy. It works in the demo. Then the bill arrives, each answer takes several seconds, and someone asks why a three-way sort needs the most expensive model on the market.

The largest model is a reasonable place to find out whether a task can be done at all. It is rarely the right place to leave it.

## The same capability costs a fraction of what it did.

Stanford’s 2025 AI Index tracked the price of a fixed level of performance. In November 2022, a model at that level cost $20.00 per million tokens. By October 2024 a model matching it cost $0.07, a drop the report puts at more than 280-fold.

Models shrank along with the price. In 2022, the smallest model to pass 60% on a standard knowledge test had 540 billion parameters. By 2024 a model with 3.8 billion did it, which the Index calls a 142-fold reduction. A model that size can run on hardware you own.

## For sorting, a small trained model beats a large general one.

Much of the work in a regulated office is sorting: this alert is fraud or it isn’t, this document is privileged or it isn’t, this referral is complete or it isn’t. Researchers have tested whether that kind of task needs a large model.

Edwards and Camacho-Collados ran 16 classification datasets and reported at a 2024 language conference that “fine-tuning smaller and more efficient language models can still outperform few-shot approaches” with larger ones. Bucher and Martini compared small models trained on task data against the leading general models of 2024 and found the trained ones ahead “in all cases.” Theirs is a preprint and has not been peer reviewed.

Both results carry the same condition. The small model wins after it has been trained on labelled examples of the task, and both papers compared it against large models that were only prompted. If you have years of decided cases, you have those examples. If you don’t, the large model is your starting point.

## Send each task to the cheapest model that passes.

You do not have to pick one model for everything. In 2023 three Stanford researchers tested a cascade: try a cheap model first, and pass the question up only when the cheap answer looks unreliable. In their best case, on a set of financial news headlines, the cascade matched the best single model at up to 98% lower cost.

AT&T described the same idea at company scale in July 2026. It runs an average of 45 billion tokens a day through a gateway that matches each task to the most cost-effective model, and says it is reducing AI costs “as much as 90%.” That is AT&T’s own figure, published without a baseline. Its other remark is the more useful one: only a small percentage of its tasks need the most capable models.

## What this looks like on one task.

For a bank, we fine-tuned a language model on the bank’s own data to classify each instance as fraud, abuse, or neither. The model does one job, it was trained on that bank’s records, and it runs where the bank chooses.

The order of work matters more than the model. Write down the task and the answers it can have. Collect cases your people have already decided. Test the smallest model that might work against those cases and against how your team does it today. Move up a size only when the small one fails, and keep the large model for the questions that have no fixed set of answers.

## Size is something you measure.

Some work needs the most capable model you can get: reading a long contract for what is unusual about it, or drafting a letter. A small model sent to do that work will fail quietly. So test for size on your own files, the way you would test anything else you buy. Most teams never run the smaller test, and pay the larger bill for years.

## Sources

- [Stanford HAI, “The 2025 AI Index Report,” Research and Development](https://hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development)
- [Stanford HAI, “The 2025 AI Index Report,” Technical Performance](https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance)
- [Edwards and Camacho-Collados, “Language Models for Text Classification: Is In-Context Learning Enough?” LREC-COLING 2024](https://arxiv.org/abs/2403.17661)
- [Bucher and Martini, “Fine-Tuned ‘Small’ LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification,” preprint, June 2024](https://arxiv.org/abs/2406.08660)
- [Chen, Zaharia and Zou, “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance,” May 2023](https://arxiv.org/abs/2305.05176)
- [Andy Markus, “The Tokenomics Equation: Balancing Cost and Performance,” AT&T, July 2026](https://about.att.com/blogs/2026/the-tokenomics-equation.html)

---
Grayhaven Industries · https://www.grayhavenindustries.com/blog/smallest-model-that-works
