While headlines chase ever-larger frontier models, a quieter trend has been reshaping how AI actually gets deployed: small language models, or SLMs — models in the single-digit-billion parameter range, distilled or trained to punch far above their size on narrow, well-defined tasks.
Why size isn't the only axis that matters
A frontier LLM is a generalist — it can write poetry, debug code, and hold a conversation, all with the same weights. That generality has a cost: higher latency, higher inference cost, and often more infrastructure to run reliably at scale. An SLM trades away some of that generality for speed and efficiency, and for a huge class of production tasks — classification, extraction, short structured responses, on-device inference — that trade is a clear win.
Where SLMs fit into a real AI product
- Real-time or high-volume tasks where latency directly affects user experience — a live voice conversation can't wait several seconds for a frontier-model round trip.
- Narrow, repetitive judgment calls — tagging, routing, short structured extraction — where a smaller fine-tuned model matches a much larger general one at a fraction of the cost.
- Privacy- or cost-sensitive contexts where running inference on smaller, sometimes even on-device, models is preferable to sending everything to a large hosted model.
The practical pattern we see across the industry — and increasingly apply ourselves — is a mixture: route the reasoning-heavy, ambiguous, high-stakes decisions to a frontier LLM, and route the high-volume, well-scoped, latency-sensitive work to smaller, cheaper models. Treating 'which model' as a routing decision, not a one-size-fits-all choice, is what actually makes an AI product affordable to run at scale.
The best model for a task is rarely the biggest one — it's the smallest one that's still reliably good enough.