In Thinking, Fast and Slow, psychologist Daniel Kahneman described two modes of human cognition: System 1, fast and automatic — the snap judgment, the instinctive answer — and System 2, slow and deliberate — the careful, step-by-step reasoning you engage when a problem actually demands it. It turns out to be a remarkably good lens for understanding where language model reasoning is headed.
Early LLMs were almost pure System 1
A standard LLM call generates its answer in a single forward pass — token after token, with no explicit backtracking or self-checking. For simple, familiar questions, that's plenty. For a genuinely hard, multi-step problem, it behaves a lot like a human blurting out the first plausible-sounding answer: fluent, confident, and sometimes wrong.
Chain-of-thought and reasoning models are System 2, built in
The industry's answer has been to make models externalize intermediate reasoning — 'thinking out loud' through a problem before committing to a final answer — rather than jumping straight to a conclusion. Modern reasoning models take this further, spending a variable, sometimes large, 'thinking budget' on harder problems before responding, deliberately trading latency for accuracy exactly when it's warranted.
This isn't free, though. Reasoning tokens consume the same budget as the actual output, and a system tuned for a fast, structured answer can end up silently truncated if a model spends its whole budget 'thinking' before it starts writing the response you actually need. Getting this balance right — enough deliberation to be accurate, without burning the whole budget on invisible reasoning — is one of the less glamorous but more important parts of building reliable AI features.
Why this matters for AI agents, specifically
An agent that scores a resume against a job description, or judges whether a candidate's answer actually satisfies a requirement, is exactly the kind of task that benefits from System-2-style deliberation: it's a judgment call with real consequences, not a reflex. Building good agents increasingly means deciding, task by task, how much 'slow thinking' a decision deserves — and making sure the system around the model gives it room to do that thinking without breaking.