Thinking mode or no thinking mode: what actually changes in an AI model
It is tempting to assume that a model with a thinking mode is simply smarter than one without it. The reality is more mundane and more interesting: the difference is not in how smart the model is but in how much computation it is willing to spend before it commits to an answer.
Both models are still next-token predictors. Neither one understands, plans or reflects in the human sense. What a thinking mode changes is the amount and the structure of the intermediate text the model is allowed to produce before the token you actually read. That single design choice cascades into everything else: accuracy on hard problems, latency, token cost and when the whole thing is worth using at all.
Thinking modes are not exclusive to cloud models. Open models run on Ollama expose the same controls and we cover how to turn them on and off in our guide to disabling thinking mode.
1. What a model without a thinking mode actually does
Strip away the language and a base large language model does one thing: given a sequence of tokens it predicts the most plausible next token, appends it and repeats. There is no hidden deliberation step and no internal scratchpad. The only place where reasoning can happen is the generated text itself.
That is not as limiting as it sounds. A model trained purely on next-token prediction has already absorbed an enormous amount of implicit reasoning during training, so it can answer many questions correctly in a single forward pass. Ask it to translate a sentence, summarise a document or continue a familiar pattern and no deliberation is needed.
Where it breaks down is when the correct answer depends on a chain of steps that must all be right. A multi-digit multiplication, a logic puzzle or a tricky bug requires holding intermediate results and a model with no room to write those results down has to guess the answer directly. It often guesses wrong, confidently and fluently at the same time.
2. What a thinking mode adds
A thinking model has been trained to first produce a long stretch of intermediate text — the chain of thought — and only then the final answer. Those intermediate tokens act as a scratchpad: the model can state a subgoal, compute a partial result, notice that it contradicts an earlier line and correct course before the user ever sees the answer.
This is the key insight behind the whole approach, first formalised in Chain-of-Thought Prompting Elicits Reasoning in Large Language Models by Wei et al. Prompting a plain model with "let's think step by step" already improves reasoning, but the gain is fragile. Reasoning models bake that behaviour into training so the model reaches for a chain of thought on its own, reliably and at much greater length.
The mechanism is inference-time compute. A conventional model spends roughly one forward pass per output token. A thinking model spends those same passes on tokens you may never read, then spends a few more on the answer. For a hard problem, throwing sequential computation at the question is often more effective than making the model bigger, as shown by Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters by Snell et al.
3. It is training first, a switch second
The popular framing of a "thinking mode button" is slightly misleading. In most cases you are not toggling a module inside one fixed brain. A reasoning model is usually the same base architecture trained differently, through reinforcement learning on problems with verifiable answers or through distillation from a stronger reasoner. DeepSeek-R1 made this concrete: the reasoning ability emerged largely from reinforcement learning rather than from supervised examples of good reasoning.
That is why disabling thinking does not turn a reasoning model into a genuinely different model. It turns off the long deliberation and asks it to answer directly, which usually makes it faster but weaker on anything that needed the scratchpad. Some families even leave the reasoning trace hidden by default and expose only the final answer, while still generating those tokens internally.
There is also a middle ground. s1: Simple Test-Time Scaling showed that controlling how long a model is allowed to think — a reasoning budget — can trade accuracy against compute. This is exactly what the effort levels in modern models do.
4. The real cost of thinking
Thinking is not free and the bill arrives in three places.
- Tokens. Every reasoning token is generated, stored and often billed. A question that a direct answer handles in 200 tokens can consume thousands once the model deliberates.
- Latency. More tokens mean more time to first useful answer. On a local GPU with a single user there is no batching to hide it, so a thinking model feels noticeably slower.
- Context. The reasoning trace occupies the context window. If it is not trimmed, it pushes out the conversation you actually care about and inflates the KV cache in memory.
There is also a failure mode in the other direction. Reasoning models can overthink a trivial question — "what is the capital of France?" — burning a page of deliberation to produce a one-word answer. A model without thinking mode simply answers and moves on.
5. Which one should you run
The two are not competitors so much as tools with different costs. The table below maps the trade-offs honestly.
| Property | No thinking mode | Thinking mode |
|---|---|---|
| Mechanism | Single forward pass per output token | Long chain of intermediate tokens before the answer |
| Best at | Chat, translation, summarising, drafting, simple lookups | Maths, logic, multi-step code, debugging, planning |
| Latency | Low and predictable | Higher, scales with the reasoning budget |
| Token cost | Low | Often several times higher |
| Error profile | Plausible but wrong answers on multi-step tasks | Fewer logic errors, but can still hallucinate and can overthink |
| Context pressure | Minimal | The trace consumes window and memory |
For everyday local use the rule of thumb is simple. Use the direct mode for conversation and anything you already know the shape of. Reach for thinking when the task has a verifiable answer and a wrong step would ruin it and keep the effort level as low as the task allows. The reasoning budget is a dial, not a binary and the skill is matching it to the problem.
6. Concrete models you can pull today
The theory is easy to check: Ollama ships both kinds of model and several reasoning families let you toggle the thinking mode on and off. Here are representative examples.
Models with a thinking mode
| Model | Ollama pull | Notes |
|---|---|---|
| DeepSeek-R1 | ollama pull deepseek-r1 |
The reference open reasoner. Long chain of thought by default; distills exist at 7B, 14B and 32B for smaller machines. |
| DeepSeek V4 Flash | ollama pull deepseek-v4-flash |
General model with a reasoning mode that can be disabled for straightforward prompts. |
| Qwen 3.6 | ollama pull qwen3.6 |
Exposes thinking and non-thinking behaviour in the same weights; effort varies with the task. |
| Gemma 4 31B | ollama pull gemma4:31b |
Google's flagship supports a thinking mode; the smaller 12B build is the lighter compromise. |
| Muse Glimmer 30B | ollama pull muse-glimmer:30b |
Meta's agent model with controllable effort (low, medium, high, xhigh). |
| GLM-5.3 | ollama pull glm-5.3 |
Reasoning-capable general model, also available as a cloud tag. |
| GPT-OSS | ollama pull gpt-oss |
OpenAI's open weights with reasoning traces you can show or hide. |
Models designed without a thinking mode
| Model | Ollama pull | Notes |
|---|---|---|
| Llama 3.x | ollama pull llama3.2 |
Classic instruct models: direct answers, no deliberation step. |
| Mistral | ollama pull mistral |
Fast, lightweight text generation with no reasoning trace. |
| Qwen 2.5 | ollama pull qwen2.5 |
The generation before the 3.x reasoning line; answers directly. |
| Gemma 3 | ollama pull gemma3 |
Earlier Gemma generation, chat and vision oriented rather than reasoning oriented. |
| Phi-4 | ollama pull phi4 |
Compact model tuned for fast, direct responses. |
The boundary is not always clean: a reasoning model run with thinking disabled behaves like the second table for that query and some general models will produce reasoning only when explicitly prompted. Tags and names also change as families evolve, so check ollama show <model> after pulling to see which parameters a build actually exposes.
Recent versions expose this directly: ollama run --think true|false or high|medium|low depending on the model, --hidethinking to hide the trace and /set nothink inside a session. See the dedicated guide for the exact commands.
Conclusion
The gap between a model with a thinking mode and one without it is not a gap in consciousness or understanding. Both are pattern predictors. The thinking model is simply allowed and trained to write down its working before answering and that extra computation buys reliability on problems where a single guess is not enough. Knowing when to pay for it is the real skill.
Sources
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Wei et al. (arXiv:2201.11903)
- Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters — Snell et al. (arXiv:2408.03314)
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (arXiv:2501.12948)
- s1: Simple Test-Time Scaling — Muennighoff et al. (arXiv:2501.19393)