Can AI Really Think? What Apple’s New Study Tells Us

What “Reasoning” Means in Language Models – Simple and Technical Explanations
Artificial intelligence has made impressive strides in recent years. Large language models (LLMs) like ChatGPT, Claude, or Gemini amaze users with human-like answers, logical reasoning, and even step-by-step solutions to complex problems. A term that keeps popping up in this context is “reasoning.”
But what does it really mean? Can these models actually think? And how should we understand reasoning in the world of AI?
What Is “Reasoning”? – A Simple Explanation
In everyday language, reasoning means logical thinking or drawing conclusions. When humans solve a problem, they usually follow a step-by-step thought process, analyzing information and reaching a conclusion based on logic.
Example:
“If it’s raining, the street will be wet. It’s raining. Is the street wet?”
A reasoning-capable model should understand this logical chain and answer: Yes.
But it doesn’t stop at yes/no questions. Reasoning includes solving math problems, interpreting text, or developing plans – in short, simulating a type of thinking process.
Reasoning from a Technical Perspective
In AI research, reasoning refers to a model’s ability to go beyond pattern-matching and engage in structured, multi-step logical processing. The goal is not just to generate statistically likely responses, but to simulate deliberate, rule-based thinking.
This is often achieved through Chain-of-Thought prompting, where the model is guided to break its responses into individual reasoning steps – much like a student writing out their work in a math exam. Ideally, the model builds internal representations of causes, conditions, and conclusions.
However, studies – including a new one from Apple – show that this type of “thinking” often breaks down under pressure. Rather than true reasoning, many LLMs appear to mimic logic based on training patterns, without robust generalization.
Apple Research Questions AI’s Ability to Reason
Why Reasoning-Language Models Collapse on Complex Problems
A recent study by Apple Research challenges one of the biggest assumptions in modern AI: Can language models like GPT, Claude, or Gemini really reason? The study looked closely at how well so-called Reasoning LLMs perform when faced with tasks of varying complexity – and the results are troubling.
The Three Performance Phases
Apple’s researchers categorized model performance into three levels:
- Simple Tasks
Standard models (without “reasoning” enhancements) performed best – fast, accurate, and efficient. - Medium Complexity
Reasoning-enhanced LLMs took the lead here, effectively applying step-by-step logic with structured outputs. - High Complexity
Here comes the shock: All models – including reasoning-focused ones – collapsed completely. They gave up mid-task, even with plenty of tokens left. Accuracy dropped to zero percent.

Why Reasoning Matters
Reasoning is crucial for advanced AI applications – from accurate medical advice to autonomous agents. But if LLMs can’t reliably solve complex problems, then better models are needed – not just bigger ones. The goal is not more data, but smarter architecture.
Reasoning is more than a buzzword – it’s a central challenge on the path to truly intelligent systems.
Overthinking & Underthinking
The researchers noted two strange behaviors:
- Overthinking on simple tasks – models make errors after initially solving them correctly.
- Underthinking on hard tasks – models give up prematurely or avoid answering.
This suggests current models don’t genuinely “think” – they mimic thinking. Their performance relies heavily on known patterns, not true logical generalization.
Apple’s Clear Conclusion
Apple’s team argues that we’ve hit fundamental scaling limits. Adding more data or parameters no longer yields real reasoning improvements. Instead, we need new architectures – designed from the ground up to handle true logical abstraction and sustained problem-solving.
“Today’s Reasoning LLMs only appear to think. When complexity increases, their structural limitations are exposed.”
– Apple Research, 2025
What This Means for AI’s Future
This research is a major reality check. Language models may sound brilliant – but they don’t think like humans. Anyone building with LLMs must understand these limits. At the same time, the findings highlight where the future of AI needs to go: beyond simulation, toward genuine reasoning.







