Google has unveiled Gemini 2.5 Deep Think, a new iteration of its AI model designed for advanced reasoning and complex problem-solving, which gained recognition last month by securing a gold medal at the International Mathematical Olympiad (IMO) — marking the first occasion an AI model achieved this accomplishment.
Yet, this is regrettably not the same gold medal-winning model. It is, in fact, a less potent “bronze” version, as noted in Google’s blog post and by Logan Kilpatrick, Product Lead for Google AI Studio.
As Kilpatrick mentioned on the social network X: “This version of our IMO gold model is quicker and better optimized for everyday use. We are also distributing the complete IMO gold model to a group of mathematicians to assess the full extent of its capabilities.”
Now accessible via the Gemini mobile app, this bronze model is available to subscribers of Google’s premium AI plan, AI Ultra, priced at $249.99 per month, with an introductory offer of $124.99/month for the first three months for new subscribers.
AI Scaling Reaches Its Limits
Power constraints, increasing token costs, and inference delays are transforming enterprise AI. Join our exclusive salon to learn how leading teams are:
- Leveraging energy as a strategic asset
- Developing efficient inference for genuine throughput improvements
- Maximizing competitive ROI with sustainable AI systems
Reserve your spot to stay ahead: https://bit.ly/4mwGngO
Google further announced in its release blog post that it plans to introduce Deep Think with and without tool usage integrations to “trusted testers” through the Gemini application programming interface (API) “in the coming weeks.”
Why ‘Deep Think’ is so powerful
Gemini 2.5 Deep Think expands on the Gemini family of large language models (LLMs), introducing new features aimed at tackling sophisticated problems.
It utilizes “parallel thinking” techniques to explore multiple ideas at once and incorporates reinforcement learning to enhance its problem-solving skills over time.
The model is tailored for applications that benefit from extended deliberation, such as mathematical conjecture testing, scientific research, algorithm design, and creative tasks like code and design refinement.
Initial testers, including mathematicians such as Michel van Garrel, have employed it to investigate unsolved problems and generate possible proofs.
AI power user and expert Ethan Mollick, a professor at the Wharton School of Business at the University of Pennsylvania, also noted on X that it could take a prompt he often uses to assess new models’ capabilities — “create something I can paste into p5js that will surprise me with its cleverness in crafting something reminiscent of a starship control panel from the distant future” — and transformed it into a 3D graphic, marking the first time any model accomplished this.
Performance benchmarks and use cases
Google emphasizes several critical application areas for Deep Think:
- Mathematics and science: The model can simulate reasoning for intricate proofs, explore conjectures, and interpret dense scientific texts
- Coding and algorithm design: It excels in tasks involving performance tradeoffs, time complexity, and multi-step logic
- Creative development: In design scenarios such as voxel art or user interface creation, Deep Think exhibits enhanced iterative improvement and detail refinement
The model also leads in benchmark evaluations like LiveCodeBench V6 (for coding proficiency) and Humanity’s Last Exam (encompassing math, science, and reasoning).
It outperformed Gemini 2.5 Pro and rival models like OpenAI’s GPT-4 and xAI’s Grok 4 by double-digit margins in several categories (Reasoning & Knowledge, Code generation, and IMO 2025 Mathematics).

Gemini 2.5 Deep Think vs. Gemini 2.5 Pro
While both Deep Think and Gemini 2.5 Pro belong to the Gemini 2.5 model family, Google positions Deep Think as a more capable and analytically skilled variant, especially in complex reasoning and multi-step problem-solving.
This advancement arises from the use of parallel thinking and reinforcement learning techniques, enabling the model to simulate deeper cognitive deliberation.
In its official statements, Google describes Deep Think as superior in handling nuanced prompts, exploring multiple hypotheses, and producing more refined outputs. This is evidenced by side-by-side comparisons in voxel art generation, where Deep Think provides more texture, structural fidelity, and compositional diversity than 2.5 Pro.
The enhancements are not merely visual or anecdotal. Google reports that Deep Think outperforms Gemini 2.5 Pro on multiple technical benchmarks related to reasoning, code generation, and cross-domain expertise. However, these improvements come with tradeoffs in responsiveness and prompt acceptance.
Here’s a breakdown:
| Capability / Attribute | Gemini 2.5 Pro | Gemini 2.5 Deep Think |
|---|---|---|
| Inference speed | Faster, low latency | Slower, extended “thinking time” |
| Reasoning complexity | Moderate | High — uses parallel thinking |
| Prompt depth and creativity | Good | More detailed and nuanced |
| Benchmark performance | Strong | State-of-the-art |
| Content safety & tone objectivity | Improved over older models | Further improved |
| Refusal rate (benign prompts) | Lower | Higher |
| Output length | Standard | Supports longer responses |
| Voxel art / design fidelity | Basic scene structure | Enhanced detail and richness |
Google acknowledges that Deep Think’s higher refusal rate is an area of active investigation. This could limit its flexibility in handling ambiguous or informal queries compared to 2.5 Pro. In contrast, 2.5 Pro is better suited for users who prioritize speed and responsiveness, particularly for lighter, general-purpose tasks.
This differentiation allows users to choose based on their priorities: 2.5 Pro for speed and fluidity, or Deep Think for rigor and reflection.
Not the gold medal winning model, just a bronze
In July, Google DeepMind made headlines when a more advanced version of the Gemini Deep Think model achieved official gold-medal status at the 2025 IMO — the world’s most prestigious mathematics competition for high school students.
The system solved five of six challenging problems and became the first AI to receive gold-level scoring from the IMO.
Demis Hassabis, CEO of Google DeepMind, announced the achievement on X, stating the model had solved problems end-to-end in natural language — without needing translation into formal programming syntax.
The IMO board confirmed the model scored 35 out of a possible 42 points, well above the gold threshold. Gemini 2.5 Deep Think’s solutions were described by competition president Gregor Dolinar as clear, precise, and in many cases, easier to follow than those of human competitors.
However, the Gemini 2.5 Deep Think released to users is not that same competition model, rather, a lower performing but apparently faster version.
How to access Deep Think now
Gemini 2.5 Deep Think is available exclusively on the Google Gemini mobile app for iOS and Android at this time to users on the Google AI Ultra plan, part of the Google One subscription lineup, with pricing as follows.
- Promotional offer: $124.99/month for 3 months, then it increases to…
- Standard rate: $249.99/month
- Included features: 30 TB of storage, access to the Gemini app with Deep Think and Veo 3, as well as tools like Flow, Whisk, and 12,500 monthly AI credits
Subscribers can activate Deep Think in the Gemini app by selecting the 2.5 Pro model and toggling the “Deep Think” option.
It supports a fixed number of prompts per day and is integrated with capabilities like code execution and Google Search. The model also generates longer and more detailed outputs compared to standard versions.
The lower-tier Google AI Pro plan, priced at $19.99/month (with a free trial), does not include access to Deep Think, nor does the free Gemini AI service.
Why it matters for enterprise technical decision-makers
Gemini 2.5 Deep Think represents the practical application of a major research milestone.
It allows enterprises and organizations to tap into a Math Olympiad medal-winning model and have it join their staff, albeit only through an individual user account now.
For researchers receiving the full IMO-grade model, it offers a glimpse into the future of collaborative AI in mathematics. For Ultra subscribers, Deep Think provides a powerful step toward more capable and context-aware AI assistance, now running in the palm of their hand.
