OpenAI Unveils New Reasoning Models o3 and o4-mini, Raising the Bar in AI Performance
Advancing AI Reasoning: The Rise of o3 and o4-mini
OpenAI describes its new flagship model, o3, as its most capable reasoning model to date, surpassing its predecessors in areas such as mathematics, science, coding, logic, and visual understanding.
“OpenAI calls o3 its most advanced reasoning model ever,” the company stated, noting its superiority in performance metrics compared to previous releases. Meanwhile, the o4-mini model offers a balance that’s particularly attractive to developers, with a solid compromise between performance, speed, and cost-efficiency.
Multimodal Abilities and Chain-of-Thought Reasoning
Both o3 and o4-mini break new ground by combining powerful reasoning with multimodal tools. These models can now analyze images, write and execute Python code directly in ChatGPT via the Canvas feature, and even browse the internet when tasked with current event queries.
“OpenAI claims that o3 and o4-mini are its first models that can ‘think with images,'” allowing users to submit whiteboard diagrams, scanned pages, or blurry visuals. The models can interpret these inputs during their reasoning process — a function referred to as the “chain-of-thought” phase. Additionally, they can manipulate visuals through zooming and rotating, enhancing their utility in technical and design tasks.
Performance Benchmarks: State-of-the-Art Results
The new models aren’t just more flexible — they also outperform rival technologies. On SWE-bench verified (without custom scaffolding), a benchmark that evaluates code generation and problem-solving:
-
o3 scored 69.1%
-
o4-mini scored 68.1%
-
Claude 3.7 Sonnet scored 62.3%
-
OpenAI’s previous model, o3-mini, scored 49.3%
These figures underscore the technical leap forward, reinforcing OpenAI’s positioning in a competitive market crowded with powerful offerings from Anthropic, Meta, Google, xAI, and DeepSeek.
Why o3 Almost Didn’t Launch
Interestingly, o3 was nearly held back from integration into ChatGPT. CEO Sam Altman had previously indicated in February that resources might be redirected toward a next-generation system incorporating o3’s core technologies. However, industry momentum and the competitive threat from rival labs appear to have influenced the company’s decision to release the model sooner than expected.
Pricing and Availability
OpenAI has made these models available across multiple platforms:
-
ChatGPT Pro, Plus, and Team subscribers can access o3, o4-mini, and an enhanced version called o4-mini-high, which takes longer to respond but aims for greater reliability.
-
Developers can integrate the models via OpenAI’s Chat Completions API and Responses API, with usage-based pricing.
Here’s how pricing breaks down:
-
o3: $10 per million input tokens and $40 per million output tokens
-
o4-mini: $1.10 per million input tokens and $4.40 per million output tokens — the same as o3-mini
Notably, one million tokens is equivalent to approximately 750,000 words — more than the full Lord of the Rings trilogy.
What’s Next? GPT-5 on the Horizon
Looking ahead, OpenAI has confirmed plans to release o3-pro, a more powerful variant of o3, in the coming weeks. This version will be exclusive to ChatGPT Pro users and promises even more refined performance, thanks to greater computing resources.
In what might signal the end of the standalone reasoning model era, CEO Sam Altman has hinted that o3 and o4-mini may be the final such models before the launch of GPT-5. The future model is expected to consolidate OpenAI’s traditional GPT architecture with its specialized reasoning models, marking a major step toward unified AI systems.




0 Comments