GPT-5 is OpenAI's flagship language model, generally available since August 2025. It runs as the default ChatGPT model across all plans and ships in four API variants: gpt-5, gpt-5-mini, gpt-5-pro, and a 'thinking' mode picked automatically by a router. Per OpenAI's GPT-5 system card, the model hallucinates roughly 80% less than GPT-4o and scores 74.9% on SWE-bench Verified, the highest result of any major foundation model in 2026.

## May 2026 Update

GPT-5 has been generally available since August 2025 and is now nine months into production. ChatGPT users got a smarter default model. Developers got three API tiers (mini, standard, pro) plus a router that picks 'thinking mode' automatically when a question is hard.

Runbear now runs on GPT-5 via the OpenAI API, powering context assembly and automated actions across 2,000+ connected tools in Slack. The improvement in instruction following (GPT-5 scored 99% on COLLIE vs earlier baselines) shows up directly in how precisely Runbear agents route requests, draft responses, and take action.

[See the Aloware case study](/content/posts/aloware-zoom-transcript-agent-case-study/index.html) for a real example: a Zoom transcript agent that logs CRM deals with a single emoji, built on this model layer.

The rest of this post covers everything you need to know: what GPT-5 is, how it compares to prior models, availability, pricing, and what it means for ops teams using AI in Slack.

## What Is GPT-5?

GPT-5 is OpenAI's most advanced language model yet—built to be more robust, reliable, and helpful than any of its predecessors. It delivers state-of-the-art accuracy on real-world tasks while significantly reducing hallucinations and deceptive responses. Designed with real workflows in mind, it powers real-time AI agents, developer tools, and enterprise copilots across platforms like Slack, HubSpot, and Microsoft Teams.

## Key Improvements in GPT-5

- State-of-the-art performance on leading AI benchmarks:
- **Fewer hallucinations**: GPT-5 responses are up to **80%** more factual than previous models like GPT-4o or o3.
- **Safer and more honest**: Reduced deceptive behavior in edge cases or underspecified prompts.
- **Improved instruction following**: More reliable for agents, workflows, and multi-step tasks.
- **Multimodal reasoning**: Enhanced capabilities across **text, image, video, and charts**.

## What makes GPT-5 different?

Rather than a single monolithic model, GPT-5 introduces a router-based model architecture:

- **GPT-5**: Fast, general-purpose response generation
- **GPT-5 Thinking**: Deep reasoning for complex tasks
- **GPT-5 Pro**: Long-context, high-precision model tuned for advanced applications

These models are deployed intelligently behind the scenes, depending on query difficulty, user intent, and context. This allows GPT-5 to feel both fast and deeply capable—delivering expert-level answers where needed, while still handling everyday queries with speed and grace.

## Evaluations: How Smart Is GPT-5?

### Math Mastery

GPT-5 dominates math evaluations:

- **AIME 2025**: 94.6% (no tools) — new state-of-the-art.
- **HMMT**: 96.7% (no tools), 100% (with tools).
- **FrontierMath**: 26.3% → 32.1% with tool support.
- **GPQA (PhD-level)**: 88.4% (no tools) → 89.4% (with thinking).

### Real-World Coding

On practical software engineering and code editing benchmarks:

- **SWE-bench Verified**: 74.9% accuracy (GPT-5) vs 69.1% (OpenAI o3) and 30.8% (GPT-4o).
- **Aider Polyglot**: 88% accuracy (GPT-5), far ahead of all previous models.

### Instruction Following & Tool Use

GPT-5 is vastly better at multi-step reasoning and agentic behaviors:

- **Scale MultiChallenge (multi-turn)**: 69.6% vs GPT-4o's 40.3%.
- **BrowseComp (search + browsing)**: 54.9% vs 49.7%.
- **COLLIE (freeform following)**: 99.0% accuracy.

### Multimodal Understanding

GPT-5 outperforms on visual, video, and diagram-based tasks:

- **MMMU**: 84.2% vs GPT-4o's 72.2%.
- **ERQA**: 65.7% vs 35.2% (GPT-4o).

### Health Conversations

GPT-5 is the most accurate and least hallucinatory model for medical applications:

- **HealthBench**: 67.2% vs GPT-4o's 32.0%.
- **Hallucination Rate**: 1.6% (with thinking), down from GPT-4o's 15.8%.

### Economically Valuable Tasks

GPT-5 beats both o3 and ChatGPT Agent on complex, real-world professional tasks:

- Achieves **47.1%** wins over industry experts in internal benchmarks.

### Faster, More Efficient Thinking

GPT-5 achieves better performance with fewer output tokens than OpenAI o3:

- **50–80%** fewer tokens across reasoning, coding, and scientific benchmarks.

### Detailed Benchmarks

## Who is GPT-5 for?

- **Business leaders & teams**: Analyze documents, monitor operations, and plan strategy with better context awareness and safer outputs.
- **Developers**: Use GPT-5 to generate production-quality frontends, debug codebases, and automate software tasks.
- **Healthcare professionals & patients**: Understand diagnoses and treatment options with greater clarity and proactive reasoning.
- **Knowledge workers & writers**: Get help crafting thoughtful reports, articles, or even poetry—with deeper structure and emotional impact.

## Final Thoughts

GPT-5 sets a new standard in reasoning, reliability, and real-world AI performance. With better benchmarks and lower hallucination rates, it's shaping up to be the most capable foundation model to date.

Runbear agents are powered by GPT-5 via the OpenAI API. That means every context assembly, Slack response, and automated action your team runs through Runbear benefits from GPT-5's improved reasoning and lower hallucination rates.
