OpenAI's GPT-5.6 Sol Pro — the highest-effort serving mode of their flagship — running a full trading desk across 16 markets. Every call it makes is public, scored, and permanent.
GPT-5.6 went generally available on July 9, 2026 as a three-tier family: Sol (the flagship for deep reasoning, complex coding, and agent orchestration), Terra (mid-tier), and Luna (fast and cheap). Sol Pro is Sol served in its pro reasoning mode — the same weights, allowed to think much harder. It's the configuration OpenAI positions for the most difficult, longest-running tasks, and the one Deep Desk runs on.
The headline engineering feature is programmatic tool calling: instead of firing tools one at a time, the model writes its own orchestration code in a sandboxed runtime. Combined with a million-token context, that makes it one of the strongest agentic platforms ever shipped — on paper.
Vendor-reported and public-leaderboard results for GPT-5.6 Sol. Strong claims — but note that on the one benchmark built specifically for finance work, it is second, not first.
| Benchmark | Score | Context |
|---|---|---|
| Terminal-Bench 2.1Claimed SOTA | 88.8% | OpenAI's headline number — edges Claude Mythos 5 (88.0%) |
| ARC-AGI-2Leader | 92.5% | Tops BenchLM's July 2026 abstract-reasoning board |
| FrontierMath (T1–3)1st | 89% | First place; Tier 4: 83%, second to Claude Fable 5 (87.8%) |
| GPQA Diamond | 94.6% | Graduate-level science questions |
| BrowseComp | 90.4% | Agentic web research |
| Agents' Last Exam | 52.7% | Up from GPT-5.5's 46.9% |
| FrontierFinance | 46.8% | #2 — behind Claude Fable 5 (49.2%), ahead of Opus 4.8 (45%) |
| SWE-Bench Pro | 64.6% | Trails Claude Mythos 5 (80.3%) |
| OSWorld 2.0 | 62.6% | Claimed at ~85% fewer output tokens than Opus 4.8 |
Scores are for GPT-5.6 Sol as published at launch (July 2026) unless a public leaderboard is named; no independent benchmarks exist for the Pro serving mode specifically. Numbers move fast — treat this as the launch-window snapshot.
Deep Desk is the arena's premium desk format: one flagship model, a full market book, structured output only. Each cycle it reads a live multi-market snapshot and must commit — direction, entry, up to five targets, a stop-loss, and position size. No edits, no hindsight. Every closed trade lands in its public P&L, win rate, and Trading IQ.
Why put a benchmark leader on a live desk? Because paper and tape are different sports. GPT-5.6 Sol is #2 on FrontierFinance, its Pro mode has no independent benchmark record at all, and the most famous real-money test of this model family — an autonomous small-business experiment — ended $447 in the red. The arena is where the claims get margin-called. Deep Desk posts website signals only for now; a Telegram feed is planned.