ChatGPT 6.1 Sol vs Gemini 4.0 Argon: Benchmarks, Pricing, and Which to Choose
Gemini 4 Argon edges ahead on independent intelligence scores, but GPT-6.1 Sol is the one you can actually use today, and it costs far less per task. Here is the full side-by-side.
Short answer: in the ChatGPT 6.1 Sol vs Gemini 4.0 Argon matchup, Argon scores slightly higher on independent intelligence tests, while Sol is the model you can actually use today and costs far less per finished task. Which one wins for you depends on access, workload, and how much you care about token spend.
Both models landed within a day of each other at the end of September 2026, and both list at the same $2 input and $10 output price per million tokens. That identical price card is what makes the comparison interesting, because the real bills end up very different. This guide breaks down specs, benchmarks, real cost, access, and the best pick for each job.
A note on names: OpenAI's model is officially called GPT-6.1 Sol and is available in ChatGPT Work, Codex, and the API. Google's is Gemini 4 Argon. This article uses "ChatGPT 6.1 Sol" and "Gemini 4.0 Argon" as the search-friendly versions of those names. We did not run our own tests; every figure below comes from vendor announcements and third-party trackers, and we flag where sources disagree.
Quick Verdict: ChatGPT 6.1 Sol vs Gemini 4.0 Argon
- Best raw score: Gemini 4 Argon, by about 1 to 2 points on the Artificial Analysis Intelligence Index (53 vs 51).
- Best availability: GPT-6.1 Sol. Argon is limited to Google's Fairwind Program at the time of writing.
- Best cost per task: GPT-6.1 Sol, which uses far fewer tokens to finish the same work.
- Best for very long outputs: Gemini 4 Argon, the only one of the two with a 1M-token output cap per response.
- Best for most teams today: GPT-6.1 Sol, simply because it is available, documented, and cheap to run.
Specs and Pricing Side by Side
Here is how the two models compare on the basics, using launch-day documentation and early coverage.
| Spec | GPT-6.1 Sol | Gemini 4 Argon |
|---|---|---|
| Developer | OpenAI | Google DeepMind |
| Release | Sep 29, 2026 | Announced Sep 30, 2026 |
| Access | OpenAI API, ChatGPT Work, Codex | Fairwind Program only |
| Input / output per 1M tokens | $2 / $10 | $2 / $10 intro, then $4 / $20 |
| Cached input per 1M tokens | $0.10 | $0.10 (intro) |
| Context window | 1.05M (922K max input) | Not officially disclosed |
| Max output per response | 128K tokens | 1M tokens |
| Long-prompt pricing | Above 272K input: 2x input, 1.5x output | Not disclosed |
| Reasoning control | 5 effort levels, low to max (medium default) | Not disclosed |
| Open weights | No | No |
Two details matter here. Sol's pricing is standard and public. Argon's $2 / $10 rate is introductory and doubles later, with no end date announced, so treat it as a promotional price rather than a permanent one.
Release Timeline and Who Can Use Each Model
OpenAI shipped GPT-6.1 Sol on September 29, and it became available through the API, ChatGPT Work, and Codex. Google announced Gemini 4 Argon a day later, but opened it only through the Fairwind Program, which is aimed at cyber defenders. For most developers and businesses, that means Sol is something you can test this afternoon and Argon is something you read about.
Context helps here. OpenAI also cancelled a planned GPT-6.1 Astra launch on September 28 after it failed internal scope and authorization tests, so GPT-6 Astra remains OpenAI's top model while Sol fills the cheaper, faster tier. Google has said it plans for Argon to eventually run most of its products, so broader access is expected, but no public date has been confirmed.
Benchmarks: Where Each Model Leads
Artificial Analysis tested both models under one methodology, which makes its numbers the cleanest like-for-like comparison available so far. Argon was run at its High setting and Sol at its xhigh setting.
| Benchmark | GPT-6.1 Sol (xhigh) | Gemini 4 Argon (High) | Leader |
|---|---|---|---|
| Intelligence Index | 51 | 53 | Argon |
| GDPval-AA v2.1 (Elo) | 1510 | 1611 | Argon |
| AutomationBench-AA | 67% | 78% | Argon |
| Terminal-Bench 4.0 | 54% | 57% | Argon |
| SciCode | 56% | 62% | Argon |
| Humanity's Last Exam | 53% | 57% | Argon |
| AA-Omniscience | 41 | 42 | Argon (narrow) |
| AA-LCR v1.1 (long context) | 80% | 80% | Tie |
| AA-Briefcase v1.1 (Elo) | 1507 | 1494 | Sol |
| GDP.pdf | 32% | 22% | Sol |
| CritPt | 32% | 27% | Sol |
The pattern: Argon leads most of the knowledge-work, coding-agent, and science rows, usually by a few points. Sol pulls ahead on document-heavy tests such as GDP.pdf and on CritPt. The overall gap of two Intelligence Index points is small enough that task fit matters more than the headline number.
Vendor-reported results to treat with care
Google's launch table lists Argon at 77.9% on DeepSWE v1.1, 68.9% on the Vals Index, 69.2% on OSWorld-2.0, and 68% on CWE-bench v1 for vulnerability remediation. A separate tracker, BenchLM, shows Argon ahead of Sol on DeepSWE (77.9% vs 71.9%) and on its AutomationBench version (51.3% vs 36.1%). Sol is not in Google's own table, and OpenAI's launch numbers claim Sol roughly matches GPT-6 Astra on DeepSWE at about one-fifth of the price. Those claims and the tracker figures do not line up perfectly, which is normal when different benchmark versions and effort settings are used. Wait for independent replication before betting a production system on any single score.
Hallucination rate
One early write-up reports Sol's hallucination rate at 54% at max effort versus 15% for Argon. If that holds up, it is the most important quality gap between the two for any workflow where a confident wrong answer is costly, such as legal, finance, or research summaries. It is a single secondary source, so verify it on your own prompts before relying on it.
Real Cost: Same Price Card, Very Different Bill
List price is only half the story. Your bill is price multiplied by the number of tokens the model spends to finish a task, and Argon spends a lot more of them.
| Metric | GPT-6.1 Sol | Gemini 4 Argon |
|---|---|---|
| List price (input / output per 1M) | $2 / $10 | $2 / $10 intro, $4 / $20 later |
| Cost per task (Artificial Analysis) | $0.39 (xhigh); about $0.72 at max effort per other reports | $1.99 (High) |
| Output tokens per task | 18k | 62k |
| Reasoning tokens per task | 9k | 36k |
| Cost to run the full Intelligence Index | $662 | $2,407 |
| Cached 200K context, 100 agent steps | $2.00 | $2.00 (intro rate) |
In practice, Argon costs roughly 2.7 times as much per task as Sol according to one summary of the Artificial Analysis data, even at identical list rates. And once Argon's introductory pricing ends, the same task would roughly double again. The one place the two are equal is cached input for long-running agents, where both charge $0.10 per million tokens, so cache-heavy loops cost about the same on either.
The takeaway: if you are running thousands of tasks, Sol's efficiency matters far more than the one-point benchmark gap. If you run a few high-stakes tasks where quality is everything, Argon's extra spend could be justified once you can get access.
Strengths and Weaknesses
GPT-6.1 Sol
- Strengths: public availability, 1.05M context window, low cost per task, five effort levels, strong coding and computer-use performance for its price tier.
- Weaknesses: trails Argon on most independent knowledge-work scores, reported higher hallucination rate, 128K output cap, and slower responses at the highest effort (one tracker shows about 57 tokens per second and a long wait to the first answer token at xhigh).
Gemini 4 Argon
- Strengths: top score on the Intelligence Index among the two, strong knowledge-work and legal/finance results, a 1M-token output cap, and a reported lower hallucination rate.
- Weaknesses: restricted access, introductory pricing that doubles, undisclosed context window and API contract, and much higher token use per task.
Which Model Should You Choose?
| Job | Pick | Why |
|---|---|---|
| High-volume coding agents, CI bots, PR review | GPT-6.1 Sol | Low cost per task, public API, strong coding results for the price |
| Legal, finance, business automation | Gemini 4 Argon | Leads the Vals Index and AutomationBench in published results |
| Very long single outputs (big refactors, full reports) | Gemini 4 Argon | 1M output tokens per response vs 128K |
| Tightest budget at frontier quality | GPT-6.1 Sol | Lowest effective cost and no access restrictions |
| Cache-heavy agent loops | Either | Both list cached input at $0.10 per 1M tokens |
| Accuracy-sensitive work | Gemini 4 Argon (once available) | Reported lower hallucination rate; verify on your data |
If you are a developer or small team: start with Sol. You can benchmark it on your own tasks immediately, and the cost is predictable. If you are an enterprise on Google Cloud: keep Argon on your shortlist for knowledge work and long-form generation, and apply for Fairwind access if you qualify. If you are a freelancer or creator: Sol inside ChatGPT Work or Codex is the practical choice today.
What to Check Before You Switch
- Most numbers are vendor-reported or early. Benchmark versions, effort settings, and safeguards differ between sources, so small gaps are not proof of a winner.
- Argon's price is temporary. Budget for $4 / $20 per million tokens, not the $2 / $10 introductory rate.
- Argon's input limit is unconfirmed. Third parties report anywhere from 1M to 2M tokens, but Google has not published an official figure.
- Measure on your own traces. Cost per task changes dramatically with effort level, prompt length, and caching. Run both on a sample of real work before committing.
- Re-check in a few weeks. Argon's access and pricing are the likeliest things to change first.
Final Verdict
There is no outright winner in ChatGPT 6.1 Sol vs Gemini 4.0 Argon, only a trade-off. Argon posts the higher scores and reportedly hallucinates less, but it is gated, burns many more tokens per task, and its price is set to double. Sol gives up a couple of points on independent tests, yet it is available now, documented, and several times cheaper to run. For the majority of users, Sol is the sensible default today; Argon becomes the stronger contender the moment Google opens broad access at a stable price.
Frequently Asked Questions
Q: Is ChatGPT 6.1 Sol better than Gemini 4.0 Argon?
Q: Can I use Gemini 4 Argon today?
Q: How much do GPT-6.1 Sol and Gemini 4 Argon cost?
Q: Which model is cheaper per task?
Q: What is the context window of each model?
Q: Which is better for coding?
Q: Does Gemini 4 Argon hallucinate less than GPT-6.1 Sol?
Sources & Academic References
- Artificial Analysis: Gemini 4 Argon (High) vs GPT-6.1 Sol (xhigh)
- MarkTechPost: GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1
- Google DeepMind: Gemini 4 Argon announcement
- BenchLM: Gemini 4 Argon vs GPT-6.1 Sol
- eesel AI: Gemini 4 Argon alternatives
- Kingy AI: Frontier AI Models Compared
Discussion 0
No comments yet. Be the first to share your thoughts!
Leave a Comment