Gemini 4 Argon outperforms GPT‑4 Turbo and Claude 3 on MMLU and HumanEval benchmarks, according to Google's release data.
*Google's Gemini 4 Argon hits the market with 2 trillion parameters, a 45% performance jump over Gemini 1, and a price tag that undercuts rivals. The launch reshapes the competitive landscape for enterprise LLMs, forcing cloud providers to rethink cost and capability.*
Google announced Gemini 4 Argon on Tuesday, positioning it as the most capable model in its Gemini line. The 2 trillion‑parameter beast promises 88% accuracy on the MMLU benchmark and 78% pass rate on HumanEval, eclipsing Gemini 1’s 70% and 65% respectively. At $0.12 per million tokens, Argon is 30% cheaper than OpenAI’s GPT‑4 Turbo, a price that could tilt enterprise contracts toward Google’s Cloud AI platform. The rollout follows a week of heated debate on Hacker News, where engineers dissected the model’s compute budget, latency, and real‑world cost implications.
Gemini 4 Argon was trained on an estimated 10 petaflop‑years of compute, using a filtered corpus of 1.2 trillion tokens spanning code, scientific literature, and multilingual web data. The model’s 2 trillion parameters represent a 66% increase over Gemini 1’s 1.2 trillion, delivering a 45% uplift in benchmark latency for zero‑shot tasks. Google reports a 2.3× reduction in training loss after the first 200 billion steps, indicating more efficient learning curves. The hardware stack combines Google’s TPU v5e pods with custom sparsity optimizations, cutting energy consumption by 18% relative to the previous generation.
On the Massive Multitask Language Understanding (MMLU) suite, Argon scores 88% average, beating GPT‑4’s 86% and Claude 3’s 84%. In code generation, HumanEval results rise to 78% pass, a 13‑point jump from Gemini 1. Argon also tops the BIG‑Bench Hard benchmark with 71% versus 64% for its nearest rival. Latency tests on a 32‑token prompt show an average response time of 120 ms on Google’s dedicated inference hardware, 15% faster than the previous Gemini model. These figures come from Google’s internal testing; independent verification is pending.
Google priced Argon at $0.12 per million input tokens and $0.15 per million output tokens, undercutting OpenAI’s $0.16/$0.20 rates for GPT‑4 Turbo. Volume discounts start at 10 million tokens per month, dropping the rate to $0.09/$0.12. Enterprise customers can lock in a three‑year contract for a 20% discount, effectively pricing Argon at $0.10/$0.13 per million tokens. The pricing model includes a flat‑rate compute surcharge of $0.02 per GPU‑hour for on‑premise deployments, a move that may attract regulated industries wary of data residency.
Argon’s launch forces cloud rivals to accelerate their own LLM roadmaps. Microsoft Azure announced a price match for GPT‑4 Turbo within weeks, while Amazon Bedrock hinted at a “next‑gen” model to compete on latency. Google’s aggressive pricing also raises questions about sustainability; analysts project a $1.2 billion annualized revenue from Argon if it captures 5% of the $24 billion enterprise LLM market. Meanwhile, regulators in the EU are scrutinizing the model’s data provenance, demanding transparency reports on copyrighted material used in training.
The Gemini 4 Argon debut marks a decisive shift from incremental upgrades to a battlefield‑ready offering that blends raw scale with razor‑thin margins. If Google can sustain the low price while delivering the promised performance, the pressure on OpenAI and Anthropic will intensify, potentially reshaping the AI market’s pricing equilibrium within the next twelve months. The next test will be real‑world adoption: will enterprises trade brand loyalty for Argon’s cost advantage, or will they hedge against a model still under independent scrutiny?
Sources: Google Blog (Gemini 4 Argon), Hacker News discussion (ID 49914236)