Grok 4.7 Claims Top Three Spots on VulcanBench Frontier v4 Leaderboard
xAI’s newest coding model has swept the top three ranks on VulcanBench Frontier v4 while maintaining its previous pricing structure.
AI-generated image
Grok 4.7 Claims Top Three Spots on VulcanBench Frontier v4 Leaderboard
xAI's newest coding model has swept the top three ranks on VulcanBench Frontier v4 while maintaining its previous pricing structure.

xAI Dominates VulcanBench Leaderboard
xAI's latest coding model, Grok 4.7, has secured the first, second, and third positions on the VulcanBench Frontier v4 leaderboard. Launched on September 21, 2026, the model allows users to select how much reasoning effort it applies to a task, with each setting evaluated separately.
At the extra-high effort level, Grok 4.7 achieved a score of 93.15 to claim first place. The high-effort setting followed closely with 92.71 for second place, while the medium setting secured third place with a score of 92.30.
Performance and Benchmark Details
The standout performance occurred at the extra-high effort setting, where Grok 4.7 successfully passed all 23 behavioral-reconstruction tasks included in the benchmark. VulcanBench Frontier v4 evaluates AI models on real-world software engineering problems, focusing on functional correctness, code quality, and complexity through deterministic hidden tests.
Grok 4.7's medium setting outperformed the closest named rival, Claude Fable 5.1, which recorded a score of 91.84. Built on an extended base model relative to its predecessor Grok 4.6, the new model supports text and image inputs, integrates additional tools, and features a context window of 500,000 tokens.
Pricing and Developer Utility
Pricing for Grok 4.7 remains unchanged from Grok 4.6, costing $2 per million input tokens and $6 per million output tokens. In addition to VulcanBench, xAI reported performance improvements over Grok 4.6 across other coding benchmarks, including CursorBench and Terminal-Bench.
The tiered reasoning effort levels provide developers with flexibility, enabling teams to utilize lower settings for routine tasks while reserving extra-high mode for more complex software engineering challenges.