Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Blueprint-Bench 2 Leaderboard & Scores — August 2026
2+ day, 3+ hour ago (148+ words) BenchLM mirrors the published score view for Blueprint-Bench 2. Claude Fable 5 leads the public snapshot at 38.6%, followed by Gemini 3.5 Flash (33.6%). BenchLM does not use these results to rank models overall. Google reported Blueprint-Bench 2 in the Gemini 3.5 Flash launch comparison table. BenchLM…...
GPT-4o mini vs GPT-5.4: Benchmarks & Cost
2+ day, 3+ hour ago (97+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Supported · Public rank #10 GPT-4o mini has the lower estimated token cost for this stated workload. Costs use the listed standard API rates. GPT…...
Gemma 4 31B vs GPT-5.4 mini: Benchmarks & Cost
2+ day, 3+ hour ago (237+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Supported · Public rank #50 Estimated · Public rank #84 Gemma 4 31B has the higher public score estimate, 60.09 versus 55.87, but the 90% score intervals overlap. Treat that as a…...
Claude Opus 4.7 vs Claude Sonnet 5: Benchmarks & Cost | BenchLM.ai
1+ day, 9+ hour ago (194+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Supported · Public rank #15 Estimated · Public rank #35 Claude Opus 4.7 has the higher public score estimate, 71.21 versus 64.5, but the 90% score intervals overlap. Treat that as…...
Claude Opus 5 vs MiniCPM-o 2.6: Benchmarks & Cost
2+ day, 3+ hour ago (83+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Supported · Public rank #3 MiniCPM-o 2.6 has no comparable published API token rate. Generally Available · Claude API LAB all-pass (Anthropic harness) LAB all-pass (Harvey held-out)…...
Claude Opus 5 vs MERaLiON-AudioLLM: Benchmarks & Cost
2+ day, 3+ hour ago (82+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Supported · Public rank #3 MERaLiON-AudioLLM has no comparable published API token rate. Generally Available · Claude API LAB all-pass (Anthropic harness) LAB all-pass (Harvey held-out)…...
hexgrad AI Models: Benchmarks & Pricing (August 2026)
1+ day, 11+ hour ago (41+ words) BenchLM This page groups canonical hexgrad model families and applies the same public scoring lane used by the main leaderboard. The model to choose, the cheaper alternative, and the release we would wait on....
Kimi K3 vs MiniMax M1 80k: Benchmarks & Cost
2+ day, 3+ hour ago (117+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Supported · Public rank #5 The page does not recommend a cost winner because at least one model cannot fit the stated workload in one…...
GPT-5.2 vs GPT-5.4 mini: Benchmarks & Cost
2+ day, 3+ hour ago (286+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Estimated · Public rank #74 Estimated · Public rank #84 GPT-5.2 has the higher public score estimate, 57.77 versus 55.87, but the 90% score intervals overlap. Treat that as a…...
GPT-5.6 Sol vs MiniCPM-o 2.6: Benchmarks & Cost
2+ day, 3+ hour ago (67+ words) Updated August 7, 2026. Public scores include evidence status and uncertainty. They are not guarantees for a specific workload. Supported · Public rank #4 MiniCPM-o 2.6 has no comparable published API token rate. Generally Available · OpenAI Responses API FrontierMath v2 (Tiers 1-3) FrontierMath v2 (Tier 4) The published evidence…...