#模型时代# 在推上看到有牛人把主要模型基准测试做了统计,称为The Ultimate LLM Benchmark list ,终极模型基准测试列表。
对这个议题感兴趣的可以收藏一波了。作者还跑了一个综合排名:Gemini 2.5 Pro > o3 > Sonnet 3.7 Thinking(作者:Lisan al Gaib,X账号:scaling01)
一、一般性基准测试:
SimpleBench: simple-bench.com/index.html
SOLO-Bench: github.com/jd-3d/SOLOBench
AidanBench: aidanbench.com
SEAL by Scale: scale.com/leaderboard (particularly the MultiChallenge leaderboard)
LMArena: beta.lmarena.ai/leaderboard (with Style Control)
LiveBench: livebench.ai
ARC-AGI: arcprize.org/leaderboard
Thematic Generalization by LechMazur: github.com/lechmazur/generalization
( other ones by Lech Mazur: github.com/lechmazur/elimination_game,
github.com/lechmazur/confabulations, ...)
EQBench: eqbench.com (especially the Longform writing leaderboard)
Fiction-Live Bench: fiction.live/stories/Fiction-liveBench-Mar-25-2025/oQdzQvKHw8JyXbN87
MC-Bench: mcbench.ai/leaderboard (ordered by winrate, not by Elo)
TrackingAI - IQ Bench: trackingai.org/home
Dubesor LLM: dubesor.de/benchtable.html
Balrog-AI: balrogai.com
Misguided Attention: github.com/cpldcpu/MisguidedAttention
Snake-Bench: snakebench.com
SmolAgents LLM: huggingface.co/spaces/smolagents/smolagents-leaderboard (just because of GAIA and SimpleQA)
Context-Arena (MRCR and Graphwalks): contextarena.ai
OpenCompass: rank.opencompass.org.cn/home
HHEM (Hallucination Benchmark): huggingface.co/spaces/vectara/leaderboard
二、编码、数学、Agents基准测试 Coding, Math and Agentic Benchmarks
Aider-Polyglot-Coding: aider.chat/docs/leaderboards/
BigCodeBench: bigcode-bench.github.io
WebDev-Arena: web.lmarena.ai/leaderboard
WeirdML: htihle.github.io/weirdml.html
Symflower Coding: symflower.com/en/company/blog/2025/dev-quality-eval-v1.0-anthropic-s-claude-3.7-sonnet-is-the-king-with-help-and-deepseek-r1-disappoints/
PHYBench: phybench-official.github.io/phybench-demo/
MathArena: matharena.ai
Galileo Agent: huggingface.co/spaces/galileo-ai/agent-leaderboard
XLANG Agent: arena.xlang.ai/leaderboard
三、AI起飞跟踪基准测试 Important for tracking AI take-off
METR long task benchmarks: metr.org (incl. RE Bench)
PaperBench: openai.com/index/paperbench/
SWE-Lancer: openai.com/index/swe-lancer/
MLE-Bench: github.com/openai/mle-bench
SWE-Bench: swebench.com
四、新模型发布会看的基准测试 other classics I ALWAYS want to see when a new model is released
GPQA-Diamond: github.com/idavidrein/gpqa
SimpleQA: openai.com/index/introducing-simpleqa/
Tau-bench: github.com/sierra-research/tau-bench
SciCode: github.com/scicode-bench/SciCode
MMMU: mmmu-benchmark.github.io/#leaderboard
Humanities Last Exam (HLE): github.com/centerforaisafety/hle
五、经典基准测试 Overview for classical benchmarks (GPQA, SimpleQA, AIME, MMLU, ...)
Simple-Evals: github.com/openai/simple-evals
Vellum AI: vellum.ai/llm-leaderboard
Artificial Analysis: artificialanalysis.ai
六、不太关心的基准测试Benchmarks I literally don't care about - saturated / no signal
MMLU, HumanEval, BBH, DROP, MGSM, basically all math benchmarks like GSM8K, MATH, AIME
