Model Leaderboard
Compare AI models by capability and cost-effectiveness
Popular Comparisons
Programming & Development
59/273 modelsLiveCodeBench: Real-world coding tasks
Use Cases: Code completion, debugging, code review, script generation
Logical Reasoning
61/273 modelsHLE: Complex reasoning and problem-solving
Use Cases: Complex decision-making, multi-step analysis, logical reasoning
Knowledge Q&A
63/273 modelsMMLU Pro: Broad knowledge assessment
Use Cases: Expert Q&A, fact-checking, educational tutoring
Scientific Research
67/273 modelsGPQA: Graduate-level science questions
Use Cases: Academic research, scientific writing, experiment design
Mathematical Computation
51/273 modelsAIME: Competition-level math problems
Use Cases: Financial analysis, data computation, statistical reasoning
AI Agent
46/273 modelsTau2: Autonomous task completion
Use Cases: Automated workflows, multi-tool invocation, complex task decomposition
SciCode
58/273 modelsSciCode: Scientific coding challenges
Use Cases: Scientific computing, research code, data analysis scripts
Terminal
47/273 modelsTerminal-Bench: Command-line operations
Use Cases: Shell scripting, system administration, DevOps automation
Instruction
46/273 modelsIFEval: Instruction following accuracy
Use Cases: Precise task execution, format compliance, constraint adherence