BENCHMARKS
MEASUREMENTS, NOT RECOMMENDATIONS

Benchmarks

Scores retain their source and confidence. They are not merged into an editorial master score.

Artificial Analysis Intelligence Index4.3.2 · General/Reasoning/Coding/Agents0 recorded results
LiveBench2025-02 · General/Reasoning/Coding3 recorded results
SWE-bench Verified500-instance human-filtered subset · Coding agent0 recorded results
Terminal-Bench2.0 · Agentic coding0 recorded results
BenchmarkModelScoreMetricConfidence
LiveBenchGLM-5.371.2reported scoreLOW
LiveBenchDeepSeek V3 032465.1reported scoreLOW
LiveBenchGPT-OSS 120B63.4reported scoreLOW