New models ship constantly, so we score them on what actually matters: capability, price, context, and reliability. Here is how the current field stacks up.
How we score
Our composite score blends hands-on testing with public benchmarks, weighted toward practical usefulness rather than leaderboard chasing.
Use the category winners as a shortcut: best for coding, best value, longest context, and best open weights.