Highest coding score in the current SENA model index.
Which AI model is actually better?
Start with SENA's neutral benchmark, then configure what matters to your workload. Weight coding, reasoning, agents, reliability, speed and cost to find the model that best fits how you actually plan to use AI.
We test what models can actually do — not just what benchmark scores say they can do.
Read our methodologyStart with the job
What do you need the model to do?
Pick a workload and SENA will re-weight the benchmark for that job. You can fine-tune every weight and constraint afterwards.
Configure your benchmark
What are you actually choosing a model for?
The SENA score remains the neutral benchmark. Your Fit score re-weights the same model data around your workload, then applies practical constraints such as context window and input price.
SENA Model Index
Best models for your workload
The best model depends on the job.
Highest agent-workflow score in the current SENA model index.
Highest document score in the current SENA model index.
Highest cost-adjusted value score in the current SENA model index.
Side by side
Compare AI models
| Metric |
|---|
| Your Fit score |
| SENA score |
| Coding |
| Reasoning |
| Agents |
| Documents |
| Extraction |
| Long context |
| Reliability |
| Speed |
| Value |
| Context window |
| Max output |
| Input price |
| Output price |
| Median latency |
Cost-adjusted performance
The best model isn't always the most expensive one.
SENA tracks capability alongside provider token pricing so teams can compare model quality against the economics of deploying each model. Verified runs add measured cost and latency when available.
Deployment economics
What will this model actually cost you?
Enter a rough production workload. SENA converts provider token prices into an estimated monthly API bill so you can compare capability and economics in the same decision.
Estimate uses the input/output token prices currently supplied with each benchmark model. Cached-input discounts and other provider-specific pricing tiers are not included.
Provider model catalog
A provider can have dozens of models.
The catalog answers “what can I use?” while the SENA benchmark answers “how do comparable models perform?”. Image, realtime, speech, embedding and safety models stay visible without being forced into an apples-to-oranges leaderboard.
Kept separate from SENA benchmark coverage so the page can scale as providers add model families.
GPT-5.6 Sol
Frontier professional-work model.
GPT-5.6 Terra
Balanced intelligence and cost.
GPT-5.6 Luna
Cost-sensitive GPT-5.6 option.
GPT-5.4
GPT-5.4 model in the general catalog.
GPT-5.4 Pro
GPT-5.4 model in the general catalog.
GPT-5.4 mini
GPT-5.4 model in the general catalog.
GPT-5.4 nano
GPT-5.4 model in the general catalog.
GPT-5.5
GPT-5.5 model in the general catalog.
GPT-5.5 Pro
GPT-5.5 model in the general catalog.
GPT-5.3-Codex
Codex model in the coding catalog.
GPT-5.2
GPT-5 model in the general catalog.
GPT-5.2 Pro
GPT-5 model in the general catalog.
GPT-5.1
GPT-5 model in the general catalog.
GPT-5
GPT-5 model in the general catalog.
GPT-5 mini
GPT-5 model in the general catalog.
GPT-5 nano
GPT-5 model in the general catalog.
GPT-5 Pro
GPT-5 model in the general catalog.
o3-pro
o-series model in the general catalog.
Methodology
How SENA evaluates AI models
Capability dimensions
SENA compares coding, reasoning, agents, documents, extraction and long-context capability separately.
Editorial baseline
Editorial rows are provisional SENA assessments used to make the index useful before a verified live run is available.
Official model economics
Context windows, output limits and token pricing are surfaced from provider documentation and kept separate from capability scores.
Clear verification status
Every model is labelled Editorial or Verified so readers can distinguish curated assessments from executed benchmark results.
Repeatable live tests
When verification is enabled, SENA runs the same task set across providers and stores the resulting scores, cost and latency.
Cost-aware ranking
Value is considered alongside capability so teams can compare strong models without ignoring deployment economics.
How to choose the right AI model
There is no universally best large language model. The right model depends on the workload, required accuracy, latency tolerance, context size, tool requirements and inference budget.
Coding-heavy applications may benefit from models that perform strongly on repository reasoning and tool use, while document-processing workloads may prioritize long context, retrieval accuracy and lower token cost.
SENA evaluates these dimensions separately so teams can compare models based on the work they actually need to perform rather than relying on a single generalized benchmark score.