SENA Model IndexUpdated regularly
SENA Models

Which AI model is actually better?

Start with SENA's neutral benchmark, then configure what matters to your workload. Weight coding, reasoning, agents, reliability, speed and cost to find the model that best fits how you actually plan to use AI.

SENA testing philosophy

We test what models can actually do — not just what benchmark scores say they can do.

Read our methodology

Start with the job

What do you need the model to do?

Pick a workload and SENA will re-weight the benchmark for that job. You can fine-tune every weight and constraint afterwards.

Configure your benchmark

What are you actually choosing a model for?

The SENA score remains the neutral benchmark. Your Fit score re-weights the same model data around your workload, then applies practical constraints such as context window and input price.

Importance weights
0 = ignore · 10 = critical
Coding5/10
Reasoning5/10
Agents5/10
Documents4/10
Extraction3/10
Long context3/10
Reliability5/10
Speed2/10
Value3/10
Hard requirements
Models that fail these are excluded.
Recommended for your configuration
No models match every requirement
Relax the context, price or verification requirement to see a recommendation.

SENA Model Index

Best models for your workload

Quick answers

The best model depends on the job.

Best for coding

Highest coding score in the current SENA model index.

Best for agents

Highest agent-workflow score in the current SENA model index.

Best for documents

Highest document score in the current SENA model index.

Best value

Highest cost-adjusted value score in the current SENA model index.

Side by side

Compare AI models

Metric
Your Fit score
SENA score
Coding
Reasoning
Agents
Documents
Extraction
Long context
Reliability
Speed
Value
Context window
Max output
Input price
Output price
Median latency

Cost-adjusted performance

The best model isn't always the most expensive one.

SENA tracks capability alongside provider token pricing so teams can compare model quality against the economics of deploying each model. Verified runs add measured cost and latency when available.

Deployment economics

What will this model actually cost you?

Enter a rough production workload. SENA converts provider token prices into an estimated monthly API bill so you can compare capability and economics in the same decision.

Estimate uses the input/output token prices currently supplied with each benchmark model. Cached-input discounts and other provider-specific pricing tiers are not included.

Model
Per 1K req.
Est. / month
Pricing is not available for the current benchmark rows.

Provider model catalog

A provider can have dozens of models.

The catalog answers “what can I use?” while the SENA benchmark answers “how do comparable models perform?”. Image, realtime, speech, embedding and safety models stay visible without being forced into an apples-to-oranges leaderboard.

OpenAI catalog reference
50 current entries seeded

Kept separate from SENA benchmark coverage so the page can scale as providers add model families.

Official catalog ↗
50 models shown0 currently benchmarked by SENA
OpenAI/Frontier

GPT-5.6 Sol

GPT-5.6

Frontier professional-work model.

OpenAI/Frontier

GPT-5.6 Terra

GPT-5.6

Balanced intelligence and cost.

OpenAI/Frontier

GPT-5.6 Luna

GPT-5.6

Cost-sensitive GPT-5.6 option.

OpenAI/General

GPT-5.4

GPT-5.4

GPT-5.4 model in the general catalog.

OpenAI/General

GPT-5.4 Pro

GPT-5.4

GPT-5.4 model in the general catalog.

OpenAI/General

GPT-5.4 mini

GPT-5.4

GPT-5.4 model in the general catalog.

OpenAI/General

GPT-5.4 nano

GPT-5.4

GPT-5.4 model in the general catalog.

OpenAI/General

GPT-5.5

GPT-5.5

GPT-5.5 model in the general catalog.

OpenAI/General

GPT-5.5 Pro

GPT-5.5

GPT-5.5 model in the general catalog.

OpenAI/Coding

GPT-5.3-Codex

Codex

Codex model in the coding catalog.

OpenAI/General

GPT-5.2

GPT-5

GPT-5 model in the general catalog.

OpenAI/General

GPT-5.2 Pro

GPT-5

GPT-5 model in the general catalog.

OpenAI/General

GPT-5.1

GPT-5

GPT-5 model in the general catalog.

OpenAI/General

GPT-5

GPT-5

GPT-5 model in the general catalog.

OpenAI/General

GPT-5 mini

GPT-5

GPT-5 model in the general catalog.

OpenAI/General

GPT-5 nano

GPT-5

GPT-5 model in the general catalog.

OpenAI/General

GPT-5 Pro

GPT-5

GPT-5 model in the general catalog.

OpenAI/General

o3-pro

o-series

o-series model in the general catalog.

Methodology

How SENA evaluates AI models

01

Capability dimensions

SENA compares coding, reasoning, agents, documents, extraction and long-context capability separately.

02

Editorial baseline

Editorial rows are provisional SENA assessments used to make the index useful before a verified live run is available.

03

Official model economics

Context windows, output limits and token pricing are surfaced from provider documentation and kept separate from capability scores.

04

Clear verification status

Every model is labelled Editorial or Verified so readers can distinguish curated assessments from executed benchmark results.

05

Repeatable live tests

When verification is enabled, SENA runs the same task set across providers and stores the resulting scores, cost and latency.

06

Cost-aware ranking

Value is considered alongside capability so teams can compare strong models without ignoring deployment economics.

Task templates
Live model calls
Live runs / task
Data updated

How to choose the right AI model

There is no universally best large language model. The right model depends on the workload, required accuracy, latency tolerance, context size, tool requirements and inference budget.

Coding-heavy applications may benefit from models that perform strongly on repository reasoning and tool use, while document-processing workloads may prioritize long context, retrieval accuracy and lower token cost.

SENA evaluates these dimensions separately so teams can compare models based on the work they actually need to perform rather than relying on a single generalized benchmark score.

SENA Intelligence

Find the best model for your actual workload.

Configure my benchmark →