| Mistral | Mistral Small 4 | Open Weight | Low-cost open-weight routing, extraction, and multimodal tasks | $0.15 | $0.60 | $0.75 per 2M tokens | 256K | Mistral's Apache 2.0 open model for lightweight multimodal and multilingual agentic workloads. Listed at $0.15 input / $0.60 output per 1M tokens on Mistral's API pricing page. | |
| OpenAI | GPT-5.6 Luna | Proprietary | High-volume OpenAI workloads that still need strong general quality | $0.20 | $1.20 | $1.40 per 2M tokens | 1.05M | OpenAI's efficient GPT-5.6 model for cost-sensitive, high-volume workloads. Short-context standard rate. Long-context rate is $0.40 input / $1.80 output per 1M tokens. Batch and Flex tiers are discounted. | |
| DeepSeek | DeepSeek V4 Flash | Open Weight | Ultra-low-cost production traffic and large-context experimentation | $0.44 | $1.32 | $1.76 per 2M tokens | 1M | DeepSeek's cost-focused V4 model with thinking and non-thinking modes plus MIT-licensed weights. Conservative peak cache-miss rate. Off-peak pricing is $0.22 input / $0.66 output per 1M tokens; weekends use off-peak rates all day from August 23, 2026. Peak cache-hit input is $0.014 per 1M tokens. | |
| Mistral | Mistral Large 3 | Open Weight | Open-weight production use with low hosted API cost | $0.50 | $1.50 | $2.00 per 2M tokens | 256K | Mistral describes Large 3 as an open-weight, general-purpose, flagship multimodal and multilingual model. Hosted API rate for mistral-large-latest; self-hosting costs depend on infrastructure. | |
| Google | Gemini 3.5 Flash-Lite | Proprietary | Budget Gemini routing, translation, and lightweight extraction | $0.30 | $2.50 | $2.80 per 2M tokens | 1M | Google's current most cost-efficient GA Gemini model for high-volume agentic tasks, translation, and simple data processing. Standard paid-tier text/image/video/audio input rate. Priority is $0.54 input / $4.50 output per 1M tokens. | |
| xAI | Grok Build 0.1 | Proprietary | Software-building workflows and lower-cost xAI usage | $1.00 | $2.00 | $3.00 per 2M tokens | 256K | xAI's build-focused model trained specifically for agentic coding workflows. Includes $0.20 cached input pricing. | |
| Google | Gemini 3 Flash Preview | Proprietary | Cost-efficient Gemini 3 apps that need speed and multimodal support | $0.50 | $3.00 | $3.50 per 2M tokens | 1M | Google's speed-focused Gemini 3 preview model with text, image, video, and audio input pricing. Standard paid-tier rate for text/image/video input. Audio input is $1 per 1M tokens. | |
| Google | Gemini 3.7 Flash | Proprietary | Fast multimodal apps, search-grounded work, and high-throughput reasoning | $0.75 | $3.75 | $4.50 per 2M tokens | 1M | Google's most capable Flash model for agentic workflows and multimodal reasoning. Promotional standard rate through December 31, 2026; it becomes $1.50 input / $7.50 output on January 1, 2027. Promotional priority pricing is $1.35 input / $6.75 output per 1M tokens. | |
| DeepSeek | DeepSeek V4 Pro | Open Weight | Very low-cost open-weight reasoning and long-context API workloads | $1.32 | $3.96 | $5.28 per 2M tokens | 1M | DeepSeek's higher-concurrency V4 model with thinking mode, JSON output, tool calls, and MIT-licensed weights. Conservative peak cache-miss rate. Off-peak pricing is $0.66 input / $1.98 output per 1M tokens; weekends use off-peak rates all day from August 23, 2026. Peak cache-hit input is $0.044 per 1M tokens. | |
| Anthropic | Claude Haiku 4.5 | Proprietary | Fast, high-volume Claude tasks, routing, chat, and extraction | $1.00 | $5.00 | $6.00 per 2M tokens | 200K | Anthropic lists Haiku 4.5 as the fastest Claude model with near-frontier intelligence. Lowest current Claude text model rate in the latest comparison table. | |
| xAI | Grok 4.6 | Proprietary | Current xAI flagship coding, chat, and agent workflows | $2.00 | $6.00 | $8.00 per 2M tokens | 500K | xAI's current frontier model for coding, agentic tasks, and knowledge work. Short-context rate for prompts below 200K tokens. At 200K prompt tokens or more, xAI charges $4 input / $12 output per 1M tokens; cached input is $0.50 short-context or $1 long-context. | |
| Mistral | Mistral Medium 3.5 | Open Weight | European open-weight option for reasoning, coding, and agent workflows | $1.50 | $7.50 | $9.00 per 2M tokens | 256K | Mistral's current open model for state-of-the-art performance, enterprise deployment, reasoning, coding, and agents. Listed at $1.50 input / $7.50 output per 1M tokens on Mistral's pricing page. | |
| Google | Gemini 2.5 Pro | Proprietary | Stable long-context Gemini workloads and complex reasoning | $1.25 | $10.00 | $11.25 per 2M tokens | 1M | Google still lists Gemini 2.5 Pro as an advanced model for complex tasks, deep reasoning, and coding. Rate for prompts <= 200K tokens. Prompts > 200K are $2.50 input / $15 output per 1M tokens. | |
| Anthropic | Claude Sonnet 5 | Proprietary | Production agents, coding assistants, analysis, and quality writing | $2.00 | $10.00 | $12.00 per 2M tokens | 1M | Anthropic describes Sonnet 5 as the best combination of speed and intelligence in the current Claude comparison table. Anthropic made the $2 input / $10 output per MTok rate standard on August 10, 2026 and cancelled the previously announced September increase. | |
| OpenAI | GPT-5.6 Terra | Proprietary | Lower-cost OpenAI frontier workloads and production agents | $2.00 | $12.00 | $14.00 per 2M tokens | 1.05M | OpenAI's GPT-5.6 model that balances intelligence and cost for production workloads. Short-context standard rate. Long-context rate is $4 input / $18 output per 1M tokens. Batch and Flex tiers are discounted. | |
| Google | Gemini 3.1 Pro Preview | Proprietary | Long-context multimodal reasoning and Google ecosystem integrations | $2.00 | $12.00 | $14.00 per 2M tokens | 1M | Google's Gemini 3.1 Pro preview for multimodal understanding, agentic capability, and coding. Standard rate for prompts <= 200K tokens. Prompts > 200K are $4 input / $18 output per 1M tokens. | |
| Anthropic | Claude Opus 5 | Proprietary | Premium coding, deep reasoning, and enterprise knowledge work | $5.00 | $25.00 | $30.00 per 2M tokens | 1M | Anthropic's most capable Opus-tier model for complex reasoning, agentic coding, and enterprise knowledge work. Model table lists $5 input / $25 output per million tokens. Microsoft Foundry context may be 200K. | |
| OpenAI | GPT-5.6 Sol | Proprietary | Frontier OpenAI reasoning, agent workflows, and high-quality generation | $5.00 | $30.00 | $35.00 per 2M tokens | 1.05M | OpenAI's current flagship GPT-5.6 model for complex reasoning, coding, and professional work. Short-context standard rate. Long-context rate is $10 input / $45 output per 1M tokens. Batch and Flex tiers are discounted. | |
| Anthropic | Claude Fable 5 | Proprietary | Highest-end Anthropic reasoning, coding, and long-running agent tasks | $10.00 | $50.00 | $60.00 per 2M tokens | 1M | Anthropic's most capable widely released Claude model for demanding reasoning and long-horizon agentic work. Generally available on the Claude API as of June 9, 2026. Mythos 5 has the same listed rate but limited availability. | |
| OpenAI | gpt-oss-120b | Open Weight | Customizable open-weight reasoning on H100-class infrastructure | Self-host | Self-host | Self-host | 131K | OpenAI's most powerful open-weight reasoning model; 117B parameters with 5.1B active parameters and Apache 2.0 licensing. Self-hosted or third-party-hosted cost depends on infrastructure. Official OpenAI pricing page does not list a token API rate for this model. | |
| OpenAI | gpt-oss-20b | Open Weight | Local and lower-infrastructure open-weight reasoning experiments | Self-host | Self-host | Self-host | 131K | Smaller OpenAI open-weight reasoning model released with the gpt-oss family under Apache 2.0 licensing. Self-hosted or third-party-hosted cost depends on infrastructure. Official OpenAI pricing page does not list a token API rate for this model. | |
| Google | Gemma 4 31B | Open Weight | Open-weight Google model for local or private-cloud reasoning and coding | Self-host | Self-host | Self-host | 256K | Google DeepMind's 31B dense Gemma 4 open-weight model for coding, reasoning, and multimodal understanding. Self-hosted or Vertex/third-party deployment cost depends on infrastructure; Gemma 4 is distributed as open weights. | |
| Meta | Llama 4 Scout | Open Weight | Very long-context open-weight multimodal workloads | Self-host | Self-host | Self-host | 10M | Meta's 17B-active, 109B-total open-weight multimodal MoE model with a very long advertised context window. Self-hosted or partner-hosted cost depends on infrastructure. Distributed under the Llama 4 Community License. | |
| Meta | Llama 4 Maverick | Open Weight | High-quality open-weight multimodal chat, reasoning, and coding | Self-host | Self-host | Self-host | 1M | Meta's 17B-active, 400B-total open-weight multimodal MoE model focused on strong quality-to-cost performance. Self-hosted or partner-hosted cost depends on infrastructure. Distributed under the Llama 4 Community License. | |
| Qwen | Qwen3-235B-A22B-Instruct-2507 | Open Weight | Open-weight multilingual instruction following, coding, and tool-use workloads | Self-host | Self-host | Self-host | 256K | Updated Qwen3 235B MoE instruct model with 22B active parameters, Apache 2.0 licensing, and improved long-context understanding. Self-hosted or third-party-hosted cost depends on infrastructure. | |