OpenRouter — 424 modèles de l’API
Source unique : API OpenRouter. 424 modèles. Aucun niveau Agentique, Code, Traduction ou benchmark n'est ajouté ici. Les champs affichés sont ceux collectés depuis la fiche API : prix, contexte, sortie maximale, modalités, paramètres supportés et description. N/D = le provider ne publie pas le champ.
Source unique : API OpenRouter. 424 modèles. Aucun niveau Agentique, Code, Traduction ou benchmark n'est ajouté ici. Les champs affichés sont ceux collectés depuis la fiche API : prix, contexte, sortie maximale, modalités, paramètres supportés et description. N/D = le provider ne publie pas le champ.
OpenRouter — 424 modèles de l’API
| # | Modèle / ID | Prix $/1M in/out | Contexte | Sortie max | Modalités | Paramètres API | Description source |
|---|---|---|---|---|---|---|---|
| 1 | Meta: Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | $0.100 / $0.200 | 1.0M | 944k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, repetition_penalty, response_format, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information... |
| 2 | Meta: Muse Spark 1.3meta/muse-spark-1.3 | $1.250 / $4.250 | 1.0M | 944k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, repetition_penalty, response_format, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through... |
| 3 | Google: Gemini 3.8 Flashgoogle/gemini-3.8-flash | $0.750 / $3.750 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. |
| 4 | Google: Gemini 3.8 Flash (batch)google/gemini-3.8-flash:batch | $0.375 / $1.875 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, tool_choice, tools | Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. |
| 5 | Anthropic: Claude Fable 5.1anthropic/claude-fable-5.1 | $10.000 / $50.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tools, verbosity | Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual... |
| 6 | Anthropic: Claude Fable 5.1 (batch)anthropic/claude-fable-5.1:batch | $5.000 / $25.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual... |
| 7 | Inception: Mercury 2.5 Previewinception/mercury-2.5-preview | $0.040 / $0.150 | 260k | 66k | text->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools | Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving... |
| 8 | IBM: Granite 4.2 8Bibm-granite/granite-4.2-8b | $0.100 / $0.150 | 131k | 118k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,... |
| 9 | Tencent: Hy4 previewtencent/hy4-preview | $0.834 / $2.501 | 1.0M | 64k | text->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools | Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that... |
| 10 | Ling 3.0 Flash Fin (free)inclusionai/ling-3.0-flash-fin:free | FREE | 262k | 33k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, repetition_penalty, seed, stop, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment... |
| 11 | Z.ai: GLM Flash Latest~z-ai/glm-flash-latest | $0.075 / $0.250 | 1.3M | 944k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | This model always redirects to the latest model in the GLM Flash family. |
| 12 | Qwen: Qwen3.8 Flashqwen/qwen3.8-flash | $0.150 / $0.470 | 1.0M | 131k | text+image+video->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. |
| 13 | Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash | $0.075 / $0.250 | 1.3M | 131k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while... |
| 14 | Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch | $0.150 / $0.500 | 1.0M | 944k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while... |
| 15 | Meta: Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributor | $0.100 / $0.200 | 1.0M | 944k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, repetition_penalty, response_format, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark... |
| 16 | DeepSeek: DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp | $0.220 / $0.660 | 1.0M | 384k | text+image->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,... |
| 17 | Tencent: Hy-MT2-1.8Btencent/hy-mt2-1.8b | $0.044 / $0.177 | 8k | 4k | text->text | max_completion_tokens, max_tokens, stop, temperature | Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided... |
| 18 | Tencent: Hy-MT2-30B-A3Btencent/hy-mt2-30b-a3b | $0.074 / $0.295 | 8k | 4k | text->text | max_completion_tokens, max_tokens, response_format, stop, structured_outputs, temperature | Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and... |
| 19 | Z.ai: GLM Latest~z-ai/glm-latest | $1.092 / $3.432 | 1.3M | 944k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | This model always redirects to the latest GLM model from Z.ai. |
| 20 | Tencent: Hy-MT2-7Btencent/hy-mt2-7b | $0.074 / $0.295 | 8k | 4k | text->text | max_completion_tokens, max_tokens, response_format, stop, structured_outputs, temperature | Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation. |
| 21 | Z.ai: GLM 5.3z-ai/glm-5.3 | $1.400 / $4.400 | 1.3M | 131k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves... |
| 22 | Qwen: Qwen3.8 27Bqwen/qwen3.8-27b | $0.425 / $2.550 | 1.0M | 131k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be... |
| 23 | Dots Studio: Dots3-Note Preview (free)dots-studio/dots-3-note-preview:free | FREE | 512k | 461k | text+image->text | include_reasoning, max_tokens, reasoning, response_format, structured_outputs, temperature, tool_choice, tools, top_p | Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is... |
| 24 | Google: Gemini 3.7 Flashgoogle/gemini-3.7-flash | $0.750 / $3.750 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step... |
| 25 | Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch | $0.375 / $1.875 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, tool_choice, tools | Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step... |
| 26 | ByteDance Seed: Seed 2.1 Turbobytedance-seed/seed-2-1-turbo | $0.500 / $2.500 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p | Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and... |
| 27 | Qwen: Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b | $2.000 / $6.000 | 1.0M | 262k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is... |
| 28 | Qwen: Qwen3.8 2.4T A95B (batch)qwen/qwen3.8-2.4t-a95b:batch | $2.000 / $6.000 | 1.0M | 909k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is... |
| 29 | ByteDance Seed: Seed-2.0-Codebytedance-seed/seed-2.0-code | $0.500 / $3.000 | 262k | 131k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p | Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude... |
| 30 | DeepSeek: DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | $1.115 / $3.346 | 1.0M | 384k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro. |
| 31 | DeepSeek: DeepSeek V4 Pro 0813 (batch)deepseek/deepseek-v4-pro-0813:batch | $1.320 / $3.960 | 1.0M | 944k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro. |
| 32 | SpaceXAI: Grok 4.6x-ai/grok-4.6 | $2.000 / $6.000 | 500k | 450k | text+image+file->text | include_reasoning, logprobs, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM. |
| 33 | LiquidAI: LFM2.5-2.6B (free)liquid/lfm-2.5-2.6b:free | FREE | 66k | 8k | text->text | frequency_penalty, include_reasoning, logprobs, max_completion_tokens, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or... |
| 34 | NVIDIA: Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning | $0.080 / $0.200 | 262k | 131k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that... |
| 35 | NVIDIA: Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | FREE | 1.0M | 66k | text->text | include_reasoning, max_tokens, reasoning, seed, temperature, tool_choice, tools, top_p | NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that... |
| 36 | Sakana: Sakana Namazusakana/sakana-namazu | $0.950 / $4.000 | 262k | 66k | text+image+file->text | include_reasoning, reasoning, reasoning_effort, structured_outputs, tool_choice, tools, web_search_options | Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,... |
| 37 | Upstage: Solar Pro 4upstage/solar-pro4 | $0.030 / $0.120 | 524k | 131k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, response_format, structured_outputs, temperature, tool_choice, tools, top_p | Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive... |
| 38 | Meta: Muse Glimmer 30Bmeta/muse-glimmer-30b | $0.300 / $1.100 | 131k | 118k | text+image->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon... |
| 39 | Meta: Muse Glimmer 30B (batch)meta/muse-glimmer-30b:batch | $0.350 / $1.500 | 131k | 118k | text+image->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon... |
| 40 | Meta: Muse Spark 1.2meta/muse-spark-1.2 | $1.250 / $4.250 | 1.0M | 944k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, repetition_penalty, response_format, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context... |
| 41 | Qwen: Qwen3.8 Maxqwen/qwen3.8-max | $2.000 / $6.000 | 1.0M | 131k | text+image+video->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,... |
| 42 | DeepSeek V4 Flash Latest~deepseek/deepseek-v4-flash-latest | $0.050 / $0.160 | 1.3M | 393k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, top_p | This model always redirects to the latest model in the DeepSeek V4 Flash family. |
| 43 | DeepSeek: DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | $0.065 / $0.180 | 1.3M | 944k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, top_p | DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.... |
| 44 | DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek/deepseek-v4-flash-0731:batch | $0.140 / $0.280 | 1.0M | 944k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.... |
| 45 | Thinking Machines: Inkling Smallthinkingmachines/inkling-small | $0.450 / $1.200 | 1.0M | 262k | text+image+audio->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, seed, stop, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of... |
| 46 | Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | $0.500 / $1.200 | 524k | 472k | text+image+audio->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, stop, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of... |
| 47 | Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | FREE | 1.0M | 262k | text+image+audio->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, reasoning_effort, seed, stop, temperature, tools, top_p | Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of... |
| 48 | Qwen: Qwen3.7 Flashqwen/qwen3.7-flash | $0.030 / $0.130 | 1.0M | 66k | text+image+video->text | include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, temperature, tool_choice, tools, top_logprobs, top_p | Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world... |
| 49 | Claude Opus 5anthropic/claude-opus-5 | $5.000 / $25.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, verbosity | Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis... |
| 50 | Claude Opus 5 (batch)anthropic/claude-opus-5:batch | $2.500 / $12.500 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis... |
| 51 | Ling-3.0-flashinclusionai/ling-3.0-flash | $0.021 / $0.063 | 262k | 33k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_logprobs, top_p | *Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers... |
| 52 | Poolside: Laguna S 2.1poolside/laguna-s-2.1 | $0.090 / $0.180 | 1.0M | 131k | text->text | include_reasoning, max_tokens, reasoning, temperature, tool_choice, tools | Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and... |
| 53 | Poolside: Laguna S 2.1 (free)poolside/laguna-s-2.1:free | FREE | 262k | 33k | text->text | include_reasoning, max_tokens, reasoning, temperature, tool_choice, tools | Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and... |
| 54 | Google: Gemini 3.6 Flashgoogle/gemini-3.6-flash | $0.750 / $3.750 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and... |
| 55 | Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | $0.375 / $1.875 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, tool_choice, tools | Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and... |
| 56 | Google: Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | $0.300 / $2.500 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows. |
| 57 | Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | $0.150 / $1.250 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, tool_choice, tools | Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows. |
| 58 | Meituan: LongCat 2.0meituan/longcat-2.0 | $0.300 / $1.200 | 1.0M | 262k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, seed, stop, temperature, tool_choice, tools, top_k, top_p | LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic... |
| 59 | Thinking Machines: Inklingthinkingmachines/inkling | $1.000 / $4.050 | 1.0M | 472k | text+image+audio->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,... |
| 60 | Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | $1.000 / $4.050 | 524k | 472k | text+image+audio->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, stop, temperature, tool_choice, tools, top_k, top_p | Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,... |
| 61 | Thinking Machines: Inkling (free)thinkingmachines/inkling:free | FREE | 1.0M | 262k | text+image+audio->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, reasoning_effort, seed, stop, temperature, tools, top_p | Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,... |
| 62 | Auto Router (Beta)openrouter/auto-beta | N/D | 2.0M | N/D | text+image+file+audio+video->text+image | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, prediction, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, top_p, web_search_options | Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your... |
| 63 | MoonshotAI: Kimi K3moonshotai/kimi-k3 | $3.000 / $15.000 | 1.0M | 944k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at... |
| 64 | MoonshotAI: Kimi K3 (batch)moonshotai/kimi-k3:batch | $3.000 / $15.000 | 1.0M | 944k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at... |
| 65 | Meta: Muse Spark 1.1meta/muse-spark-1.1 | $1.250 / $4.250 | 1.0M | 944k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, repetition_penalty, response_format, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context... |
| 66 | Kwaipilot: KAT-Coder-Pro V2.5kwaipilot/kat-coder-pro-v2.5 | $0.740 / $2.960 | 262k | 236k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make... |
| 67 | OpenAI: GPT-5.6 Luna Proopenai/gpt-5.6-luna-pro | $0.200 / $1.200 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode |
| 68 | OpenAI: GPT-5.6 Luna Pro (batch)openai/gpt-5.6-luna-pro:batch | $0.100 / $0.600 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode |
| 69 | OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna | $0.200 / $1.200 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for... |
| 70 | OpenAI: GPT-5.6 Luna (batch)openai/gpt-5.6-luna:batch | $0.100 / $0.600 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for... |
| 71 | OpenAI: GPT-5.6 Terra Proopenai/gpt-5.6-terra-pro | $2.000 / $12.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode |
| 72 | OpenAI: GPT-5.6 Terra Pro (batch)openai/gpt-5.6-terra-pro:batch | $1.000 / $6.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode |
| 73 | OpenAI: GPT-5.6 Terraopenai/gpt-5.6-terra | $2.000 / $12.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic... |
| 74 | OpenAI: GPT-5.6 Terra (batch)openai/gpt-5.6-terra:batch | $1.000 / $6.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic... |
| 75 | OpenAI: GPT-5.6 Sol Proopenai/gpt-5.6-sol-pro | $2.000 / $10.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode |
| 76 | OpenAI: GPT-5.6 Sol Pro (batch)openai/gpt-5.6-sol-pro:batch | $1.000 / $5.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode |
| 77 | OpenAI: GPT-5.6 Solopenai/gpt-5.6-sol | $2.000 / $10.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks... |
| 78 | OpenAI: GPT-5.6 Sol (batch)openai/gpt-5.6-sol:batch | $1.000 / $5.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks... |
| 79 | SpaceXAI: Grok 4.5x-ai/grok-4.5 | $2.000 / $6.000 | 500k | 450k | text+image+file->text | include_reasoning, logprobs, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. |
| 80 | xAI: Grok Latest~x-ai/grok-latest | $2.000 / $6.000 | 500k | 450k | text+image+file->text | include_reasoning, logprobs, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | This model always redirects to the latest Grok model from xAI. |
| 81 | AionLabs: Aion-3.0-Miniaion-labs/aion-3.0-mini | $0.700 / $1.400 | 131k | 33k | text->text | include_reasoning, max_tokens, reasoning, response_format, temperature, tool_choice, tools, top_p | Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each... |
| 82 | AionLabs: Aion-3.0aion-labs/aion-3.0 | $3.000 / $6.000 | 131k | 33k | text->text | include_reasoning, max_tokens, reasoning, response_format, temperature, tool_choice, tools, top_p | Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute... |
| 83 | Tencent: Hy3tencent/hy3 | $0.132 / $0.528 | 262k | 128k | text->text | frequency_penalty, include_reasoning, logit_bias, max_completion_tokens, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:... |
| 84 | Poolside: Laguna XS 2.1poolside/laguna-xs-2.1 | $0.060 / $0.120 | 262k | 33k | text->text | include_reasoning, max_tokens, reasoning, temperature, tool_choice, tools | Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines... |
| 85 | Poolside: Laguna XS 2.1 (free)poolside/laguna-xs-2.1:free | FREE | 262k | 33k | text->text | include_reasoning, max_tokens, reasoning, temperature, tool_choice, tools | Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines... |
| 86 | Anthropic: Claude Sonnet 5anthropic/claude-sonnet-5 | $2.000 / $10.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,... |
| 87 | Anthropic: Claude Sonnet 5 (batch)anthropic/claude-sonnet-5:batch | $1.000 / $5.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,... |
| 88 | Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)google/gemini-3.1-flash-lite-image | $0.250 / $1.500 | 66k | 59k | text+image->text+image | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, temperature, top_p | Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation... |
| 89 | Nex AGI: Nex-N2-Mininex-agi/nex-n2-mini | $0.025 / $0.100 | 262k | 236k | text+image->text | include_reasoning, logprobs, max_tokens, reasoning, response_format, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is built for coding, tool use,... |
| 90 | Sakana: Fugu Ultrasakana/fugu-ultra | $5.000 / $30.000 | 1.0M | 128k | text+image->text | include_reasoning, reasoning, reasoning_effort, structured_outputs, tool_choice, tools, web_search_options | Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route... |
| 91 | Google: Nano Banana 2 (Gemini 3.1 Flash Image)google/gemini-3.1-flash-image | $0.500 / $3.000 | 131k | 33k | text+image->text+image | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, top_p | Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced... |
| 92 | Google: Nano Banana Pro (Gemini 3 Pro Image)google/gemini-3-pro-image | $2.000 / $12.000 | 131k | 33k | text+image->text+image | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and... |
| 93 | Cohere: North Mini Code (free)cohere/north-mini-code:free | FREE | 256k | 64k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, seed, stop, temperature, tool_choice, tools, top_k, top_p | North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized... |
| 94 | Z.ai: GLM 5.2z-ai/glm-5.2 | $0.966 / $3.036 | 1.0M | 131k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,... |
| 95 | Z.ai: GLM 5.2 (free)z-ai/glm-5.2:free | FREE | 256k | 230k | text->text | frequency_penalty, include_reasoning, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,... |
| 96 | OpenRouter: Fusionopenrouter/fusion | N/D | 1.0M | N/D | text->text | N/D | Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a... |
| 97 | MoonshotAI: Kimi K2.7 Codemoonshotai/kimi-k2.7-code | $0.660 / $3.400 | 262k | 236k | text+image->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts... |
| 98 | Anthropic: Claude Fable Latest~anthropic/claude-fable-latest | $10.000 / $50.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tools, verbosity | This model always redirects to the latest model in the Claude Fable family. |
| 99 | Anthropic: Claude Fable 5anthropic/claude-fable-5 | $10.000 / $50.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and... |
| 100 | Anthropic: Claude Fable 5 (batch)anthropic/claude-fable-5:batch | $5.000 / $25.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and... |
| 101 | Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | $0.250 / $1.000 | 262k | 236k | text+image->text | frequency_penalty, include_reasoning, logprobs, max_tokens, reasoning, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces... |
| 102 | NVIDIA: Nemotron 3.5 Content Safety (free)nvidia/nemotron-3.5-content-safety:free | FREE | 128k | 8k | text+image->text | include_reasoning, max_tokens, reasoning, seed, temperature, top_p | NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting... |
| 103 | NVIDIA: Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | $0.600 / $2.400 | 262k | 183k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it... |
| 104 | NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | FREE | 1.0M | 66k | text->text | include_reasoning, max_tokens, reasoning, reasoning_effort, seed, temperature, tool_choice, tools, top_p | NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it... |
| 105 | Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | $0.320 / $1.280 | 1.0M | 131k | text+image->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its... |
| 106 | MiniMax: MiniMax M3minimax/minimax-m3 | $0.300 / $1.200 | 1.0M | 512k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,... |
| 107 | MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch | $0.300 / $1.200 | 524k | 472k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,... |
| 108 | MiniMax: MiniMax M3 (free)minimax/minimax-m3:free | FREE | 1.0M | 944k | text+image+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, temperature, tool_choice, tools, top_p | MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,... |
| 109 | StepFun: Step 3.7 Flashstepfun/step-3.7-flash | $0.200 / $1.150 | 262k | 230k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters... |
| 110 | Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8 | $5.000 / $25.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, verbosity | Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token... |
| 111 | Anthropic: Claude Opus 4.8 (batch)anthropic/claude-opus-4.8:batch | $2.500 / $12.500 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token... |
| 112 | Qwen: Qwen3.7 Maxqwen/qwen3.7-max | $1.475 / $4.425 | 1.0M | 131k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,... |
| 113 | SpaceXAI: Grok Build 0.1x-ai/grok-build-0.1 | $1.000 / $2.000 | 256k | 230k | text+image+file->text | include_reasoning, logprobs, max_tokens, reasoning, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding... |
| 114 | Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash | $1.500 / $9.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution... |
| 115 | Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | $0.750 / $4.500 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution... |
| 116 | Perceptron: Perceptron Mk1perceptron/perceptron-mk1 | $0.150 / $1.500 | 33k | 8k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, structured_outputs, temperature, top_k, top_p | Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding... |
| 117 | Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | $0.250 / $1.500 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... |
| 118 | Google: Gemini 3.1 Flash Lite (batch)google/gemini-3.1-flash-lite:batch | $0.125 / $0.750 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... |
| 119 | OpenAI: GPT Chat Latestopenai/gpt-chat-latest | $5.000 / $30.000 | 400k | 128k | text+image+file->text | max_tokens, response_format, seed, structured_outputs, tool_choice, tools | GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates... |
| 120 | SpaceXAI: Grok 4.3x-ai/grok-4.3 | $1.250 / $2.500 | 1.0M | 900k | text+image+file->text | include_reasoning, logprobs, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual... |
| 121 | IBM: Granite 4.1 8Bibm-granite/granite-4.1-8b | $0.050 / $0.100 | 131k | 118k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window and is designed for enterprise tasks... |
| 122 | Mistral: Mistral Medium 3.5mistralai/mistral-medium-3-5 | $1.500 / $7.500 | 262k | 210k | text+image+file->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex... |
| 123 | Mistral: Mistral Medium 3.5 (batch)mistralai/mistral-medium-3-5:batch | $0.750 / $3.750 | 262k | 210k | text+image+file->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex... |
| 124 | NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | FREE | 256k | 66k | text+image+audio+video->text | include_reasoning, max_tokens, reasoning, seed, temperature, tool_choice, tools, top_p | NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and... |
| 125 | Anthropic Claude Haiku Latest~anthropic/claude-haiku-latest | $1.000 / $5.000 | 200k | 64k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | This model always redirects to the latest model in the Anthropic Claude Haiku family. |
| 126 | OpenAI GPT Mini Latest~openai/gpt-mini-latest | $0.750 / $4.500 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | This model always redirects to the latest model in the OpenAI GPT Mini family. |
| 127 | Google Gemini Pro Latest~google/gemini-pro-latest | $2.000 / $12.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | This model always redirects to the latest model in the Google Gemini Pro family. |
| 128 | MoonshotAI Kimi Latest~moonshotai/kimi-latest | $2.550 / $12.750 | 1.0M | 944k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | This model always redirects to the latest model in the MoonshotAI Kimi family. |
| 129 | Google Gemini Flash Latest~google/gemini-flash-latest | $0.750 / $3.750 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | This model always redirects to the latest model in the Google Gemini Flash family. |
| 130 | Anthropic Claude Sonnet Latest~anthropic/claude-sonnet-latest | $2.000 / $10.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | This model always redirects to the latest model in the Anthropic Claude Sonnet family. |
| 131 | OpenAI GPT Latest~openai/gpt-latest | $2.000 / $10.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | This model always redirects to the latest model in the OpenAI GPT family. |
| 132 | Qwen: Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | $0.300 / $1.800 | 1.0M | 66k | text+image+video->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This... |
| 133 | Qwen: Qwen3.6 Flashqwen/qwen3.6-flash | $0.188 / $1.125 | 1.0M | 66k | text+image+video->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in... |
| 134 | Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | $0.100 / $0.900 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated... |
| 135 | Qwen: Qwen3.6 Max Previewqwen/qwen3.6-max-preview | $1.027 / $6.162 | 262k | 66k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and... |
| 136 | Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | $0.600 / $3.600 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs... |
| 137 | OpenAI: GPT-5.5 Proopenai/gpt-5.5-pro | $30.000 / $180.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for... |
| 138 | OpenAI: GPT-5.5 Pro (batch)openai/gpt-5.5-pro:batch | $15.000 / $90.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for... |
| 139 | OpenAI: GPT-5.5openai/gpt-5.5 | $5.000 / $30.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token... |
| 140 | OpenAI: GPT-5.5 (batch)openai/gpt-5.5:batch | $2.500 / $15.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token... |
| 141 | DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | $1.042 / $2.085 | 1.0M | 384k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_completion_tokens, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,... |
| 142 | DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | $0.089 / $0.177 | 1.0M | 384k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_completion_tokens, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, top_p | DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and... |
| 143 | Tencent: Hy3 previewtencent/hy3-preview | $0.180 / $0.600 | 262k | 236k | text->text | include_reasoning, max_tokens, reasoning, reasoning_effort, seed, temperature, tool_choice, tools, top_p | Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to... |
| 144 | Xiaomi: MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | $0.435 / $0.870 | 1.1M | 131k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro.... |
| 145 | Xiaomi: MiMo-V2.5xiaomi/mimo-v2.5 | $0.140 / $0.280 | 1.1M | 131k | text+image+audio+video->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding... |
| 146 | OpenAI: GPT-5.4 Image 2openai/gpt-5.4-image-2 | $8.000 / $15.000 | 272k | 128k | text+image+file->text+image | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, presence_penalty, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, top_logprobs | [GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and... |
| 147 | Anthropic: Claude Opus Latest~anthropic/claude-opus-latest | $5.000 / $25.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, verbosity | This model always redirects to the latest model in the Claude Opus family. |
| 148 | Pareto Code Routeropenrouter/pareto-code | N/D | 2.0M | N/D | text->text | N/D | The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how... |
| 149 | MoonshotAI: Kimi K2.6moonshotai/kimi-k2.6 | $0.950 / $4.000 | 262k | 236k | text+image->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and... |
| 150 | Anthropic: Claude Opus 4.7anthropic/claude-opus-4.7 | $5.000 / $25.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on... |
| 151 | Anthropic: Claude Opus 4.7 (batch)anthropic/claude-opus-4.7:batch | $2.500 / $12.500 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, tool_choice, tools, verbosity | Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on... |
| 152 | Z.ai: GLM 5.1z-ai/glm-5.1 | $0.966 / $3.036 | 205k | 128k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on... |
| 153 | Google: Gemma 4 26B A4B google/gemma-4-26b-a4b-it | $0.070 / $0.340 | 262k | 16k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at... |
| 154 | Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | FREE | 262k | 33k | text+image+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, temperature, tool_choice, tools, top_p | Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at... |
| 155 | Google: Gemma 4 31Bgoogle/gemma-4-31b-it | $0.090 / $0.340 | 262k | 16k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function... |
| 156 | Google: Gemma 4 31B (batch)google/gemma-4-31b-it:batch | $0.390 / $0.970 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function... |
| 157 | Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | FREE | 262k | 33k | text+image+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, temperature, tool_choice, tools, top_p | Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function... |
| 158 | Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | $0.325 / $1.950 | 1.0M | 66k | text+image+video->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers... |
| 159 | Z.ai: GLM 5V Turboz-ai/glm-5v-turbo | $1.200 / $4.000 | 203k | 131k | text+image+video->text | include_reasoning, max_tokens, reasoning, response_format, temperature, tool_choice, tools, top_k, top_p | GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,... |
| 160 | Arcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinking | $0.250 / $0.800 | 262k | 80k | text->text | include_reasoning, max_tokens, reasoning, temperature, tool_choice, tools, top_k, top_p | Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7... |
| 161 | SpaceXAI: Grok 4.20 Multi-Agentx-ai/grok-4.20-multi-agent | $1.250 / $2.500 | 2.0M | 1.8M | text+image+file->text | include_reasoning, logprobs, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, top_logprobs, top_p | Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information... |
| 162 | SpaceXAI: Grok 4.20x-ai/grok-4.20 | $1.250 / $2.500 | 2.0M | 1.8M | text+image+file->text | include_reasoning, logprobs, max_tokens, reasoning, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering... |
| 163 | Google: Lyria 3 Pro Previewgoogle/lyria-3-pro-preview | FREE | 1.0M | 66k | text+image->text+audio | max_tokens, response_format, seed, temperature, top_p | Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz... |
| 164 | Google: Lyria 3 Clip Previewgoogle/lyria-3-clip-preview | FREE | 1.0M | 66k | text+image->text+audio | max_tokens, response_format, seed, temperature, top_p | 30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate... |
| 165 | Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | $0.300 / $1.200 | 262k | 144k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,... |
| 166 | Reka Edgerekaai/reka-edge | $0.100 / $0.100 | 16k | 15k | text+image+video->text | frequency_penalty, logprobs, max_tokens, presence_penalty, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,... |
| 167 | MiniMax: MiniMax M2.7minimax/minimax-m2.7 | $0.300 / $1.200 | 205k | 131k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent... |
| 168 | MiniMax: MiniMax M2.7 (free)minimax/minimax-m2.7:free | FREE | 197k | 177k | text->text | include_reasoning, max_tokens, reasoning, response_format, seed, temperature, tool_choice, tools, top_p | MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent... |
| 169 | OpenAI: GPT-5.4 Nanoopenai/gpt-5.4-nano | $0.200 / $1.250 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency... |
| 170 | OpenAI: GPT-5.4 Nano (batch)openai/gpt-5.4-nano:batch | $0.100 / $0.625 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency... |
| 171 | OpenAI: GPT-5.4 Miniopenai/gpt-5.4-mini | $0.750 / $4.500 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,... |
| 172 | OpenAI: GPT-5.4 Mini (batch)openai/gpt-5.4-mini:batch | $0.375 / $2.250 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,... |
| 173 | Mistral: Mistral Small 4mistralai/mistral-small-2603 | $0.150 / $0.600 | 262k | 210k | text+image->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from... |
| 174 | Z.ai: GLM 5 Turboz-ai/glm-5-turbo | $1.200 / $4.000 | 203k | 131k | text->text | include_reasoning, max_tokens, reasoning, response_format, temperature, tool_choice, tools, top_k, top_p | GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows... |
| 175 | NVIDIA: Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | $0.085 / $0.400 | 1.0M | 16k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer... |
| 176 | NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | FREE | 262k | 236k | text->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer... |
| 177 | ByteDance Seed: Seed-2.0-Litebytedance-seed/seed-2.0-lite | $0.250 / $2.000 | 262k | 131k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p | Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across... |
| 178 | Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | $0.100 / $0.150 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design... |
| 179 | Qwen: Qwen3.5-9B (batch)qwen/qwen3.5-9b:batch | $0.170 / $0.250 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design... |
| 180 | OpenAI: GPT-5.4 Proopenai/gpt-5.4-pro | $30.000 / $180.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K... |
| 181 | OpenAI: GPT-5.4 Pro (batch)openai/gpt-5.4-pro:batch | $15.000 / $90.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K... |
| 182 | OpenAI: GPT-5.4openai/gpt-5.4 | $2.500 / $15.000 | 1.1M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for... |
| 183 | OpenAI: GPT-5.4 (batch)openai/gpt-5.4:batch | $1.250 / $7.500 | 1.1M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for... |
| 184 | Inception: Mercury 2inception/mercury-2 | $0.250 / $0.750 | 128k | 50k | text->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools | Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving... |
| 185 | Google: Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | $0.250 / $1.500 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across... |
| 186 | ByteDance Seed: Seed-2.0-Minibytedance-seed/seed-2.0-mini | $0.100 / $0.400 | 262k | 131k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p | Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,... |
| 187 | Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)google/gemini-3.1-flash-image-preview | $0.500 / $3.000 | 66k | 59k | text+image->text+image | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, top_p | Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines... |
| 188 | Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | $0.250 / $1.250 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall... |
| 189 | Qwen: Qwen3.5-27Bqwen/qwen3.5-27b | $0.195 / $1.560 | 262k | 66k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of... |
| 190 | Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | $0.290 / $2.400 | 262k | 82k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of... |
| 191 | Qwen: Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | $0.065 / $0.260 | 1.0M | 66k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the... |
| 192 | Google: Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools | $2.000 / $12.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party... |
| 193 | OpenAI: GPT-5.3-Codexopenai/gpt-5.3-codex | $1.750 / $14.000 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results... |
| 194 | AionLabs: Aion-2.0aion-labs/aion-2.0 | $0.800 / $1.600 | 131k | 33k | text->text | include_reasoning, max_tokens, reasoning, response_format, temperature, tool_choice, tools, top_p | Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and conflict into stories, making narratives feel more engaging.... |
| 195 | Google: Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | $2.000 / $12.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation... |
| 196 | Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | $1.000 / $6.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation... |
| 197 | Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | $3.000 / $15.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p, verbosity | Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with... |
| 198 | Anthropic: Claude Sonnet 4.6 (batch)anthropic/claude-sonnet-4.6:batch | $1.500 / $7.500 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p, verbosity | Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with... |
| 199 | Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | $0.260 / $1.560 | 1.0M | 66k | text+image+video->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of... |
| 200 | Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | $0.550 / $3.500 | 262k | 236k | text+image+video->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers... |
| 201 | MiniMax: MiniMax M2.5minimax/minimax-m2.5 | $0.270 / $1.080 | 205k | 128k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1... |
| 202 | Z.ai: GLM 5z-ai/glm-5 | $0.600 / $1.920 | 205k | 128k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading... |
| 203 | Qwen: Qwen3 Max Thinkingqwen/qwen3-max-thinking | $0.780 / $3.900 | 262k | 66k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it... |
| 204 | Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6 | $5.000 / $25.000 | 1.0M | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p, verbosity | Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective... |
| 205 | Anthropic: Claude Opus 4.6 (batch)anthropic/claude-opus-4.6:batch | $2.500 / $12.500 | 1.0M | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p, verbosity | Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective... |
| 206 | Qwen: Qwen3 Coder Nextqwen/qwen3-coder-next | $0.120 / $0.800 | 262k | 236k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per... |
| 207 | Free Models Routeropenrouter/free | FREE | 200k | N/D | text+image->text | frequency_penalty, include_reasoning, logprobs, max_completion_tokens, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that... |
| 208 | StepFun: Step 3.5 Flashstepfun/step-3.5-flash | $0.100 / $0.300 | 262k | 66k | text->text | frequency_penalty, include_reasoning, max_tokens, reasoning, temperature, tool_choice, tools, top_k, top_p | Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token.... |
| 209 | MoonshotAI: Kimi K2.5moonshotai/kimi-k2.5 | $0.450 / $2.250 | 262k | 236k | text+image->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed... |
| 210 | Upstage: Solar Pro 3upstage/solar-pro-3 | $0.150 / $0.600 | 131k | 118k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, response_format, structured_outputs, temperature, tool_choice, tools, top_p | Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized... |
| 211 | MiniMax: MiniMax M2-herminimax/minimax-m2-her | $0.300 / $1.200 | 66k | 2k | text->text | max_tokens, temperature, top_p | MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations. Designed to stay consistent in tone and personality, it supports rich message... |
| 212 | Writer: Palmyra X5writer/palmyra-x5 | $0.600 / $6.000 | 1.0M | 8k | text->text | max_tokens, stop, temperature, top_k, top_p | Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading speed and efficiency on context windows up to 1 million... |
| 213 | OpenAI: GPT Audioopenai/gpt-audio | $2.500 / $10.000 | 128k | 16k | text+audio->text+audio | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced... |
| 214 | OpenAI: GPT Audio Miniopenai/gpt-audio-mini | $0.600 / $2.400 | 128k | 16k | text+audio->text+audio | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million... |
| 215 | Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash | $0.060 / $0.400 | 203k | 16k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,... |
| 216 | OpenAI: GPT-5.2-Codexopenai/gpt-5.2-codex | $1.750 / $14.000 | 400k | 128k | text+image->text | include_reasoning, max_completion_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... |
| 217 | ByteDance Seed: Seed 1.6 Flashbytedance-seed/seed-1.6-flash | $0.075 / $0.300 | 262k | 33k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p | Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of... |
| 218 | ByteDance Seed: Seed 1.6bytedance-seed/seed-1.6 | $0.250 / $2.000 | 262k | 33k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_p | Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window. |
| 219 | MiniMax: MiniMax M2.1minimax/minimax-m2.1 | $0.300 / $1.200 | 205k | 131k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world... |
| 220 | Z.ai: GLM 4.7z-ai/glm-4.7 | $0.400 / $1.750 | 205k | 131k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_p | GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while... |
| 221 | Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | $0.500 / $3.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool... |
| 222 | Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | $0.250 / $1.500 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool... |
| 223 | NVIDIA: Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | $0.050 / $0.200 | 262k | 236k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully... |
| 224 | OpenAI: GPT-5.2 Chatopenai/gpt-5.2-chat | $1.750 / $14.000 | 128k | 32k | text+image+file->text | max_completion_tokens, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on... |
| 225 | OpenAI: GPT-5.2 Proopenai/gpt-5.2-pro | $21.000 / $168.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,... |
| 226 | OpenAI: GPT-5.2 Pro (batch)openai/gpt-5.2-pro:batch | $10.500 / $84.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,... |
| 227 | OpenAI: GPT-5.2openai/gpt-5.2 | $1.750 / $14.000 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly... |
| 228 | OpenAI: GPT-5.2 (batch)openai/gpt-5.2:batch | $0.875 / $7.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly... |
| 229 | Mistral: Devstral 2 2512mistralai/devstral-2512 | $0.400 / $2.000 | 262k | 210k | text+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring... |
| 230 | Relace: Relace Searchrelace/relace-search | $1.000 / $3.000 | 256k | 128k | text->text | max_tokens, response_format, seed, stop, temperature, tool_choice, tools, top_p | The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic... |
| 231 | Z.ai: GLM 4.6Vz-ai/glm-4.6v | $0.300 / $0.900 | 131k | 33k | text+image+video->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts... |
| 232 | Body Builder (beta)openrouter/bodybuilder | N/D | 128k | N/D | text->text | N/D | Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and Body Builder will construct the appropriate API calls. Example:... |
| 233 | OpenAI: GPT-5.1-Codex-Maxopenai/gpt-5.1-codex-max | $1.250 / $10.000 | 400k | 128k | text+image->text | include_reasoning, max_completion_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic... |
| 234 | Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | $0.300 / $2.500 | 1.0M | 66k | text+image+file+video->text | include_reasoning, max_tokens, reasoning, stop, temperature, tool_choice, tools, top_k, top_p | Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing... |
| 235 | Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | $0.200 / $0.200 | 262k | 210k | text+image->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language... |
| 236 | Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | $0.150 / $0.150 | 262k | 210k | text+image->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities. |
| 237 | Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | $0.100 / $0.100 | 131k | 105k | text+image->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities. |
| 238 | Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | $0.500 / $1.500 | 262k | 210k | text+image+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license. |
| 239 | DeepSeek: DeepSeek V3.2deepseek/deepseek-v3.2 | $0.269 / $0.400 | 164k | 66k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism... |
| 240 | Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | $5.000 / $25.000 | 200k | 64k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, verbosity | Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and... |
| 241 | Anthropic: Claude Opus 4.5 (batch)anthropic/claude-opus-4.5:batch | $2.500 / $12.500 | 200k | 64k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, verbosity | Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and... |
| 242 | Google: Nano Banana Pro (Gemini 3 Pro Image Preview)google/gemini-3-pro-image-preview | $2.000 / $12.000 | 66k | 33k | text+image->text+image | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, top_p | Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and... |
| 243 | OpenAI: GPT-5.1openai/gpt-5.1 | $1.250 / $10.000 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning... |
| 244 | OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | $0.625 / $5.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning... |
| 245 | OpenAI: GPT-5.1-Codexopenai/gpt-5.1-codex | $1.250 / $10.000 | 400k | 128k | text+image->text | include_reasoning, max_completion_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... |
| 246 | OpenAI: GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | $0.250 / $2.000 | 400k | 128k | text+image->text | include_reasoning, max_completion_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex |
| 247 | MoonshotAI: Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | $0.600 / $2.500 | 262k | 100k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillion-parameter Mixture-of-Experts (MoE) architecture introduced in... |
| 248 | Amazon: Nova Premier 1.0amazon/nova-premier-v1 | $2.500 / $12.500 | 1.0M | 32k | text+image->text | max_tokens, stop, temperature, tools, top_k, top_p | Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models. |
| 249 | Perplexity: Sonar Pro Searchperplexity/sonar-pro-search | $3.000 / $15.000 | 200k | 8k | text+image->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, structured_outputs, temperature, top_k, top_p, web_search_options | Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based... |
| 250 | Mistral: Voxtral Small 24B 2507mistralai/voxtral-small-24b-2507 | $0.100 / $0.300 | 33k | 26k | text+file+audio->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio... |
| 251 | OpenAI: gpt-oss-safeguard-20bopenai/gpt-oss-safeguard-20b | $0.075 / $0.300 | 131k | 66k | text->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model offers lower latency for safety tasks like content classification, LLM filtering, and trust... |
| 252 | MiniMax: MiniMax M2minimax/minimax-m2 | $0.255 / $1.020 | 205k | 131k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,... |
| 253 | Qwen: Qwen3 VL 32B Instructqwen/qwen3-vl-32b-instruct | $0.104 / $0.416 | 131k | 33k | text+image->text | frequency_penalty, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text... |
| 254 | IBM: Granite 4.0 Microibm-granite/granite-4.0-h-micro | $0.017 / $0.112 | 131k | 118k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, temperature, top_k, top_logprobs, top_p | Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long... |
| 255 | OpenAI: GPT-5 Image Miniopenai/gpt-5-image-mini | $2.500 / $2.000 | 400k | 128k | text+image+file->text+image | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, top_logprobs, top_p | GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text... |
| 256 | Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | $1.000 / $5.000 | 200k | 64k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance... |
| 257 | Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | $0.500 / $2.500 | 200k | 64k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance... |
| 258 | Qwen: Qwen3 VL 8B Thinkingqwen/qwen3-vl-8b-thinking | $0.180 / $2.100 | 131k | 33k | text+image->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and... |
| 259 | Qwen: Qwen3 VL 8B Instructqwen/qwen3-vl-8b-instruct | $0.117 / $0.455 | 262k | 33k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon... |
| 260 | OpenAI: GPT-5 Imageopenai/gpt-5-image | $10.000 / $10.000 | 400k | 128k | text+image+file->text+image | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, top_logprobs, top_p | [GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,... |
| 261 | Google: Nano Banana (Gemini 2.5 Flash Image)google/gemini-2.5-flash-image | $0.300 / $2.500 | 33k | 8k | text+image->text+image | max_tokens, response_format, seed, stop, structured_outputs, temperature, top_p | Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,... |
| 262 | Qwen: Qwen3 VL 30B A3B Thinkingqwen/qwen3-vl-30b-a3b-thinking | $0.200 / $2.400 | 262k | 33k | text+image->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels... |
| 263 | Qwen: Qwen3 VL 30B A3B Instructqwen/qwen3-vl-30b-a3b-instruct | $0.150 / $0.600 | 262k | 16k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception... |
| 264 | OpenAI: GPT-5 Proopenai/gpt-5-pro | $15.000 / $120.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and... |
| 265 | OpenAI: GPT-5 Pro (batch)openai/gpt-5-pro:batch | $7.500 / $60.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and... |
| 266 | Z.ai: GLM 4.6z-ai/glm-4.6 | $0.550 / $2.200 | 205k | 131k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex... |
| 267 | Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | $3.000 / $15.000 | 1.0M | 64k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with... |
| 268 | Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | $1.500 / $7.500 | 1.0M | 64k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with... |
| 269 | DeepSeek: DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | $0.270 / $0.410 | 164k | 66k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism... |
| 270 | TheDrummer: Cydonia 24B V4.1thedrummer/cydonia-24b-v4.1 | $0.300 / $0.500 | 131k | 118k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_logprobs, top_p | Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence. |
| 271 | Relace: Relace Apply 3relace/relace-apply-3 | $0.850 / $1.250 | 256k | 128k | text->text | max_tokens, seed, stop | Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o, Claude, and others into your files at... |
| 272 | Qwen: Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | $0.400 / $4.000 | 131k | 33k | text+image->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math.... |
| 273 | Qwen: Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | $0.210 / $1.900 | 262k | 33k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table... |
| 274 | Qwen: Qwen3 Maxqwen/qwen3-max | $0.780 / $3.900 | 262k | 66k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It... |
| 275 | Qwen: Qwen3 Coder Plusqwen/qwen3-coder-plus | $0.650 / $3.250 | 1.0M | 66k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and... |
| 276 | DeepSeek: DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | $0.270 / $1.000 | 164k | 33k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's... |
| 277 | Qwen: Qwen3 Coder Flashqwen/qwen3-coder-flash | $0.195 / $0.975 | 1.0M | 66k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling... |
| 278 | Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | $0.150 / $1.200 | 262k | 33k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic... |
| 279 | Qwen: Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | $0.100 / $1.100 | 262k | 236k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual... |
| 280 | Qwen: Qwen Plus 0728qwen/qwen-plus-2025-07-28 | $0.260 / $0.780 | 1.0M | 33k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination. |
| 281 | MoonshotAI: Kimi K2 0905moonshotai/kimi-k2-0905 | $0.600 / $2.500 | 262k | 100k | text->text | frequency_penalty, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32... |
| 282 | Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | $0.200 / $2.400 | 82k | 33k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated... |
| 283 | Nous: Hermes 4 70Bnousresearch/hermes-4-70b | $0.130 / $0.400 | 131k | 118k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, temperature, top_k, top_p | Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either... |
| 284 | Nous: Hermes 4 405Bnousresearch/hermes-4-405b | $1.000 / $3.000 | 131k | 118k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, temperature, top_k, top_p | Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with... |
| 285 | DeepSeek: DeepSeek V3.1deepseek/deepseek-chat-v3.1 | $0.250 / $0.950 | 164k | 33k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context... |
| 286 | Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | $0.400 / $2.000 | 131k | 105k | text+image+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances... |
| 287 | Z.ai: GLM 4.5Vz-ai/glm-4.5v | $0.600 / $1.800 | 66k | 16k | text+image->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,... |
| 288 | OpenAI: GPT-5openai/gpt-5 | $1.250 / $10.000 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy... |
| 289 | OpenAI: GPT-5 (batch)openai/gpt-5:batch | $0.625 / $5.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy... |
| 290 | OpenAI: GPT-5 Miniopenai/gpt-5-mini | $0.250 / $2.000 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost.... |
| 291 | OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | $0.125 / $1.000 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost.... |
| 292 | OpenAI: GPT-5 Nanoopenai/gpt-5-nano | $0.050 / $0.400 | 400k | 128k | text+image+file->text | include_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger... |
| 293 | OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch | $0.025 / $0.200 | 400k | 128k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger... |
| 294 | OpenAI: gpt-oss-120bopenai/gpt-oss-120b | $0.037 / $0.170 | 131k | 118k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, top_p | gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized... |
| 295 | OpenAI: gpt-oss-120b (batch)openai/gpt-oss-120b:batch | $0.150 / $0.600 | 131k | 118k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized... |
| 296 | OpenAI: gpt-oss-20bopenai/gpt-oss-20b | $0.030 / $0.130 | 131k | 118k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for... |
| 297 | OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch | $0.050 / $0.200 | 131k | 118k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, stop, structured_outputs, temperature, top_k, top_p | gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for... |
| 298 | Anthropic: Claude Opus 4.1anthropic/claude-opus-4.1 | $15.000 / $75.000 | 200k | 32k | text+image+file->text | include_reasoning, max_tokens, reasoning, stop, temperature, tool_choice, tools, top_k, top_p | Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains... |
| 299 | Anthropic: Claude Opus 4.1 (batch)anthropic/claude-opus-4.1:batch | $7.500 / $37.500 | 200k | 32k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, stop, structured_outputs, temperature, tool_choice, tools | Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains... |
| 300 | Mistral: Codestral 2508mistralai/codestral-2508 | $0.300 / $0.900 | 256k | 205k | text+file->text | frequency_penalty, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08) |
| 301 | Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | $0.070 / $0.280 | 262k | 236k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the... |
| 302 | Qwen: Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | $0.048 / $0.193 | 262k | 32k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and... |
| 303 | Z.ai: GLM 4.5z-ai/glm-4.5 | $0.600 / $2.200 | 131k | 98k | text->text | include_reasoning, max_tokens, reasoning, response_format, temperature, tool_choice, tools, top_k, top_p | GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly... |
| 304 | Z.ai: GLM 4.5 Airz-ai/glm-4.5-air | $0.130 / $0.850 | 131k | 98k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, seed, stop, temperature, tool_choice, tools, top_k, top_p | GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter... |
| 305 | Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | $0.230 / $2.300 | 131k | 118k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144... |
| 306 | Qwen: Qwen3 Coder 480B A35Bqwen/qwen3-coder | $0.300 / $1.000 | 262k | 66k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over... |
| 307 | ByteDance: UI-TARS 7B bytedance/ui-tars-1.5-7b | $0.100 / $0.200 | 128k | 2k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, seed, stop, structured_outputs, temperature, top_k, top_logprobs, top_p | UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement... |
| 308 | Google: Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-lite | $0.100 / $0.400 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance... |
| 309 | Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | $0.050 / $0.200 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance... |
| 310 | Qwen: Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | $0.087 / $0.350 | 262k | 236k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,... |
| 311 | MoonshotAI: Kimi K2 0711moonshotai/kimi-k2 | $0.570 / $2.300 | 131k | 100k | text->text | frequency_penalty, max_tokens, presence_penalty, repetition_penalty, seed, stop, temperature, tool_choice, tools, top_k, top_p | Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for... |
| 312 | Venice: Uncensoredcognitivecomputations/dolphin-mistral-24b-venice-edition | $0.200 / $0.900 | 128k | 8k | text->text | frequency_penalty, max_tokens, presence_penalty, response_format, stop, temperature, top_k, top_p | Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving... |
| 313 | Tencent: Hunyuan A13B Instructtencent/hunyuan-a13b-instruct | $0.140 / $0.570 | 131k | 118k | text->text | frequency_penalty, include_reasoning, max_tokens, reasoning, response_format, structured_outputs, temperature, top_k, top_p | Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark... |
| 314 | Morph: Morph V3 Largemorph/morph-v3-large | $0.900 / $1.900 | 262k | 131k | text->text | logprobs, max_tokens, response_format, stop, structured_outputs, temperature, top_logprobs | Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>... |
| 315 | Morph: Morph V3 Fastmorph/morph-v3-fast | $0.800 / $1.200 | 82k | 38k | text->text | max_tokens, stop, temperature | Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>... |
| 316 | Baidu: ERNIE 4.5 VL 424B A47B baidu/ernie-4.5-vl-424b-a47b | $0.420 / $1.250 | 123k | 16k | text+image->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, seed, stop, temperature, top_k, top_p | ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data... |
| 317 | Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | $0.075 / $0.200 | 131k | 16k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on... |
| 318 | MiniMax: MiniMax M1minimax/minimax-m1 | $0.550 / $2.200 | 1.0M | 40k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, seed, stop, temperature, tool_choice, tools, top_k, top_p | MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (MoE) architecture paired with a custom "lightning attention" mechanism, allowing it... |
| 319 | Google: Gemini 2.5 Flashgoogle/gemini-2.5-flash | $0.300 / $2.500 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater... |
| 320 | Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | $0.150 / $1.250 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater... |
| 321 | Google: Gemini 2.5 Progoogle/gemini-2.5-pro | $1.250 / $10.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy... |
| 322 | Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | $0.625 / $5.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy... |
| 323 | OpenAI: o3 Proopenai/o3-pro | $20.000 / $80.000 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently... |
| 324 | Google: Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | $1.250 / $10.000 | 1.0M | 66k | text+image+file+audio->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy... |
| 325 | DeepSeek: R1 0528deepseek/deepseek-r1-0528 | $0.500 / $2.150 | 164k | 33k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active... |
| 326 | Anthropic: Claude Opus 4anthropic/claude-opus-4 | $15.000 / $75.000 | 200k | 32k | text+image+file->text | include_reasoning, max_tokens, reasoning, stop, temperature, tool_choice, tools, top_p | Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in... |
| 327 | Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | $3.000 / $15.000 | 1.0M | 64k | text+image+file->text | include_reasoning, max_tokens, reasoning, stop, temperature, tool_choice, tools, top_k, top_p | Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),... |
| 328 | Mistral: Mistral Medium 3mistralai/mistral-medium-3 | $0.400 / $2.000 | 131k | 105k | text+image+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost... |
| 329 | Google: Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | $1.250 / $10.000 | 1.0M | 66k | text+image+file+audio+video->text | include_reasoning, max_tokens, reasoning, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy... |
| 330 | Meta: Llama Guard 4 12Bmeta-llama/llama-guard-4-12b | $0.180 / $0.180 | 164k | 16k | text+image->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, temperature, top_k, top_p | Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM... |
| 331 | Qwen: Qwen3 30B A3Bqwen/qwen3-30b-a3b | $0.120 / $0.500 | 131k | 16k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique... |
| 332 | Qwen: Qwen3 8Bqwen/qwen3-8b | $0.117 / $0.455 | 131k | 8k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,... |
| 333 | Qwen: Qwen3 14Bqwen/qwen3-14b | $0.228 / $0.910 | 131k | 8k | text->text | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for... |
| 334 | Qwen: Qwen3 32Bqwen/qwen3-32b | $0.080 / $0.280 | 131k | 16k | text->text | frequency_penalty, include_reasoning, logit_bias, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for... |
| 335 | Qwen: Qwen3 235B A22Bqwen/qwen3-235b-a22b | $0.455 / $1.820 | 131k | 8k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, response_format, seed, stop, temperature, tool_choice, tools, top_k, top_p | Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and... |
| 336 | OpenAI: o4 Mini Highopenai/o4-mini-high | $1.100 / $4.400 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining... |
| 337 | OpenAI: o3openai/o3 | $2.000 / $8.000 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following.... |
| 338 | OpenAI: o3 (batch)openai/o3:batch | $1.000 / $4.000 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following.... |
| 339 | OpenAI: o4 Miniopenai/o4-mini | $1.100 / $4.400 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning... |
| 340 | OpenAI: o4 Mini (batch)openai/o4-mini:batch | $0.550 / $2.200 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning... |
| 341 | OpenAI: GPT-4.1openai/gpt-4.1 | $2.000 / $8.000 | 1.0M | 33k | text+image+file->text | max_completion_tokens, max_tokens, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and... |
| 342 | OpenAI: GPT-4.1 (batch)openai/gpt-4.1:batch | $1.000 / $4.000 | 1.0M | 33k | text+image+file->text | max_tokens, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and... |
| 343 | OpenAI: GPT-4.1 Miniopenai/gpt-4.1-mini | $0.400 / $1.600 | 1.0M | 33k | text+image+file->text | max_completion_tokens, max_tokens, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard... |
| 344 | OpenAI: GPT-4.1 Mini (batch)openai/gpt-4.1-mini:batch | $0.200 / $0.800 | 1.0M | 33k | text+image+file->text | max_tokens, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard... |
| 345 | OpenAI: GPT-4.1 Nanoopenai/gpt-4.1-nano | $0.100 / $0.400 | 1.0M | 33k | text+image+file->text | max_completion_tokens, max_tokens, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million... |
| 346 | OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch | $0.050 / $0.200 | 1.0M | 33k | text+image+file->text | max_tokens, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p | For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million... |
| 347 | Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | $0.200 / $0.696 | 1.0M | 115k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward... |
| 348 | Meta: Llama 4 Scoutmeta-llama/llama-4-scout | $0.100 / $0.300 | 1.3M | 16k | text+image->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input... |
| 349 | DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | $0.250 / $1.000 | 164k | 147k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well... |
| 350 | OpenAI: o1-proopenai/o1-pro | $150.000 / $600.000 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs | The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide... |
| 351 | Mistral: Mistral Small 3.1 24Bmistralai/mistral-small-3.1-24b-instruct | $0.351 / $0.555 | 128k | 102k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, seed, stop, temperature, top_k, top_logprobs, top_p | Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and... |
| 352 | Google: Gemma 3 4Bgoogle/gemma-3-4b-it | $0.050 / $0.100 | 131k | 16k | text+image->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,... |
| 353 | Google: Gemma 3 12Bgoogle/gemma-3-12b-it | $0.050 / $0.150 | 131k | 16k | text+image->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,... |
| 354 | Cohere: Command Acohere/command-a | $2.500 / $10.000 | 256k | 8k | text->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary... |
| 355 | Reka Flash 3rekaai/reka-flash-3 | $0.100 / $0.200 | 66k | 59k | text->text | frequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, seed, stop, structured_outputs, temperature, top_k, top_logprobs, top_p | Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a... |
| 356 | Google: Gemma 3 27Bgoogle/gemma-3-27b-it | $0.080 / $0.450 | 131k | 118k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,... |
| 357 | TheDrummer: Skyfall 36B V2thedrummer/skyfall-36b-v2 | $0.550 / $0.800 | 33k | 29k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_logprobs, top_p | Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent storytelling. |
| 358 | Perplexity: Sonar Reasoning Properplexity/sonar-reasoning-pro | $2.000 / $8.000 | 128k | 115k | text+image->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, temperature, top_k, top_p, web_search_options | Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-reasoning-pro-and-sonar-pro) Sonar Reasoning Pro is a premier reasoning model powered by DeepSeek R1 with Chain of Thought (CoT). Designed for... |
| 359 | Perplexity: Sonar Properplexity/sonar-pro | $3.000 / $15.000 | 200k | 8k | text+image->text | frequency_penalty, max_tokens, presence_penalty, temperature, top_k, top_p, web_search_options | Note: Sonar Pro pricing includes Perplexity search pricing. See [details here](https://docs.perplexity.ai/guides/pricing#detailed-pricing-breakdown-for-sonar-reasoning-pro-and-sonar-pro) For enterprises seeking more advanced capabilities, the Sonar Pro API can handle in-depth, multi-step queries with added extensibility, like... |
| 360 | Perplexity: Sonar Deep Researchperplexity/sonar-deep-research | $2.000 / $8.000 | 128k | 115k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, temperature, top_k, top_p, web_search_options | Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autonomously searches, reads, and evaluates sources, refining its approach as it gathers... |
| 361 | Mistral: Sabamistralai/mistral-saba | $0.200 / $0.600 | 33k | 26k | text+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant responses while maintaining efficient performance. Trained on curated regional... |
| 362 | OpenAI: o3 Mini Highopenai/o3-mini-high | $1.100 / $4.400 | 200k | 100k | text+file->text | include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools | OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and... |
| 363 | AionLabs: Aion-RP 1.0 (8B)aion-labs/aion-rp-llama-3.1-8b | $0.800 / $1.600 | 33k | 29k | text->text | max_tokens, temperature, top_p | Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses. It is a fine-tuned base model... |
| 364 | Qwen: Qwen2.5 VL 72B Instructqwen/qwen2.5-vl-72b-instruct | $0.250 / $0.750 | 128k | 29k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_logprobs, top_p | Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images. |
| 365 | Qwen: Qwen-Plusqwen/qwen-plus | $0.260 / $0.780 | 1.0M | 33k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination. |
| 366 | OpenAI: o3 Miniopenai/o3-mini | $1.100 / $4.400 | 200k | 100k | text+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to... |
| 367 | OpenAI: o3 Mini (batch)openai/o3-mini:batch | $0.550 / $2.200 | 200k | 100k | text+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to... |
| 368 | Mistral: Mistral Small 3mistralai/mistral-small-24b-instruct-2501 | $0.050 / $0.080 | 33k | 16k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed... |
| 369 | Perplexity: Sonarperplexity/sonar | $1.000 / $1.000 | 127k | 114k | text+image->text | frequency_penalty, max_tokens, presence_penalty, temperature, top_k, top_p, web_search_options | Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for companies seeking to integrate lightweight question-and-answer features... |
| 370 | DeepSeek: R1 Distill Llama 70Bdeepseek/deepseek-r1-distill-llama-70b | $0.800 / $0.800 | 8k | 7k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, seed, stop, temperature, top_k, top_p | DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across... |
| 371 | DeepSeek: R1deepseek/deepseek-r1 | $0.700 / $2.500 | 64k | 16k | text->text | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.... |
| 372 | MiniMax: MiniMax-01minimax/minimax-01 | $0.200 / $1.100 | 1.0M | 900k | text+image->text | max_tokens, temperature, top_p | MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context... |
| 373 | Microsoft: Phi 4microsoft/phi-4 | $0.070 / $0.140 | 16k | 15k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | [Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion... |
| 374 | DeepSeek: DeepSeek V3deepseek/deepseek-chat | $0.257 / $1.029 | 164k | 16k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations... |
| 375 | Sao10K: Llama 3.3 Euryale 70Bsao10k/l3.3-euryale-70b | $0.650 / $0.750 | 131k | 16k | text->text | frequency_penalty, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_logprobs, top_p | Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b). |
| 376 | OpenAI: o1openai/o1 | $15.000 / $60.000 | 200k | 100k | text+image+file->text | include_reasoning, max_tokens, reasoning, response_format, seed, structured_outputs, tool_choice, tools | The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason... |
| 377 | Cohere: Command R7B (12-2024)cohere/command-r7b-12-2024 | $0.037 / $0.150 | 128k | 4k | text->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning... |
| 378 | Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | $0.100 / $0.320 | 131k | 16k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model... |
| 379 | Amazon: Nova Lite 1.0amazon/nova-lite-v1 | $0.060 / $0.240 | 300k | 5k | text+image->text | max_tokens, stop, temperature, tools, top_k, top_p | Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite... |
| 380 | Amazon: Nova Micro 1.0amazon/nova-micro-v1 | $0.035 / $0.140 | 128k | 5k | text->text | max_tokens, stop, temperature, tools, top_k, top_p | Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length... |
| 381 | Amazon: Nova Pro 1.0amazon/nova-pro-v1 | $0.800 / $3.200 | 300k | 5k | text+image->text | max_tokens, stop, temperature, tools, top_k, top_p | Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December... |
| 382 | OpenAI: GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | $2.500 / $10.000 | 128k | 16k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded... |
| 383 | Mistral Large 2407mistralai/mistral-large-2407 | $2.000 / $6.000 | 131k | 105k | text+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/).... |
| 384 | Qwen2.5 Coder 32B Instructqwen/qwen-2.5-coder-32b-instruct | $0.660 / $1.000 | 33k | 29k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, seed, stop, temperature, top_k, top_p | Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reasoning**... |
| 385 | TheDrummer: UnslopNemo 12Bthedrummer/unslopnemo-12b | $0.400 / $0.400 | 1.0M | 26k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios. |
| 386 | Magnum v4 72Banthracite-org/magnum-v4-72b | $2.500 / $5.000 | 33k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_a, top_k, top_logprobs, top_p | This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet) and Opus(https://openrouter.ai/anthropic/claude-3-opus). The model is fine-tuned on top of [Qwen2.5 72B](https://openrouter.ai/qwen/qwen-2.5-72b-instruct). |
| 387 | Qwen: Qwen2.5 7B Instructqwen/qwen-2.5-7b-instruct | $0.100 / $0.200 | 33k | 29k | text->text | frequency_penalty, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and... |
| 388 | Meta: Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instruct | $0.027 / $0.201 | 60k | 54k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, seed, stop, temperature, top_k, top_p | Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate... |
| 389 | Meta: Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instruct | $0.050 / $0.330 | 131k | 118k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, seed, stop, structured_outputs, temperature, top_k, top_logprobs, top_p | Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it... |
| 390 | Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | $0.360 / $0.400 | 33k | 16k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and... |
| 391 | Cohere: Command R (08-2024)cohere/command-r-08-2024 | $0.150 / $0.600 | 128k | 4k | text->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation (RAG) and tool use. More broadly, it is better at math, code and reasoning and... |
| 392 | Cohere: Command R+ (08-2024)cohere/command-r-plus-08-2024 | $2.500 / $10.000 | 128k | 4k | text->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as compared to the previous Command R+ version, while keeping the hardware footprint... |
| 393 | Sao10K: Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70b | $0.850 / $0.850 | 131k | 16k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_p | Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b). |
| 394 | Nous: Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70b | $0.700 / $0.700 | 131k | 16k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the... |
| 395 | Nous: Hermes 3 405B Instructnousresearch/hermes-3-llama-3.1-405b | $1.000 / $1.000 | 131k | 16k | text->text | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the... |
| 396 | Sao10K: Llama 3 8B Lunarissao10k/l3-lunaris-8b | $0.040 / $0.050 | 8k | 7k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_logprobs, top_p | Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3. It's a strategic merge of multiple models, designed to balance creativity with improved logic and general knowledge.... |
| 397 | OpenAI: GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | $2.500 / $10.000 | 128k | 16k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_completion_tokens, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/). GPT-4o ("o" for "omni") is... |
| 398 | Meta: Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | $0.400 / $0.400 | 131k | 16k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong... |
| 399 | Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | $0.050 / $0.080 | 131k | 118k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to... |
| 400 | Mistral: Mistral Nemomistralai/mistral-nemo | $0.019 / $0.030 | 131k | 16k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,... |
| 401 | OpenAI: GPT-4o-miniopenai/gpt-4o-mini | $0.150 / $0.600 | 128k | 16k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_completion_tokens, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable... |
| 402 | OpenAI: GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | $0.150 / $0.600 | 128k | 16k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable... |
| 403 | OpenAI: GPT-4o-mini (batch)openai/gpt-4o-mini:batch | $0.075 / $0.300 | 128k | 16k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable... |
| 404 | Google: Gemma 2 27Bgoogle/gemma-2-27b-it | $0.650 / $0.650 | 8k | 2k | text->text | frequency_penalty, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_p | Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-suited for a variety of... |
| 405 | OpenAI: GPT-4oopenai/gpt-4o | $2.500 / $10.000 | 128k | 16k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_completion_tokens, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as... |
| 406 | OpenAI: GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | $5.000 / $15.000 | 128k | 4k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_completion_tokens, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as... |
| 407 | OpenAI: GPT-4o (batch)openai/gpt-4o:batch | $1.250 / $5.000 | 128k | 16k | text+image+file->text | frequency_penalty, logit_bias, logprobs, max_tokens, prediction, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p, web_search_options | GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as... |
| 408 | Mistral: Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | $2.000 / $6.000 | 66k | 52k | text+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offering unparalleled cost efficiency for its size. Its strengths include: - strong math, coding,... |
| 409 | WizardLM-2 8x22Bmicrosoft/wizardlm-2-8x22b | $0.620 / $0.620 | 66k | 8k | text->text | frequency_penalty, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, temperature, top_k, top_p | WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is... |
| 410 | OpenAI: GPT-4 Turboopenai/gpt-4-turbo | $10.000 / $30.000 | 128k | 4k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023. |
| 411 | OpenAI: GPT-4 Turbo (batch)openai/gpt-4-turbo:batch | $5.000 / $15.000 | 128k | 4k | text+image->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023. |
| 412 | Anthropic: Claude 3 Haikuanthropic/claude-3-haiku | $0.250 / $1.250 | 200k | 4k | text+image->text | max_tokens, stop, temperature, tool_choice, tools, top_k, top_p | Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal |
| 413 | Mistral Largemistralai/mistral-large | $2.000 / $6.000 | 128k | 102k | text+file->text | frequency_penalty, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_p | This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/).... |
| 414 | OpenAI: GPT-3.5 Turbo (older v0613)openai/gpt-3.5-turbo-0613 | $1.000 / $2.000 | 4k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_completion_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021. |
| 415 | OpenAI: GPT-4 Turbo Previewopenai/gpt-4-turbo-preview | $10.000 / $30.000 | 128k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training data: up to Dec 2023. **Note:** heavily rate limited by OpenAI while... |
| 416 | Auto Routeropenrouter/auto | N/D | 2.0M | N/D | text+image+file+audio+video->text+image | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, prediction, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, top_p, web_search_options | The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community collectively spends on... |
| 417 | OpenAI: GPT-3.5 Turbo Instructopenai/gpt-3.5-turbo-instruct | $1.500 / $2.000 | 4k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, top_logprobs, top_p | This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021. |
| 418 | OpenAI: GPT-3.5 Turbo 16kopenai/gpt-3.5-turbo-16k | $3.000 / $4.000 | 16k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_completion_tokens, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up... |
| 419 | Mancer: Weaver (alpha)mancer/weaver | $0.400 / $0.750 | 8k | 6k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, temperature, top_a, top_k, top_logprobs, top_p | An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations. |
| 420 | ReMM SLERP 13Bundi95/remm-slerp-l2-13b | $0.450 / $0.650 | 6k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_a, top_k, top_logprobs, top_p | A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge |
| 421 | MythoMax 13Bgryphe/mythomax-l2-13b | $0.060 / $0.060 | 8k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_a, top_k, top_logprobs, top_p | One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge |
| 422 | OpenAI: GPT-3.5 Turboopenai/gpt-3.5-turbo | $0.500 / $1.500 | 16k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021. |
| 423 | OpenAI: GPT-3.5 Turbo (batch)openai/gpt-3.5-turbo:batch | $0.250 / $0.750 | 16k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021. |
| 424 | OpenAI: GPT-4openai/gpt-4 | $30.000 / $60.000 | 8k | 4k | text->text | frequency_penalty, logit_bias, logprobs, max_completion_tokens, max_tokens, presence_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_p | OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning... |