Records 401–800 — AI model deprecation and retirement dates by provider
Records 401 to 800 of 943. Each links to its own page, which carries the source URL and the quoted line that states the value.
- gpt-realtime-2.1| gpt-realtime-2.1 | 2026-07-07 | Preview | 2027-06-25 | — |
- gpt-realtime-2.1-mini| gpt-realtime-2.1-mini | 2026-07-07 | Preview | 2027-06-25 | — |
- gpt-realtime-mini| Jan 20, 2027 | `gpt-realtime-mini` | `gpt-realtime-2.1-mini` |
- gpt-realtime-mini-2025-10-06| July 23, 2026 | `gpt-realtime-mini-2025-10-06` | `gpt-realtime-2.1-mini` |
- grok-3| grok-3 | 1 | Retired | 2026-05-01 | grok-4 |
- grok-3-mini| grok-3-mini | 1 | Retired | 2026-05-01 | grok-4-1-fast-reasoning |
- grok-4-1-fast-non-reasoningAs we continue advancing Grok, we are retiring several earlier models to focus fully on our newest generation. Effective May 15, 2026 at 12:00 PM PT, the following models will be r
- grok-4-1-fast-reasoningAs we continue advancing Grok, we are retiring several earlier models to focus fully on our newest generation. Effective May 15, 2026 at 12:00 PM PT, the following models will be r
- grok-4-20-non-reasoning| grok-4-20-non-reasoning | 1 | Preview | 2027-04-06 | — |
- grok-4-20-reasoning| grok-4-20-reasoning | 1 | Preview | 2027-04-06 | — |
- grok-4-0709As we continue advancing Grok, we are retiring several earlier models to focus fully on our newest generation. Effective May 15, 2026 at 12:00 PM PT, the following models will be r
- grok-4-fast-non-reasoning| grok-4-fast-non-reasoning | 1 | Retired | 2026-05-01 | grok-4-1-fast-non-reasoning |
- grok-4-fast-reasoning| grok-4-fast-reasoning | 1 | Retired | 2026-05-01 | grok-4-1-fast-reasoning |
- grok-code-fast-1As we continue advancing Grok, we are retiring several earlier models to focus fully on our newest generation. Effective May 15, 2026 at 12:00 PM PT, the following models will be r
- grok-imagine-image-proAs we continue advancing Grok, we are retiring several earlier models to focus fully on our newest generation. Effective May 15, 2026 at 12:00 PM PT, the following models will be r
- grok-imagine-image-qualitygrok-imagine-image-2.0 now covers everything grok-imagine-image-quality was used for, at a lower price. Effective November 2, 2026, the grok-imagine-image-quality model slug is ret
- gryphe-mythomax-l2-13b| 2025-06-13 | `gryphe-mythomax-l2-13b` | No |
- gryphe-mythomax-l2-13b-lite| 2025-06-13 | `gryphe-mythomax-l2-13b-lite` | No |
- gummy-realtime-v1Audio Series Mainline Models:qwen-tts、qwen-tts-realtime、qwen-voice-enrollment、qwen-voice-design、gummy-realtime-v1
- hazyresearch/M2-BERT-2k-Retrieval-Encoder-V1| 2024-08-22 | `hazyresearch/M2-BERT-2k-Retrieval-Encoder-V1` | No |
- ibm/granite-3-2-8b-instructgranite-3-2-8b-instruct ibm/granite-3-2-8b-instruct | 20 February 2025 | 24 November 2025 | 22 February 2026 | granite-4-h-small |
- ibm/granite-3-3-8b-instructgranite-3-3-8b-instruct ibm/granite-3-3-8b-instruct | 16 April 2025 | 24 November 2025 | 22 February 2026 | granite-4-h-small |
- ibm/granite-3-8b-instructgranite-3-8b-instruct ibm/granite-3-8b-instruct | 21 October 2024 | 24 November 2025 | 22 February 2026 | granite-4-h-small |
- ibm/granite-8b-code-instructgranite-8b-code-instruct (Dallas, Sydney, AWS Mumbai data centers) ibm/granite-8b-code-instruct | 13 February 2025 | 8 May 2026 | 8 August 2026 | granite-8b-code-instruct on Deploy
- ibm/granite-13b-instruct-v2granite-13b-instruct-v2 ibm/granite-13b-instruct-v2 | 1 December 2023 | 19 June 2025 | 15 October 2025 |
- ibm/granite-embedding-107m-multilingualgranite-embedding-107m-multilingual ibm/granite-embedding-107m-multilingual | 6 January 2025 | 13 August 2025 | 12 November 2025 | granite-278m-multilingual-embedding |
- ibm/granite-guardian-3-8bgranite-guardian-3-8b (Dallas, Sydney, AWS Mumbai, AWS GovCloud data centers) ibm/granite-guardian-3-8b | 21 October 2024 | 8 May 2026 | 8 August 2026 | |
- ibm/slate-30m-english-rtrvrslate-30m-english-rtrvr ibm/slate-30m-english-rtrvr | 9 August 2024 | 13 August 2025 | 12 November 2025 | slate-30m-english-rtrvr-v2 Deprecated |
- ibm/slate-30m-english-rtrvr-v2slate-30m-english-rtrvr-v2 (Frankfurt, London, Tokyo, Sydney, Toronto, AWS Mumbai, AWS GovCloud data centers) ibm/slate-30m-english-rtrvr-v2 | 15 August 2024 | 8 May 2026 | 12 Janu
- ibm/slate-125m-english-rtrvrslate-125m-english-rtrvr ibm/slate-125m-english-rtrvr | 11 April 2024 | 13 August 2025 | 12 November 2025 | slate-125m-english-rtrvr-v2 Deprecated |
- ibm/slate-125m-english-rtrvr-v2slate-125m-english-rtrvr-v2 (Dallas, Frankfurt, London, Tokyo, Sydney, Toronto, AWS Mumbai, AWS GovCloud data centers) ibm/slate-125m-english-rtrvr-v2 | 15 August 2024 | 8 May 2026
- imagen-3.0-generate-002imagen-3.0-generate-002 | February 6, 2025 | November 10, 2025 | imagen-4.0-generate-001 |
- imagen-4.0-fast-generate-001imagen-4.0-fast-generate-001 | June 24, 2025 | August 17, 2026 | gemini-3.1-flash-image |
- imagen-4.0-generate-001Imagen models Model | Release date | Shutdown date | Recommended replacement | imagen-4.0-generate-001 | June 24, 2025 | August 17, 2026 | gemini-3.1-flash-image |
- imagen-4.0-generate-preview-06-06imagen-4.0-generate-preview-06-06 | June 24, 2025 | February 17, 2026 | imagen-4.0-generate-001 |
- imagen-4.0-ultra-generate-001imagen-4.0-ultra-generate-001 | June 24, 2025 | August 17, 2026 | gemini-3.1-flash-image |
- imagen-4.0-ultra-generate-preview-06-06imagen-4.0-ultra-generate-preview-06-06 | June 24, 2025 | February 17, 2026 | imagen-4.0-ultra-generate-001 |
- imagetextimagetext | June 7, 2023 | September 24, 2025 | gemini-2.5-flash-image |
- intfloat/multilingual-e5-large-instruct| 2026-09-14 | `intfloat/multilingual-e5-large-instruct` | | No |
- jamba-1.5-largeJamba 1.5 Large is deprecated as of August 27, 2025 and will be shut down on February 27, 2026. Jamba 1.5 Large is available to existing customers only.
- jamba-1.5-miniJamba 1.5 Mini is deprecated as of August 27, 2025 and will be shut down on February 27, 2026. Jamba 1.5 Mini is available to existing customers only.
- Kimi-K2| `Kimi-K2` | `Kimi-K2-0905` | Same architecture, improved post-training |
- kimi-k2-0711-previewThe `kimi-k2` series models were officially discontinued on **May 25, 2026** and are no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-
- kimi-k2-0905-previewThe `kimi-k2` series models were officially discontinued on **May 25, 2026** and are no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-
- kimi-k2-thinkingThe `kimi-k2` series models were officially discontinued on **May 25, 2026** and are no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-
- kimi-k2-thinking-maaskimi-k2-thinking-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/moonshotai/model-garden/kimi-k2> Self-deploy Kimi K2 on Model
- kimi-k2-thinking-turboThe `kimi-k2` series models were officially discontinued on **May 25, 2026** and are no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-
- kimi-k2-turbo-previewThe `kimi-k2` series models were officially discontinued on **May 25, 2026** and are no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-
- Kimi-K2.5| Kimi-K2.5 | 1 | Preview | 2027-01-26 | — |
- Kimi-K2.6| Kimi-K2.6 | 2026-04-20 | Preview | 2027-04-16 | — |
- Kimi-K2.7-Code| Kimi-K2.7-Code | 2026-06-12 | Preview | 2026-10-03 | — |
- kimi-latest`kimi-latest` was officially discontinued on **January 28, 2026** and is no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-k3-quickstar
- kimi-thinking-preview`kimi-thinking-preview` was officially discontinued on **November 11, 2025** and is no longer maintained or supported. We recommend upgrading to the latest model [kimi-k3](/docs/gu
- labs-devstral-small-2512<https://docs.mistral.ai/models/devstral-small-2-25-12> Devstral Small 2 ↗ | 25.12 | labs-devstral-small-2512 | 2/27/2026 3/31/2026 | <https://docs.mistral.ai/models/mistral-medium
- labs-leanstral-2603Model | Version | API | Deprecation Retirement | Alternative | Scroll for more <https://docs.mistral.ai/models/leanstral-26-03> Leanstral ↗ | 26.03 | labs-leanstral-2603 | 5/22/202
- labs-mistral-small-creative<https://docs.mistral.ai/models/mistral-small-creative-25-12> Mistral Small Creative ↗ | 25.12 | labs-mistral-small-creative | 3/31/2026 4/30/2026 | <https://docs.mistral.ai/models
- liveportraitliveportrait | / | / | wan2.7-r2v | $0.086012/sec | / |
- liveportrait-detectliveportrait-detect | / | / | qwen-image-3.0 | / | $0.03/image |
- llama3-8b-8192| llama3-8b-8192 | 08/30/25 | llama-3.1-8b-instant |
- llama3-70b-8192August 30, 2025: llama3-70b-8192 and llama3-8b-8192 In line with our commitment to bringing you cutting-edge models, on May 31, 2025, we emailed users to announce the deprecation o
- llama3-groq-8b-8192-tool-use-preview| llama3-groq-8b-8192-tool-use-preview | 1/6/25 | llama-3.3-70b-versatile |
- llama3-groq-70b-8192-tool-use-preview| llama3-groq-70b-8192-tool-use-preview | 1/6/25 | llama-3.3-70b-versatile |
- llama3.1-8b2026-05-27 Deprecated llama3.1-8b and qwen-3-235b-a22b-instruct-2507 We recommend migrating to GPT OSS 120B .
- llama-3.1-8b-instantAugust 16, 2026: llama-3.1-8b-instant and llama-3.3-70b-versatile In line with our commitment to bringing you cutting-edge models, on June 17, 2026, we emailed users to announce th
- llama3.1-70b2025-01-17 The llama3.1-70b model has been deprecated. We recommend transitioning to llama-3.3-70b, a more powerful model released by Meta.
- llama-3.1-70b-specdec| llama-3.1-70b-specdec | 01/24/25 | llama-3.3-70b-specdec |
- llama-3.1-70b-versatileJanuary 24, 2025: Llama 3.1 70B and Llama 3.1 70B (Speculative Decoding) On December 6, 2024, in partnership with Meta, we released llama-3.3-70b-versatile and llama-3.3-70b-specde
- Llama-3.1-Swallow-70B-Instruct-v0.3| `Llama-3.1-Swallow-70B-Instruct-v0.3` | 4/14/2025 | `Llama-3.3-Swallow-70B-Instruct-v0.4` |
- Llama-3.1-Tulu-3-405BLlama-3.1-Tulu-3-405B | 4/14/2025 | DeepSeek-V3-0324, Meta-Llama-3.1-405B-Instruct |
- llama-3.2-1b-previewApril 14, 2025: Multiple Model Deprecations In line with our commitment to bringing you cutting-edge models, on April 7, 2025, we emailed users to announce the deprecation of sever
- llama-3.2-3b-preview| llama-3.2-3b-preview | 04/14/25 | llama-3.1-8b-instant |
- llama-3.2-11b-text-preview| llama-3.2-11b-text-preview | 10/28/24 | llama-3.2-11b-vision-preview llama-3.1-8b-instant (text-only workloads) |
- Llama-3.2-11B-Vision-Instruct| Llama-3.2-11B-Vision-Instruct | — | Retired | 2026-06-13 | — |
- llama-3.2-11b-vision-preview| llama-3.2-11b-vision-preview | 04/14/25 | meta-llama/llama-4-scout-17b-16e-instruct |
- llama-3.2-90b-text-preview| llama-3.2-90b-text-preview | 11/25/24 | llama-3.2-90b-vision-preview llama-3.1-70b-versatile (text-only workloads) |
- Llama-3.2-90B-Vision-Instruct| Llama-3.2-90B-Vision-Instruct | — | Retired | 2026-06-13 | — |
- llama-3.2-90b-vision-preview| llama-3.2-90b-vision-preview | 04/14/25 | meta-llama/llama-4-scout-17b-16e-instruct |
- llama-3.3-70b2026-02-16 Deprecated qwen-3-32b and llama-3.3-70b We recommend migrating to GPT OSS 120B .
- llama-3.3-70b-instruct-maasllama-3.3-70b-instruct-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/meta/model-garden/llama3-3> Self-deploy Llama 3.3 on Mo
- llama-3.3-70b-specdec| llama-3.3-70b-specdec | 04/14/25 | meta-llama/llama-4-scout-17b-16e-instruct llama-3.3-70b-versatile |
- llama-3.3-70b-versatile| llama-3.3-70b-versatile | 08/16/26 | openai/gpt-oss-120b or qwen/qwen3.6-27b |
- llama-4-maverick-17b-128e-instruct2025-10-15 Deprecated llama-4-maverick-17b-128e-instruct We recommend migrating to GPT OSS 120B or Llama 3.3 70B .
- llama-4-scout-17b-16e-instruct2025-11-03 Deprecated llama-4-scout-17b-16e-instruct We recommend migrating to GPT OSS 120B or Llama 3.3 70B .
- llama-guard-3-8bJune 6, 2025: Llama Guard 3 In line with our commitment to bringing you cutting-edge models, on May 9, 2025, we emailed users to announce the deprecation of llama-guard-3-8b in fav
- llava-v1.5-7b-4096-preview| llava-v1.5-7b-4096-preview | 10/28/24 | llama-3.2-11b-vision-preview |
- lmsys/vicuna-7b-v1.5| 2024-08-22 | `lmsys/vicuna-7b-v1.5` | No |
- lmsys/vicuna-13b-v1.5| 2024-08-22 | `lmsys/vicuna-13b-v1.5` | No |
- lyria-3-clip-previewlyria-3-clip-preview | March 25, 2026 | No shutdown date announced | |
- lyria-3-pro-previewlyria-3-pro-preview | March 25, 2026 | No shutdown date announced | lyria-3.5 |
- lyria-3.5Lyria models Model | Release date | Shutdown date | Recommended replacement | lyria-3.5 | September 3, 2026 | No shutdown date announced | |
- lyria-realtime-explyria-realtime-exp | May 20, 2025 | No shutdown date announced | |
- magistral-medium-2506<https://docs.mistral.ai/models/magistral-medium-1-0-25-06> Magistral Medium 1.0 ↗ | 25.06 | magistral-medium-2506 | 10/31/2025 11/30/2025 | <https://docs.mistral.ai/models/mistral
- magistral-medium-2507<https://docs.mistral.ai/models/magistral-medium-1-1-25-07> Magistral Medium 1.1 ↗ | 25.07 | magistral-medium-2507 | 10/31/2025 11/30/2025 | <https://docs.mistral.ai/models/mistral
- magistral-medium-2509<https://docs.mistral.ai/models/magistral-medium-1-2-25-09> Magistral Medium 1.2 ↗ | 25.09 | magistral-medium-2509 | 5/22/2026 7/31/2026 | <https://docs.mistral.ai/models/mistral-m
- magistral-small-2506<https://docs.mistral.ai/models/magistral-small-1-0-25-06> Magistral Small 1.0 ↗ | 25.06 | magistral-small-2506 | 10/31/2025 11/30/2025 | <https://docs.mistral.ai/models/mistral-sm
- magistral-small-2507<https://docs.mistral.ai/models/magistral-small-1-1-25-07> Magistral Small 1.1 ↗ | 25.07 | magistral-small-2507 | 10/31/2025 11/30/2025 | <https://docs.mistral.ai/models/mistral-sm
- magistral-small-2509<https://docs.mistral.ai/models/magistral-small-1-2-25-09> Magistral Small 1.2 ↗ | 25.09 | magistral-small-2509 | 4/30/2026 7/31/2026 | <https://docs.mistral.ai/models/mistral-smal
- MAI-Code-1-FlashModel | Deprecation date | Suggested alternative | MAI-Code-1-Flash | 9-10-2026 | MAI-Code-1.1-Flash |
- MAI-Image-2| MAI-Image-2 | 2026-02-20 | Retired | 2026-08-15 | MAI-Image-2.5 |
- MAI-Image-2e| MAI-Image-2e | 2026-04-09 | Retired | 2026-08-15 | MAI-Image-2.5-Flash |
- MAI-Image-2.5| MAI-Image-2.5 | 2026-06-02 | Preview | 2026-10-01 | — |
- MAI-Image-2.5-Flash| MAI-Image-2.5-Flash | 2026-06-02 | Preview | 2026-10-01 | — |
- MAI-Image-2.5-Pro| MAI-Image-2.5-Pro | 2026-06-19 | Preview | 2026-10-01 | — |
- MAI-Transcribe-1| MAI-Transcribe-1 | 2026-01-23 | Preview | 2026-09-15 | MAI-Transcribe-1.5 |
- marin-community/Marin-8B-Instruct| 2026-02-25 | `marin-community/Marin-8B-Instruct` | No |
- Meta Llama 2 7BMeta Llama 2 7B | Provisioned throughput: February 27, 2026 | Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size. |
- Meta Llama 2 13BMeta Llama 2 13B | Provisioned throughput: February 27, 2026 | Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size. |
- meta.llama-2-70bmeta.llama-2-70b | 2024-01-22 | 2024-10-02 |
- Meta Llama 3 8BMeta Llama 3 8B | Provisioned throughput: February 27, 2026 | Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size. |
- Meta Llama 3 (70B)Meta Llama 3 (70B) | Pay-per-token: July 23, 2024 (Meta-Llama-3-70B-Instruct); December 11, 2024 (Meta-Llama-3.1-70B-Instruct) Provisioned throughput: February 27, 2026 | Pay-per-t
- meta.llama-3-70b-instructmeta.llama-3-70b-instruct | 2024-06-04 | 2024-11-12 |
- Meta-Llama-3.1-8B| Meta-Llama-3.1-8B | — | Retired | 2026-06-13 | — |
- Meta-Llama-3.1-8B-Instruct| Meta-Llama-3.1-8B-Instruct | — | Retired | 2026-06-13 | — |
- meta.llama-3.1-70b-instructmeta.llama-3.1-70b-instruct | 2024-09-19 | 2025-07-10 |
- Meta-Llama-3.1-405B-Instruct| Meta-Llama-3.1-405B-Instruct | — | Retired | 2026-06-13 | — |
- Meta-Llama-3.2-1B-Instruct| `Meta-Llama-3.2-1B-Instruct` | 6/25/2025 | `Meta-Llama-3.1-8B-Instruct` |
- Meta-Llama-3.2-3B-Instruct| `Meta-Llama-3.2-3B-Instruct` | 6/25/2025 | `Meta-Llama-3.1-8B-Instruct` |
- meta.llama-3.2-11b-vision-instructmeta.llama-3.2-11b-vision-instruct | 2024-11-14 | 2026-08-03 |
- meta.llama-3.2-90b-vision-instructmeta.llama-3.2-90b-vision-instruct | 2024-11-14 | 2026-08-03 |
- Meta-Llama-Guard-3-8B| `Meta-Llama-Guard-3-8B` | 6/25/2025 | `N/A` |
- meta-llama/Llama-2-7b-chat-hf| 2024-08-22 | `meta-llama/Llama-2-7b-chat-hf` | No |
- meta-llama/Llama-2-7b-hf| 2024-08-22 | `meta-llama/Llama-2-7b-hf` | No |
- meta-llama/llama-2-13b-chatllama-2-13b-chat meta-llama/llama-2-13b-chat | 8 January 2025 | – | 30 January 2026 |
- meta-llama/Llama-2-13b-chat-hf| 2025-04-24 | `meta-llama/Llama-2-13b-chat-hf` | No |
- meta-llama/Llama-2-13b-hf| 2024-08-22 | `meta-llama/Llama-2-13b-hf` | No |
- meta-llama/Llama-2-70b-chat-hf| 2024-08-22 | `meta-llama/Llama-2-70b-chat-hf` | No |
- meta-llama-llama-3-2-1b-instruct-lora| 2025-04-24 | `meta-llama-llama-3-2-1b-instruct-lora` | No |
- meta-llama-llama-3-2-3b-instruct-turbo-lora| 2025-05-16 | `meta-llama-llama-3-2-3b-instruct-turbo-lora` | No |
- meta-llama-llama-3-3-70b-instruct-lora| 2025-08-28 | `meta-llama-llama-3-3-70b-instruct-lora` | No |
- meta-llama/Llama-3-8b-chat-hf| 2025-08-28 | `meta-llama/Llama-3-8b-chat-hf` | No |
- meta-llama/Llama-3-8b-hf| 2024-08-22 | `meta-llama/Llama-3-8b-hf` | No |
- meta-llama/llama-3-405b-instructllama-3-405b-instruct meta-llama/llama-3-405b-instruct | 23 July 2024 | 24 November 2025 | 31 March 2026 | llama-4-maverick-17b-128e-instruct-fp8 |
- meta-llama/Llama-3.2-3B-Instruct-Turbo| 2026-03-06 | `meta-llama/Llama-3.2-3B-Instruct-Turbo` | No |
- meta-llama/Llama-3.2-3B-Instruct-Turbo-Classifier| 2026-02-25 | `meta-llama/Llama-3.2-3B-Instruct-Turbo-Classifier` | No |
- meta-llama/Llama-3.2-11B-Vision-Instruct-Turbo| 2025-08-28 | `meta-llama/Llama-3.2-11B-Vision-Instruct-Turbo` | No |
- meta-llama/Llama-3.2-90B-Vision-Instruct-Turbo| 2025-08-28 | `meta-llama/Llama-3.2-90B-Vision-Instruct-Turbo` | No |
- meta-llama/Llama-3.3-70B-Instruct| [`meta-llama/Llama-3.3-70B-Instruct`](https://tokenfactory.nebius.com/models/catalog/text2text/meta-llama%2FLlama-3.3-70B-Instruct) | [`nvidia/Nemotron-3_5-Lightning`
- meta-llama/Llama-3.3-70B-Instruct-Turbo-Free| 2025-11-13 | `meta-llama/Llama-3.3-70B-Instruct-Turbo-Free` | No |
- meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8| 2026-03-31 | `meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8` | Yes |
- meta-llama/llama-guard-3-11b-visionllama-guard-3-11b-vision (Dallas, Tokyo, Sydney data centers) meta-llama/llama-guard-3-11b-vision | 22 January 2025 | 8 May 2026 | 8 August 2026 | llama-3-2-90b-vision-instruct on
- meta-llama/Llama-Guard-3-11B-Vision-Turbo| 2026-02-25 | `meta-llama/Llama-Guard-3-11B-Vision-Turbo` | No |
- meta-llama/llama-guard-4-12bMarch 5, 2026: meta-llama/llama-guard-4-12b In line with our commitment to bringing you cutting-edge models, on February 10, 2026, we emailed users to announce the deprecation of m
- Meta-Llama/Llama-Guard-7b| 2025-03-08 | `Meta-Llama/Llama-Guard-7b` | No |
- meta-llama/Llama-Vision-Free| 2025-08-28 | `meta-llama/Llama-Vision-Free` | No |
- meta-llama/LlamaGuard-2-8b| 2026-02-25 | `meta-llama/LlamaGuard-2-8b` | No |
- meta-llama-meta-llama-3-1-8b-instruct-turbo-lora| 2025-04-24 | `meta-llama-meta-llama-3-1-8b-instruct-turbo-lora` | No |
- meta-llama-meta-llama-3-1-70b-instruct-turbo-lora| 2025-04-24 | `meta-llama-meta-llama-3-1-70b-instruct-turbo-lora` | No |
- meta-llama/Meta-Llama-3-8B-Instruct| 2025-08-28 | `meta-llama/Meta-Llama-3-8B-Instruct` | No |
- meta-llama/Meta-Llama-3-8B-Instruct-Lite| 2026-07-10 | `meta-llama/Meta-Llama-3-8B-Instruct-Lite` | No |
- meta-llama-meta-llama-3-8b-instruct-turbo| 2025-05-16 | `meta-llama-meta-llama-3-8b-instruct-turbo` | No |
- meta-llama/Meta-Llama-3-70B-Instruct-Lite| 2025-03-11 | `meta-llama/Meta-Llama-3-70B-Instruct-Lite` | No |
- meta-llama/Meta-Llama-3-70B-Instruct-Turbo| 2025-12-23 | `meta-llama/Meta-Llama-3-70B-Instruct-Turbo` | No |
- meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo| 2026-03-06 | `meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo` | Yes |
- meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo| 2026-02-25 | `meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo` | No |
- meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo| 2026-02-06 | `meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo` | No |
- microsoft/phi-2| 2024-08-22 | `microsoft/phi-2` | No |
- microsoft-wizardlm-2-8x22b| 2025-04-24 | `microsoft-wizardlm-2-8x22b` | No |
- minimax-m2-maasminimax-m2-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/minimaxai/model-garden/minimax-m2> Self-deploy MiniMax M2 on Model
- MiniMax-M2.5| `MiniMax-M2.5` | 5/18/2026 | `MiniMax-M2.7` |
- MiniMaxAI/MiniMax-M2.5-fast| `MiniMaxAI/MiniMax-M2.5-fast` |
- ministral-3b-2410<https://docs.mistral.ai/models/ministral-3b-24-1> Ministral 3B ↗ | 24.1 | ministral-3b-2410 | 12/2/2025 12/31/2025 | <https://docs.mistral.ai/models/ministral-3-3b-25-12> Ministra
- ministral-8b-2410<https://docs.mistral.ai/models/ministral-8b-24-1> Ministral 8B ↗ | 24.1 | ministral-8b-2410 | 12/2/2025 12/31/2025 | <https://docs.mistral.ai/models/ministral-3-8b-25-12> Ministra
- Mistral 7BMistral 7B | Provisioned throughput: February 27, 2026 | Comparable model on the same offering, like Llama 3.2, 3.3, or 4 model of similar size. |
- mistral-document-ai-2505| mistral-document-ai-2505 | 1 | Retired | 2026-07-20 | mistral-document-ai-2512, mistral-ocr-4-0 |
- mistral-large-2402<https://docs.mistral.ai/models/mistral-large-1-0-24-02> Mistral Large 1.0 ↗ | 24.02 | mistral-large-2402 | 11/30/2024 6/16/2025 | <https://docs.mistral.ai/models/mistral-large-3-2
- mistral-large-2407<https://docs.mistral.ai/models/mistral-large-2-0-24-07> Mistral Large 2.0 ↗ | 24.07 | mistral-large-2407 | 11/30/2024 3/30/2025 | <https://docs.mistral.ai/models/mistral-large-3-2
- mistral-large-2411<https://docs.mistral.ai/models/mistral-large-2-1-24-11> Mistral Large 2.1 ↗ | 24.11 | mistral-large-2411 | 2/27/2026 5/31/2026 | <https://docs.mistral.ai/models/mistral-medium-3-5
- mistral-medium-2312<https://docs.mistral.ai/models/mistral-medium-1-0-23-12> Mistral Medium 1.0 ↗ | 23.12 | mistral-medium-2312 | 11/30/2024 6/16/2025 | <https://docs.mistral.ai/models/mistral-medium
- mistral-medium-2505<https://docs.mistral.ai/models/mistral-medium-3-25-05> Mistral Medium 3 ↗ | 25.05 | mistral-medium-2505 | 5/22/2026 8/31/2026 | <https://docs.mistral.ai/models/mistral-medium-3-5-
- mistral-medium-2508<https://docs.mistral.ai/models/mistral-medium-3-1-25-08> Mistral Medium 3.1 ↗ | 25.08 | mistral-medium-2508 | 5/22/2026 8/31/2026 | <https://docs.mistral.ai/models/mistral-medium-
- mistral-moderation-2411<https://docs.mistral.ai/models/mistral-moderation-24-11> Mistral Moderation ↗ | 24.11 | mistral-moderation-2411 | 3/31/2026 6/30/2026 | <https://docs.mistral.ai/models/mistral-mod
- mistral-ocr-2503<https://docs.mistral.ai/models/ocr-25-03> OCR ↗ | 25.03 | mistral-ocr-2503 | 12/2/2025 12/31/2025 | <https://docs.mistral.ai/models/ocr-4-1> OCR 4.1 |
- mistral-ocr-2505<https://docs.mistral.ai/models/ocr-2-25-05> OCR 2 ↗ | 25.05 | mistral-ocr-2505 | 2/27/2026 5/31/2026 | <https://docs.mistral.ai/models/ocr-4-1> OCR 4.1 |
- mistral-saba-24b| mistral-saba-24b | 07/30/25 | qwen/qwen3-32b |
- mistral-saba-2502<https://docs.mistral.ai/models/mistral-saba-25-02> Mistral Saba ↗ | 25.02 | mistral-saba-2502 | 6/10/2025 9/30/2025 | <https://docs.mistral.ai/models/mistral-small-4-0-26-03> Mist
- mistral-small-2402<https://docs.mistral.ai/models/mistral-small-1-0-24-02> Mistral Small 1.0 ↗ | 24.02 | mistral-small-2402 | 11/30/2024 6/16/2025 | <https://docs.mistral.ai/models/mistral-small-4-0
- mistral-small-2409<https://docs.mistral.ai/models/mistral-small-2-0-24-09> Mistral Small 2.0 ↗ | 24.09 | mistral-small-2409 | 11/6/2025 11/30/2025 | <https://docs.mistral.ai/models/mistral-small-4-0
- mistral-small-2501<https://docs.mistral.ai/models/mistral-small-3-0-25-01> Mistral Small 3.0 ↗ | 25.01 | mistral-small-2501 | 11/6/2025 11/30/2025 | <https://docs.mistral.ai/models/mistral-small-4-0
- mistral-small-2503<https://docs.mistral.ai/models/mistral-small-3-1-25-03> Mistral Small 3.1 ↗ | 25.03 | mistral-small-2503 | 11/6/2025 11/30/2025 | <https://docs.mistral.ai/models/mistral-small-4-0
- mistral-small-2506<https://docs.mistral.ai/models/mistral-small-3-2-25-06> Mistral Small 3.2 ↗ | 25.06 | mistral-small-2506 | 4/30/2026 7/31/2026 | <https://docs.mistral.ai/models/mistral-small-4-0-
- mistralai/Ministral-3-14B-Instruct| 2026-02-25 | `mistralai/Ministral-3-14B-Instruct` | No |
- mistralai/Mistral-7B-Instruct-v0.1| 2025-11-13 | `mistralai/Mistral-7B-Instruct-v0.1` | No |
- mistralai/Mistral-7B-Instruct-v0.1-json| 2024-08-22 | `mistralai/Mistral-7B-Instruct-v0.1-json` | No |
- mistralai/Mistral-7B-Instruct-v0.1-tools| 2024-08-22 | `mistralai/Mistral-7B-Instruct-v0.1-tools` | No |
- mistralai/Mistral-7B-Instruct-v0.3| `mistralai/Mistral-7B-Instruct-v0.3` | `mistralai/Ministral-3-14B-Instruct-2512` | Same lineage, upgraded version |
- mistralai/Mistral-7B-v0.1| 2025-03-27 | `mistralai/Mistral-7B-v0.1` | No |
- mistralai/mistral-large-2512mistral-large-2512 (Dallas data center) mistralai/mistral-large-2512 | 12 December 2025 | 8 May 2026 | 12 January 2027 | mistral-large-2512 on Deploy on demand |
- mistralai/Mistral-Small-24B-Instruct-2501| 2026-04-02 | `mistralai/Mistral-Small-24B-Instruct-2501` | No |
- mistralai/Mixtral-8x7B-Instruct-v0.1| 2026-04-16 | `mistralai/Mixtral-8x7B-Instruct-v0.1` | Yes |
- mistralai-mixtral-8x7b-v0-1| 2025-06-13 | `mistralai-mixtral-8x7b-v0-1` | No |
- mistralai/Mixtral-8x22B| 2024-08-22 | `mistralai/Mixtral-8x22B` | No |
- mistralai-mixtral-8x22b-instruct-v0-1| 2025-06-13 | `mistralai-mixtral-8x22b-instruct-v0-1` | No |
- mistralai/Voxtral-Mini-3B-2507| 2026-06-11 | `mistralai/Voxtral-Mini-3B-2507` | No |
- mixedbread-ai/Mxbai-Rerank-Large-V2| 2026-03-06 | `mixedbread-ai/Mxbai-Rerank-Large-V2` | No |
- mixtral-8x7b-32768March 20, 2025: Mixtral 8x7B On March 5, 2025, we emailed all users of the mixtral-8x7b-32768 model that we would be deprecating this model ID in favor of newer, more performant mo
- moonshot-v1-8kThe `moonshot-v1` series models (including `moonshot-v1-auto` and the `-vision-preview` variants) were officially discontinued on **August 31, 2026** and are no longer maintained o
- moonshot-v1-8k-vision-previewThe `moonshot-v1` series models (including `moonshot-v1-auto` and the `-vision-preview` variants) were officially discontinued on **August 31, 2026** and are no longer maintained o
- moonshot-v1-32kThe `moonshot-v1` series models (including `moonshot-v1-auto` and the `-vision-preview` variants) were officially discontinued on **August 31, 2026** and are no longer maintained o
- moonshot-v1-32k-vision-previewThe `moonshot-v1` series models (including `moonshot-v1-auto` and the `-vision-preview` variants) were officially discontinued on **August 31, 2026** and are no longer maintained o
- moonshot-v1-128kThe `moonshot-v1` series models (including `moonshot-v1-auto` and the `-vision-preview` variants) were officially discontinued on **August 31, 2026** and are no longer maintained o
- moonshot-v1-128k-vision-previewThe `moonshot-v1` series models (including `moonshot-v1-auto` and the `-vision-preview` variants) were officially discontinued on **August 31, 2026** and are no longer maintained o
- moonshot-v1-autoThe `moonshot-v1` series models (including `moonshot-v1-auto` and the `-vision-preview` variants) were officially discontinued on **August 31, 2026** and are no longer maintained o
- moonshotai/kimi-k2-instructOctober 10, 2025: moonshotai/kimi-k2-instruct In line with our commitment to bringing you cutting-edge models, on September 10, 2025, we emailed users to announce the deprecation o
- moonshotai/kimi-k2-instruct-0905April 15, 2026: moonshotai/kimi-k2-instruct-0905 In line with our commitment to bringing you cutting-edge models, on March 23, 2026, we emailed users to announce the deprecation of
- moonshotai/Kimi-K2.5-fast| `moonshotai/Kimi-K2.5-fast` |
- multilingual-e5-large-instruct-maasmultilingual-e5-large-instruct-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/intfloat/model-garden/e5> Self-deploy E5 Text E
- multilingual-e5-small-maasmultilingual-e5-small-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/intfloat/model-garden/e5> Self-deploy E5 Text Embedding
- multimodalembedding@001multimodalembedding@001 | February 12, 2024 | April 1, 2027 | |
- Nexusflow/NexusRaven-V2-13B| 2024-08-22 | `Nexusflow/NexusRaven-V2-13B` | No |
- nlp-rag-rewrite-onenlp-rag-rewrite-one | / | / | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- NousResearch/Hermes-3-Llama-3.1-405B-Turbo| 2024-10-07 | `NousResearch/Hermes-3-Llama-3.1-405B-Turbo` | No |
- NousResearch/Hermes-4-70B| [`NousResearch/Hermes-4-70B`](https://tokenfactory.nebius.com/models/catalog/text2text/NousResearch%2FHermes-4-70B) | [`nvidia/Nemotron-3_5-Lightning`
- NousResearch/Nous-Capybara-7B-V1p9| 2024-08-22 | `NousResearch/Nous-Capybara-7B-V1p9` | No |
- NousResearch/Nous-Hermes-2-Mistral-7B-DPO| 2024-08-22 | `NousResearch/Nous-Hermes-2-Mistral-7B-DPO` | No |
- NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO| 2025-08-28 | `NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO` | No |
- NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT| 2024-08-22 | `NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT` | No |
- NousResearch/Nous-Hermes-2-Yi-34B| 2024-10-07 | `NousResearch/Nous-Hermes-2-Yi-34B` | No |
- NousResearch/Nous-Hermes-llama-2-7b| 2024-08-22 | `NousResearch/Nous-Hermes-llama-2-7b` | No |
- NousResearch/Nous-Hermes-Llama2-13b| 2024-08-22 | `NousResearch/Nous-Hermes-Llama2-13b` | No |
- nvidia/Cosmos3-Super-Reasoner| [`nvidia/Cosmos3-Super-Reasoner`](https://tokenfactory.nebius.com/models/catalog/image2text/nvidia%2FCosmos3-Super-Reasoner) | [`MiniMaxAI/MiniMax-M3`](https:/
- nvidia/Llama-3_1-Nemotron-Ultra-253B-v1| [`nvidia/Llama-3_1-Nemotron-Ultra-253B-v1`](https://tokenfactory.nebius.com/models/catalog/text2text/nvidia%2FLlama-3_1-Nemotron-Ultra-253B-v1) | [`nvidia/nemotron-3-super-120b-a
- nvidia/Llama-3.1-Nemotron-70B-Instruct-HF| 2025-08-28 | `nvidia/Llama-3.1-Nemotron-70B-Instruct-HF` | No |
- nvidia/Nemotron-3-Nano-Omni| [`nvidia/Nemotron-3-Nano-Omni`](https://tokenfactory.nebius.com/models/catalog/text2text/nvidia%2FNemotron-3-Nano-Omni) | [`nvidia/Nemotron-3_5-Lightning`
- nvidia/Nemotron-3-ultra-550b-a55b| 2026-08-27 | `nvidia/Nemotron-3-ultra-550b-a55b` | No |
- nvidia/Nvidia-Nemotron-Nano-9B-v2| 2026-02-25 | `nvidia/Nvidia-Nemotron-Nano-9B-v2` | No |
- o1| o1 | 2024-12-17 | Deprecated | 2026-11-19 | gpt-5.6-sol |
- o1-2024-12-17| October 23, 2026 | `o1-2024-12-17` \| `o1` | `gpt-5.6-sol` |
- o1-mini| 2025-10-27 | `o1-mini` | `o4-mini` |
- o1-preview| 2025-07-28 | `o1-preview` | `o3` |
- o1-pro| o1-pro | 2025-03-19 | Deprecated | 2026-11-19 | gpt-5.6-sol |
- o1-pro-2025-03-19October 23, 2026 | o1-pro-2025-03-19 | o1-pro | gpt-5.6-sol (reasoning.mode: pro) |
- o3| o3 | 2025-04-16 | Deprecated | 2026-11-19 | gpt-5.6-sol |
- o3-2025-04-16| Dec 11, 2026 | `o3-2025-04-16` | `gpt-5.6-sol` |
- o3-deep-research| o3-deep-research | 2025-06-26 | Deprecated | 2026-11-19 | gpt-5.6-sol |
- o3-deep-research-2025-06-26| July 23, 2026 | `o3-deep-research-2025-06-26` \| `o3-deep-research` | `gpt-5.6-sol` |
- o3-mini| o3-mini | 2025-01-31 | Deprecated | 2026-11-19 | gpt-5.6-terra |
- o3-mini-2025-01-31| October 23, 2026 | `o3-mini-2025-01-31` \| `o3-mini` | `gpt-5.6-sol` |
- o3-pro| o3-pro | 2025-06-10 | Deprecated | 2026-11-19 | gpt-5.6-sol |
- o3-pro-2025-06-10Dec 11, 2026 | o3-pro-2025-06-10 | gpt-5.6-sol (reasoning.mode: pro) |
- o4-mini| o4-mini | 2025-04-16 | Deprecated | 2026-11-19 | gpt-5.6-terra |
- o4-mini-2025-04-16| October 23, 2026 | `o4-mini-2025-04-16` \| `o4-mini` | `gpt-5.6-terra` |
- o4-mini-deep-research-2025-06-26| July 23, 2026 | `o4-mini-deep-research-2025-06-26` \| `o4-mini-deep-research` | `gpt-5.6-sol` |
- open-codestral-mamba<https://docs.mistral.ai/models/codestral-mamba-7b-0-1> Codestral Mamba 7B ↗ | 0.1 | open-codestral-mamba | 6/6/2025 6/6/2025 | <https://docs.mistral.ai/models/codestral-25-08> Cod
- open-mistral-7b<https://docs.mistral.ai/models/mistral-7b-0-3> Mistral 7B ↗ | 0.3 | open-mistral-7b | 11/30/2024 3/30/2025 | <https://docs.mistral.ai/models/ministral-3-8b-25-12> Ministral 3 8B |
- open-mistral-nemo-2407<https://docs.mistral.ai/models/mistral-nemo-12b-24-07> Mistral Nemo 12B ↗ | 24.07 | open-mistral-nemo-2407 | 5/22/2026 7/31/2026 | <https://docs.mistral.ai/models/ministral-3-8b-2
- open-mixtral-8x7b<https://docs.mistral.ai/models/mixtral-8x7b-0-1> Mixtral 8x7B ↗ | 0.1 | open-mixtral-8x7b | 11/30/2024 3/30/2025 | <https://docs.mistral.ai/models/mistral-small-4-0-26-03> Mistral
- open-mixtral-8x22b<https://docs.mistral.ai/models/mixtral-8x22b-0-1-0-3> Mixtral 8x22B ↗ | 0.1-0.3 | open-mixtral-8x22b | 11/30/2024 3/30/2025 | <https://docs.mistral.ai/models/mistral-small-4-0-26-
- Open-Orca/Mistral-7B-OpenOrca| 2024-08-22 | `Open-Orca/Mistral-7B-OpenOrca` | No |
- openai/gpt-oss-20b| 2026-09-14 | `openai/gpt-oss-20b` | `Qwen/Qwen3.5-9B` | Yes |
- openai/gpt-oss-120b-fast| `openai/gpt-oss-120b-fast` |
- openchat/openchat-3.5-1210| 2024-08-22 | `openchat/openchat-3.5-1210` | No |
- DBRX / DBRX Instruct · Pay-per-tokenDBRX / DBRX Instruct | Pay-per-token: April 30, 2025 Provisioned throughput: December 19, 2025 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Comparable model on th
- Meta Llama 2 70B / Meta-Llama-2-70B-Chat · Pay-per-tokenMeta Llama 2 70B / Meta-Llama-2-70B-Chat | Pay-per-token: October 30, 2024 Provisioned throughput: February 27, 2026 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput:
- Meta Llama 3.1 405B · Pay-per-tokenMeta Llama 3.1 405B | Pay-per-token: February 15, 2026 Provisioned throughput: May 15, 2026 | OpenAI GPT OSS 120B |
- Mixtral 8x7B / Mixtral-8x7B Instruct · Pay-per-tokenMixtral 8x7B / Mixtral-8x7B Instruct | Pay-per-token: April 30, 2025 Provisioned throughput: February 27, 2026 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Compar
- MPT 7B / MPT 7B Instruct · Pay-per-tokenMPT 7B / MPT 7B Instruct | Pay-per-token: August 30, 2024 Provisioned throughput: December 19, 2025 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Comparable model
- MPT 30B / MPT 30B Instruct · Pay-per-tokenMPT 30B / MPT 30B Instruct | Pay-per-token: August 30, 2024 Provisioned throughput: December 19, 2025 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Comparable mode
- pearl-ai/gemma-4-31b-it| 2026-08-27 | `pearl-ai/gemma-4-31b-it` | No |
- perplexity-ai/r1-1776| 2025-08-28 | `perplexity-ai/r1-1776` | No |
- Phind/Phind-CodeLlama-34B-v2| 2024-08-22 | `Phind/Phind-CodeLlama-34B-v2` | No |
- pixtral-12b-2409<https://docs.mistral.ai/models/pixtral-12b-24-09> Pixtral 12B ↗ | 24.09 | pixtral-12b-2409 | 12/2/2025 12/31/2025 | <https://docs.mistral.ai/models/ministral-3-14b-25-12> Ministra
- pixtral-large-2411<https://docs.mistral.ai/models/pixtral-large-24-11> Pixtral Large ↗ | 24.11 | pixtral-large-2411 | 2/27/2026 5/31/2026 | <https://docs.mistral.ai/models/mistral-medium-3-5-26-04>
- playai-ttsDecember 31, 2025: playai-tts and playai-tts-arabic In line with our commitment to bringing you cutting-edge models, on December 23, 2025, we emailed users to announce the deprecat
- playai-tts-arabic| playai-tts-arabic | 12/31/25 | canopylabs/orpheus-arabic-saudi |
- PrimeIntellect/INTELLECT-3| `PrimeIntellect/INTELLECT-3` |
- prompthero/openjourney| 2024-08-22 | `prompthero/openjourney` | No |
- DBRX / DBRX Instruct · Provisioned throughputDBRX / DBRX Instruct | Pay-per-token: April 30, 2025 Provisioned throughput: December 19, 2025 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Comparable model on th
- Meta Llama 2 70B / Meta-Llama-2-70B-Chat · Provisioned throughputMeta Llama 2 70B / Meta-Llama-2-70B-Chat | Pay-per-token: October 30, 2024 Provisioned throughput: February 27, 2026 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput:
- Meta Llama 3.1 405B · Provisioned throughputMeta Llama 3.1 405B | Pay-per-token: February 15, 2026 Provisioned throughput: May 15, 2026 | OpenAI GPT OSS 120B |
- Mixtral 8x7B / Mixtral-8x7B Instruct · Provisioned throughputMixtral 8x7B / Mixtral-8x7B Instruct | Pay-per-token: April 30, 2025 Provisioned throughput: February 27, 2026 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Compar
- MPT 7B / MPT 7B Instruct · Provisioned throughputMPT 7B / MPT 7B Instruct | Pay-per-token: August 30, 2024 Provisioned throughput: December 19, 2025 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Comparable model
- MPT 30B / MPT 30B Instruct · Provisioned throughputMPT 30B / MPT 30B Instruct | Pay-per-token: August 30, 2024 Provisioned throughput: December 19, 2025 | Pay-per-token: Meta-Llama-4-Maverick Provisioned throughput: Comparable mode
- qvq-maxQwen Series Legacy Mainline Models:qwen-turbo、qwen-turbo-realtime、qwen-vl-max、qwen-vl-plus、qwq-plus、qvq-max、qvq-plus、qwen-math-turbo、qwen-coder-turbo、qwen-coder-plus
- qvq-plusQwen Series Legacy Mainline Models:qwen-turbo、qwen-turbo-realtime、qwen-vl-max、qwen-vl-plus、qwq-plus、qvq-max、qvq-plus、qwen-math-turbo、qwen-coder-turbo、qwen-coder-plus
- Qwen2-Audio-7B-Instruct| `Qwen2-Audio-7B-Instruct` | 6/19/2025 | `Whisper-Large-v3` |
- qwen-2.5-32b| qwen-2.5-32b | 04/14/25 | qwen-qwq-32b meta-llama/llama-4-scout-17b-16e-instruct |
- Qwen2.5-72B-InstructQwen2.5-72B-Instruct | 4/14/2025 | Meta-Llama-3.3-70B-Instruct, DeepSeek-V3-0324 |
- qwen-2.5-coder-32b| qwen-2.5-coder-32b | 04/14/25 | qwen-qwq-32b openai/gpt-oss-120b |
- Qwen2.5-Coder-32B-InstructQwen2.5-Coder-32B-Instruct | 4/14/2025 | DeepSeek-V3-0324, Llama-4-Maverick-17B-128E-Instruct |
- qwen2.5-omni-7bqwen2.5-omni-7b | Multi-billing from 0.1 | Multi-billing from 0.4 | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen-3-32b2026-02-16 Deprecated qwen-3-32b and llama-3.3-70b We recommend migrating to GPT OSS 120B .
- qwen-3-235b-a22b2025-07-29 Deprecated qwen-3-235b-a22b We recommend migrating to either Qwen 3 235B Instruct or Qwen 3 235B Thinking .
- qwen-3-235b-a22b-instruct-25072026-05-27 Deprecated llama3.1-8b and qwen-3-235b-a22b-instruct-2507 We recommend migrating to GPT OSS 120B .
- qwen3-235b-a22b-instruct-2507-maasqwen3-235b-a22b-instruct-2507-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/qwen/model-garden/qwen3> Self-deploy Qwen3 on Mo
- qwen-3-235b-a22b-thinking-25072025-11-14 Deprecated qwen-3-235b-a22b-thinking-2507 We recommend migrating to GPT OSS 120B .
- qwen3-asr-flash-2025-09-08qwen3-asr-flash-2025-09-08 | $0.000035/sec | / | fun-asr | $0.000035/sec | / |
- qwen3-asr-flash-2026-02-10qwen3-asr-flash-2026-02-10 | $0.000035/sec | / | fun-asr | $0.000035/sec | / |
- qwen3-asr-flash-filetrans-2025-11-17qwen3-asr-flash-filetrans-2025-11-17 | $0.000035/sec | / | fun-asr | $0.000035/sec | / |
- qwen3-asr-flash-realtime-2025-10-27qwen3-asr-flash-realtime-2025-10-27 | $0.00009/sec | / | fun-asr-realtime | / | / |
- qwen3-asr-flash-realtime-2026-02-10qwen3-asr-flash-realtime-2026-02-10 | $0.00009/sec | / | fun-asr-realtime | / | / |
- qwen-3-coder-480b2025-11-05 Deprecated qwen-3-coder-480b We recommend migrating to Z.ai GLM 4.7 .
- qwen3-coder-480b-a35b-instruct-maasqwen3-coder-480b-a35b-instruct-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/qwen/model-garden/qwen3-coder> Self-deploy Qwen
- qwen3-coder-plusQwen3 Series Mainline Models:qwen3.6-max-preview、qwen3-max-preview、qwen3-max、qwen3-vl-flash、qwen3-coder-plus
- qwen3-livetranslate-flash-realtimeqwen3-livetranslate-flash-realtime | Multi-billing from 1.3 | Multi-billing from 10 | qwen3.5-livetranslate-flash-realtime | Multi-billing from 0.55 | Multi-billing from 20 |
- qwen3-livetranslate-flash-realtime-2025-09-22qwen3-livetranslate-flash-realtime-2025-09-22 | Multi-billing from 1.3 | Multi-billing from 10 | qwen3.5-livetranslate-flash-realtime | Multi-billing from 0.55 | Multi-billing from
- qwen3-maxQwen3 Series Mainline Models:qwen3.6-max-preview、qwen3-max-preview、qwen3-max、qwen3-vl-flash、qwen3-coder-plus
- qwen3-max-previewQwen3 Series Mainline Models:qwen3.6-max-preview、qwen3-max-preview、qwen3-max、qwen3-vl-flash、qwen3-coder-plus
- qwen3-next-80b-a3b-instruct-maasqwen3-next-80b-a3b-instruct-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/qwen/model-garden/qwen3-next> Self-deploy Qwen3 Ne
- qwen3-next-80b-a3b-thinking-maasqwen3-next-80b-a3b-thinking-maas | July 21, 2026 | October 21, 2026 | <https://console.cloud.google.com/agent-platform/publishers/qwen/model-garden/qwen3-next> Self-deploy Qwen3 Ne
- qwen3-omni-30b-a3b-captionerqwen3-omni-30b-a3b-captioner | 3.81 | 3.06 | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen3-omni-flash-2025-09-15qwen3-omni-flash-2025-09-15 | Multi-billing from 0.43 | Multi-billing from 1.66 | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen3-omni-flash-2025-12-01qwen3-omni-flash-2025-12-01 | / | / | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen3-omni-flash-realtimeqwen3-omni-flash-realtime | Multi-billing from 0.52 | Multi-billing from 1.99 | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-omni-flash-realtime-2025-12-01qwen3-omni-flash-realtime-2025-12-01 | / | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-flash-2025-09-18qwen3-tts-flash-2025-09-18 | $0.1/10K chars | / | cosyvoice-v3.5-plus | $0.22/10K chars | / |
- qwen3-tts-flash-2025-11-27qwen3-tts-flash-2025-11-27 | $0.1/10K chars | / | cosyvoice-v3.5-plus | $0.22/10K chars | / |
- qwen3-tts-flash-realtime-2025-09-18qwen3-tts-flash-realtime-2025-09-18 | $0.13/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-flash-realtime-2025-11-27qwen3-tts-flash-realtime-2025-11-27 | $0.13/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-instruct-flashqwen3-tts-instruct-flash | $0.115/10K chars | / | cosyvoice-v3.5-plus | $0.22/10K chars | / |
- qwen3-tts-instruct-flash-2026-01-26qwen3-tts-instruct-flash-2026-01-26 | $0.115/10K chars | / | cosyvoice-v3.5-plus | $0.22/10K chars | / |
- qwen3-tts-instruct-flash-realtimeqwen3-tts-instruct-flash-realtime | $0.143/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-instruct-flash-realtime-2026-01-22qwen3-tts-instruct-flash-realtime-2026-01-22 | $0.143/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-vc-2026-01-22qwen3-tts-vc-2026-01-22 | $0.115/10K chars | / | cosyvoice-v3.5-plus | $0.22/10K chars | / |
- qwen3-tts-vc-realtime-2025-11-27qwen3-tts-vc-realtime-2025-11-27 | $0.13/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-vc-realtime-2026-01-15qwen3-tts-vc-realtime-2026-01-15 | $0.13/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-vd-2026-01-26qwen3-tts-vd-2026-01-26 | $0.115/10K chars | / | cosyvoice-v3.5-plus | $0.22/10K chars | / |
- qwen3-tts-vd-realtime-2025-12-16qwen3-tts-vd-realtime-2025-12-16 | $0.13/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-tts-vd-realtime-2026-01-15qwen3-tts-vd-realtime-2026-01-15 | $0.13/10K chars | / | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen3-vl-flashQwen3 Series Mainline Models:qwen3.6-max-preview、qwen3-max-preview、qwen3-max、qwen3-vl-flash、qwen3-coder-plus
- qwen3-vl-flash-2025-10-15-usqwen3-vl-flash-2025-10-15-us | Tiered pricing from 0.05 | Tiered pricing from 0.4 | qwen3.6-flash-us | Tiered pricing from 0.25 | Tiered pricing from 1.5 |
- qwen3-vl-flash-2026-01-22-usqwen3-vl-flash-2026-01-22-us | Tiered pricing from 0.05 | Tiered pricing from 0.4 | qwen3.6-flash-us | Tiered pricing from 0.25 | Tiered pricing from 1.5 |
- qwen3-vl-flash-usqwen3-vl-flash-us | Tiered pricing from 0.05 | Tiered pricing from 0.4 | qwen3.6-flash-us | Tiered pricing from 0.25 | Tiered pricing from 1.5 |
- qwen3-vl-plus-2025-09-23qwen3-vl-plus-2025-09-23 | Tiered pricing from 0.2 | Tiered pricing from 1.6 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen3-vl-plus-2025-12-19qwen3-vl-plus-2025-12-19 | Tiered pricing from 0.2 | Tiered pricing from 1.6 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen3.6-max-previewQwen3 Series Mainline Models:qwen3.6-max-preview、qwen3-max-preview、qwen3-max、qwen3-vl-flash、qwen3-coder-plus
- qwen-coder-plusQwen Series Legacy Mainline Models:qwen-turbo、qwen-turbo-realtime、qwen-vl-max、qwen-vl-plus、qwq-plus、qvq-max、qvq-plus、qwen-math-turbo、qwen-coder-turbo、qwen-coder-plus
- qwen-coder-turboQwen Series Legacy Mainline Models:qwen-turbo、qwen-turbo-realtime、qwen-vl-max、qwen-vl-plus、qwq-plus、qvq-max、qvq-plus、qwen-math-turbo、qwen-coder-turbo、qwen-coder-plus
- qwen-flash-2025-07-28-usqwen-flash-2025-07-28-us | Tiered pricing from 0.05 | Tiered pricing from 0.4 | qwen3.6-flash-us | Tiered pricing from 0.25 | Tiered pricing from 1.5 |
- qwen-flash-usqwen-flash-us | Tiered pricing from 0.05 | Tiered pricing from 0.4 | qwen3.6-flash-us | Tiered pricing from 0.25 | Tiered pricing from 1.5 |
- qwen-imageqwen-image | $0.035/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-editqwen-image-edit | $0.045/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-edit-maxqwen-image-edit-max | $0.075/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-edit-max-2026-01-16qwen-image-edit-max-2026-01-16 | $0.075/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-edit-plusqwen-image-edit-plus | $0.03/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-edit-plus-2025-10-30qwen-image-edit-plus-2025-10-30 | $0.03/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-edit-plus-2025-12-15qwen-image-edit-plus-2025-12-15 | $0.03/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-maxqwen-image-max | $0.075/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-max-2025-12-30qwen-image-max-2025-12-30 | $0.075/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-plusqwen-image-plus | $0.03/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-image-plus-2026-01-09qwen-image-plus-2026-01-09 | $0.03/image | / | qwen-image-3.0 | / | $0.03/image |
- qwen-long-2025-01-25qwen-long-2025-01-25 | 0.072 | 0.287 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-long-latestqwen-long-latest | 0.072 | 0.287 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-math-plusqwen-math-plus | 0.574 | 1.721 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-math-plus-0816qwen-math-plus-0816 | 0.574 | 1.721 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-math-plus-0919qwen-math-plus-0919 | 0.574 | 1.721 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-math-plus-latestqwen-math-plus-latest | 0.574 | 1.721 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-math-turboQwen Series Legacy Mainline Models:qwen-turbo、qwen-turbo-realtime、qwen-vl-max、qwen-vl-plus、qwq-plus、qvq-max、qvq-plus、qwen-math-turbo、qwen-coder-turbo、qwen-coder-plus
- qwen-mt-lite-usqwen-mt-lite-us | 0.12 $ | 0.36 | qwen3.6-flash-us | Tiered pricing from 0.25 | Tiered pricing from 1.5 |
- qwen-mt-turboqwen-mt-turbo | 0.16 | 0.49 | qwen-mt-flash | 0.16 | 0.49 |
- qwen-omni-turboqwen-omni-turbo | Multi-billing from 0.07 | Multi-billing from 0.27 | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen-omni-turbo-2025-01-19qwen-omni-turbo-2025-01-19 | Multi-billing from 0.058 | Multi-billing from 0.23 | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen-omni-turbo-2025-03-26qwen-omni-turbo-2025-03-26 | Multi-billing from 0.07 | Multi-billing from 0.27 | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen-omni-turbo-latestqwen-omni-turbo-latest | Multi-billing from 0.07 | Multi-billing from 0.27 | qwen3.5-omni-plus | Multi-billing from 1.4 | Multi-billing from 8.3 |
- qwen-omni-turbo-realtimeqwen-omni-turbo-realtime | Multi-billing from 0.27 | Multi-billing from 1.07 | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen-omni-turbo-realtime-2025-05-08qwen-omni-turbo-realtime-2025-05-08 | Multi-billing from 0.27 | Multi-billing from 1.07 | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen-omni-turbo-realtime-latestqwen-omni-turbo-realtime-latest | Multi-billing from 0.27 | Multi-billing from 1.07 | qwen3.5-omni-plus-realtime | Multi-billing from 2.1 | Multi-billing from 12.4 |
- qwen-plus-0112qwen-plus-0112 | Tiered pricing from 0.115 | Tiered pricing from 0.287 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-plus-1220qwen-plus-1220 | Tiered pricing from 0.115 | Tiered pricing from 0.287 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-plus-2025-07-28qwen-plus-2025-07-28 | Tiered pricing from 0.4 | Tiered pricing from 1.2 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-plus-2025-09-11qwen-plus-2025-09-11 | Tiered pricing from 0.115 | Tiered pricing from 0.287 | qwen3.7-plus | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-plus-2025-12-01-usqwen-plus-2025-12-01-us | Tiered pricing from 0.4 | Tiered pricing from 1.2 | qwen3.7-plus-us | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- qwen-plus-usqwen-plus-us | Tiered pricing from 0.4 | Tiered pricing from 1.2 | qwen3.7-plus-us | Tiered pricing from 0.4 | Tiered pricing from 1.6 |
- Qwen/Qwen1.5-0.5B| 2024-08-22 | `Qwen/Qwen1.5-0.5B` | No |
- Qwen/Qwen1.5-0.5B-Chat| 2024-08-22 | `Qwen/Qwen1.5-0.5B-Chat` | No |
- Qwen/Qwen1.5-1.8B| 2024-08-22 | `Qwen/Qwen1.5-1.8B` | No |
- Qwen/Qwen1.5-1.8B-Chat| 2024-08-22 | `Qwen/Qwen1.5-1.8B-Chat` | No |
- Qwen/Qwen1.5-4B| 2024-08-22 | `Qwen/Qwen1.5-4B` | No |
- Qwen/Qwen1.5-4B-Chat| 2024-08-22 | `Qwen/Qwen1.5-4B-Chat` | No |
- Qwen/Qwen1.5-7B| 2024-08-22 | `Qwen/Qwen1.5-7B` | No |
- Qwen/Qwen1.5-7B-Chat| 2024-08-22 | `Qwen/Qwen1.5-7B-Chat` | No |
- Qwen/Qwen1.5-14B| 2024-08-22 | `Qwen/Qwen1.5-14B` | No |
- Qwen/Qwen1.5-14B-Chat| 2024-08-22 | `Qwen/Qwen1.5-14B-Chat` | No |
- Qwen/Qwen1.5-32B| 2024-08-22 | `Qwen/Qwen1.5-32B` | No |
- Qwen/Qwen1.5-32B-Chat| 2024-08-22 | `Qwen/Qwen1.5-32B-Chat` | No |
- Qwen/Qwen1.5-72B| 2024-08-22 | `Qwen/Qwen1.5-72B` | No |
- Qwen/Qwen1.5-72B-Chat| 2024-10-29 | `Qwen/Qwen1.5-72B-Chat` | No |
- Qwen/Qwen1.5-110B-Chat| 2024-10-29 | `Qwen/Qwen1.5-110B-Chat` | No |
- qwen-qwen2-5-14b-instruct-lora| 2026-02-06 | `qwen-qwen2-5-14b-instruct-lora` | No |
- Qwen/Qwen2-72B-Instruct| 2025-08-28 | `Qwen/Qwen2-72B-Instruct` | No |
- Qwen/Qwen2-VL-72B-Instruct| 2025-08-28 | `Qwen/Qwen2-VL-72B-Instruct` | No |
- Qwen/Qwen2.5-7B| 2026-08-19 | `Qwen/Qwen2.5-7B` | No |
- Qwen/Qwen2.5-14B| 2025-08-28 | `Qwen/Qwen2.5-14B` | No |
- Qwen/Qwen2.5-72B-Instruct-Turbo| 2026-02-06 | `Qwen/Qwen2.5-72B-Instruct-Turbo` | No |
- Qwen/Qwen2.5-VL-72B-Instruct| 2026-01-05 | `Qwen/Qwen2.5-VL-72B-Instruct` | No |
- Qwen/Qwen3-235B-A22B-fp8-tput| 2026-02-06 | `Qwen/Qwen3-235B-A22B-fp8-tput` | No |
- Qwen/Qwen3-235B-A22B-Instruct-2507-tput| 2026-07-10 | `Qwen/Qwen3-235B-A22B-Instruct-2507-tput` | Yes |
- Qwen/Qwen3-235B-A22B-Thinking-2507-fast| `Qwen/Qwen3-235B-A22B-Thinking-2507-fast` |
- Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8| 2026-06-04 | `Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8` | Yes |
- Qwen/Qwen3-Coder-Next-FP8| 2026-05-14 | `Qwen/Qwen3-Coder-Next-FP8` | Yes |
- Qwen/Qwen3-Next-80B-A3B-Instruct| 2026-04-02 | `Qwen/Qwen3-Next-80B-A3B-Instruct` | Yes |
- Qwen/Qwen3-Next-80B-A3B-Thinking| 2026-02-25 | `Qwen/Qwen3-Next-80B-A3B-Thinking` | No |
- Qwen/Qwen3-Next-80B-A3B-Thinking-fast| `Qwen/Qwen3-Next-80B-A3B-Thinking-fast` |
- Qwen/Qwen3-VL-8B-Instruct| 2026-04-16 | `Qwen/Qwen3-VL-8B-Instruct` | Yes |
- Qwen/Qwen3-VL-32B-Instruct| 2026-02-25 | `Qwen/Qwen3-VL-32B-Instruct` | No |
- Qwen/Qwen3.5-397B-A17B| 2026-06-29 | `Qwen/Qwen3.5-397B-A17B` | Yes |
- Qwen/Qwen3.5-397B-A17B-fast| `Qwen/Qwen3.5-397B-A17B-fast` |
- qwen-qwq-32b| qwen-qwq-32b | 07/14/25 | qwen/qwen3-32b |
- Qwen/QwQ-32B-Preview| 2025-03-25 | `Qwen/QwQ-32B-Preview` | No |
- qwen-ttsAudio Series Mainline Models:qwen-tts、qwen-tts-realtime、qwen-voice-enrollment、qwen-voice-design、gummy-realtime-v1
- qwen-tts-2025-04-10Qwen Audio Series Snapshots:qwen-tts-latest、qwen-tts-2025-05-22、qwen-tts-2025-04-10、qwen-tts-realtime-latest、qwen-tts-realtime-2025-07-15
← records 1–400 · All of AI model deprecation and retirement dates by provider · records 801 on →