What Is Qwen 3.8
OpenRouter ·

Qwen 3.8 is Alibaba’s model generation released in August 2026. Four models share the name, and they differ in ways that matter before you pick one. You can download Qwen3.8 27B under Apache 2.0 and Qwen3.8 2.4T A95B under the Qwen3.8-Max License. Qwen3.8 Flash has no repository of its own, but Qwen3.8-Flash-Next, the model it is based on, is downloadable under the Qwen Community License 1.0. Qwen3.8 Max has no public weights.
Alibaba announced Qwen3.8-Max as the hosted model of the generation. The QwenLM/Qwen3.8 release notes record the Qwen3.8-2.4T-A95B weights arriving on Hugging Face on 12 August 2026 and the 27B weights on 14 August 2026. The 2.4T model card describes this as the first open release of a Qwen-Max-class model.
The model you can download is not the same as the model you can call. Alibaba’s model card describes Qwen3.8-Max as the hosted version of Qwen3.8-2.4T-A95B with added features, including vision input, non-thinking mode, a 1,000,000-token default context, and built-in tools. The same relationship holds for Flash. Alibaba describes Qwen3.8-Flash as the hosted version of Qwen3.8-Flash-Next with a 1,000,000-token default context and built-in tools.
We route all four. This post covers what each model is, which ones you can download, what the licenses require, and what each one costs to run.
The four Qwen 3.8 models
Prices are US dollars per million tokens, prompt then completion, from our model catalog on 11 September 2026. The 27B price is a range because 14 providers serve it at different rates. Six of the seven 2.4T providers charge $2 / $6, and Venice charges $2.50 / $7.50.
| Model | Weights | License | Input | Context, native / largest hosted | Price |
|---|---|---|---|---|---|
| qwen/qwen3.8-max-0902 | API only | Proprietary | Text, image, video | Not published / 1,000,000 | $2 / $6 |
| qwen/qwen3.8-2.4t-a95b | Hugging Face | Qwen3.8-Max License | Text | 262,144 / 1,048,576 | $2 / $6 |
| qwen/qwen3.8-27b | Hugging Face | Apache 2.0 | Text, image, video | 262,144 / 1,000,000 | $0.15 to $0.45 / $2.00 to $3.20 |
| qwen/qwen3.8-flash | Hugging Face, as Qwen3.8-Flash-Next | Qwen Community License 1.0, on Qwen3.8-Flash-Next | Text, image, video | 262,144 on Qwen3.8-Flash-Next / 1,000,000 | $0.15 / $0.47 |
The model ID qwen/qwen3.8-max is an alias. It currently resolves to qwen/qwen3.8-max-0902, the snapshot dated 2 September 2026. Use the dated ID if you want the request to keep hitting the same snapshot after Alibaba publishes a new one.
The context column shows two numbers because they measure different things. The first is the window the model was trained on, which Alibaba publishes for the open-weight models and does not publish for Max. The second is the largest window any provider in our catalog offers, and it is larger because the window can be extended past the trained length. The Qwen3.8-27B model card describes extending the 27B to 1,000,000 tokens with YaRN, a RoPE scaling technique, and notes that static YaRN can reduce performance on shorter inputs. Whether a provider extends the window, and how, is that provider’s decision, so the ceiling you get depends on where we route your request.
If you use prompt caching, the cache rates differ by model and provider. Max charges $0.25 per million tokens to read from the cache and $2.50 to write to it. Flash charges $0.016 to read and, on Alibaba, $0.20 to write. The 2.4T charges $0.25 to read on most providers and lists a write rate only on Alibaba, at $2.50. The 27B cache read rate ranges from $0.032 to $0.18 across the providers that offer it, and only Alibaba lists a write rate, at $0.53.
Is Qwen 3.8 open source?
Partly, and the answer depends on which model you mean.
You cannot download Max. It has no Hugging Face repository. Max runs on Alibaba’s servers and you reach it through an API.
You can download the 2.4T, and its license adds conditions at scale. Qwen/Qwen3.8-2.4T-A95B is the instruction-tuned model with 2.4 trillion total parameters and 95 billion active per token. It does not use Apache 2.0. Qwen wrote the Qwen3.8-Max License for it. The Hugging Face metadata records the license as other with the name qwen3.8-max.
The license grants the rights to use, copy, modify, distribute, sublicense, sell, deploy, host, fine-tune, and create derivative works, subject to two conditions.
- Attribution. If a commercial product or service built on the model has more than 100,000,000 monthly active users or more than US$20,000,000 in monthly revenue, the model name must be prominently displayed in its user interface.
- Separate license for large model-as-a-service businesses. If you or your affiliates run a Model as a Service or AI Work Assistant business, and your aggregate revenue exceeds US$50,000,000 in any consecutive twelve months, you need a separate license from Qwen before any commercial use. This condition does not apply to internal use that does not make the model, its outputs, or its capabilities available to a third party.
The license defines Model as a Service as giving a third party access to model inference or fine-tuning in a way that lets them control the inputs, parameters, or training data. It excludes relaying requests to models hosted by others. It defines AI Work Assistant as an independent product primarily designed for AI-assisted coding or office productivity.
You can download the 27B under Apache 2.0. Qwen/Qwen3.8-27B records apache-2.0 in its Hugging Face metadata. Apache 2.0 is on the Open Source Initiative approved list and carries no revenue or user-count conditions.
Flash has open weights under a different name and a third license. There is no repository for Qwen3.8 Flash. There is one for Qwen3.8-Flash-Next, which Alibaba describes as an experimental preview of the architecture that will underpin Qwen4. Its license is the Qwen Community License 1.0. It has the same attribution condition as the Qwen3.8-Max License, and it requires a separate license from Qwen for any Model as a Service or AI Work Assistant business before commercial use, with no revenue threshold.
Two of the three open-weight models therefore ship under licenses that Qwen wrote and that are not OSI-approved. Open weights means you can download and inspect the checkpoint. It does not mean the license is open source.
Qwen 3 is an earlier generation
Qwen 3 is not the version before Qwen 3.8. Qwen 3 was released in April 2025, and Qwen 3.5, Qwen 3.6, and Qwen 3.7 shipped between the two. Qwen 3.8 follows Qwen 3.7. A license term that applied to a Qwen 3 model does not tell you anything about the four models in this post.
Can you run Qwen 3.8 locally?
You can run the 27B on your own hardware. The 2.4T requires multi-GPU infrastructure.

Figure 1. Downloadable weights against input modality. The 2.4T is the only model in the family that reads text alone. No Qwen 3.8 model is API only and text only.
The 2.4T is a mixture-of-experts model. Only 95 billion parameters are active on any one token, but all 2.4 trillion parameters have to be loaded, because the model selects a different set of experts for each token. The number your hardware has to hold is 2.4 trillion. The README in the QwenLM/Qwen3.8 repository shows serving commands for the 27B only.
The 27B is a dense model with 27 billion parameters. At bf16, two bytes per parameter, the weights occupy about 54 GB, and about half that at fp8. Alibaba’s serving examples run it across four GPUs at the full 262,144-token native window using SGLang, vLLM, or TokenSpeed.
vllm serve Qwen/Qwen3.8-27B --port 8000 --tensor-parallel-size 4 --max-model-len 262144 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder
Serving it at 1,000,000 tokens requires the YaRN override described in the model card, and the model card advises changing the RoPE parameters only when you need the longer context.
The open weights of the large model only read text. The hosted Max accepts images and video, and so does the open-weight 27B. The open-weight 2.4T does not. When we sent an image to qwen/qwen3.8-2.4t-a95b on 11 September 2026, the request failed with HTTP 404 and the message No endpoints found that support image input. The same request to qwen/qwen3.8-27b and qwen/qwen3.8-flash succeeded.
On text, the open weights score close to Max.
| Model | Intelligence Index | Coding Index | Agentic Index |
|---|---|---|---|
| Qwen3.8 Max (0902) | 40.3 | 71.8 | 49.6 |
| Qwen3.8 2.4T A95B | 40.0 | 71.9 | 50.4 |
| Qwen3.8 27B (xhigh) | 33.9 | 68.1 | 46.5 |
These scores come from Artificial Analysis, an independent benchmarking organization, and we publish them on each model’s page. We retrieved them on 11 September 2026. Artificial Analysis measures a dated snapshot of each model, so the Max row is the 2 September 2026 snapshot and the 27B row was measured at xhigh reasoning effort. The scores will move if Artificial Analysis remeasures a model or if you run the model at a different reasoning setting.
The 2.4T is 0.3 points behind Max on the Intelligence Index, 0.1 ahead on the Coding Index, and 0.8 ahead on the Agentic Index. The 27B is behind Max by 6.4, 3.7, and 3.1 points respectively.
If you want Max-class weights to self-host a workload that reads screenshots, the 27B is the model that reads images.
Choosing a model
Use these facts to narrow the choice.
- If you need image or video input and want to hold the weights, the 27B is the only open-weight Qwen 3.8 model in our catalog that accepts them. It is Apache 2.0.
- If you need image or video input and want the highest Artificial Analysis scores in the family, Max scores highest on the Intelligence Index.
- If you want the lowest per-token price, Flash is $0.15 per million prompt tokens and $0.47 per million completion tokens, and it accepts images and video. You cannot download Qwen3.8 Flash itself, only Qwen3.8-Flash-Next.
- If you work with text and want to hold the weights, the 2.4T is the open-weight variant of Max. Check the license thresholds above against your own numbers.
- If you work with text and do not want to host anything, Max and the 2.4T cost the same through us on six of the 2.4T’s seven providers, so the difference is whether you need image input.
How to call the Qwen 3.8 models
Change the model ID and the rest of the request stays the same. This example uses the TypeScript SDK.
import { OpenRouter } from '@openrouter/sdk';
const openRouter = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY ?? '',
});
const result = await openRouter.chat.send({
chatRequest: {
model: 'qwen/qwen3.8-27b',
messages: [
{
role: 'user',
content: 'Summarize the trade-offs between fp8 and bf16 inference in two sentences.',
},
],
stream: false,
},
});
if (result instanceof ReadableStream) {
throw new Error('Expected a non-streaming response');
}
console.log(result.choices[0].message.content);
To send an image, pass an array of content parts instead of a string. See image inputs for the request shape.
const result = await openRouter.chat.send({
chatRequest: {
model: 'qwen/qwen3.8-27b',
messages: [
{
role: 'user',
content: [
{ type: 'text', text: 'What changed in this screenshot?' },
{ type: 'image_url', imageUrl: { url: screenshotUrl } },
],
},
],
stream: false,
},
});
if (result instanceof ReadableStream) {
throw new Error('Expected a non-streaming response');
}
console.log(result.choices[0].message.content);
The same request to qwen/qwen3.8-2.4t-a95b fails, because none of its providers accept image input.
What affects the cost
Two things change your bill beyond the listed rates. The models reason before they answer, and you pay for the reasoning tokens. The provider we route you to sets the price.
Reasoning is on by default
All four models reason before answering by default. Reasoning tokens are billed at the completion rate, which is the higher of the two rates. See reasoning tokens for how the reasoning parameter works.
The share of your output that is reasoning can be large. We asked the 27B through Alibaba, three times, to explain in two sentences why fp8 quantization changes a model’s output. Reasoning tokens were 74% of the completion tokens on each run, at 128 of 174, 171 of 231, and 181 of 243. On an image description request to the same model, 87 of 121 completion tokens were reasoning.
Whether you can turn reasoning off depends on the model and the endpoint. On 11 September 2026 we sent reasoning: { effort: 'none' } to each model.
| Model | Result of effort: 'none' |
|---|---|
| qwen/qwen3.8-max-0902 | HTTP 400, Reasoning is mandatory for this endpoint and cannot be disabled. |
| qwen/qwen3.8-2.4t-a95b | HTTP 400, same message |
| qwen/qwen3.8-27b | HTTP 200, zero reasoning tokens |
| qwen/qwen3.8-flash | HTTP 200, zero reasoning tokens |
Alibaba’s model card lists non-thinking mode as a feature of the hosted Max, but the Alibaba endpoint we route to rejects the request. The 27B result covers the routes we tested, which were Alibaba and Chutes. The Flash result covers Makora. Other providers can behave differently, so check the response usage field when you rely on reasoning being off.
To disable reasoning on the 27B, add one field.
const result = await openRouter.chat.send({
chatRequest: {
model: 'qwen/qwen3.8-27b',
messages: [{ role: 'user', content: 'Reply with the single word OK.' }],
reasoning: { effort: 'none' },
stream: false,
},
});
We ran the fp8 question three times with reasoning on and three times with it off, both through Alibaba. With reasoning on, the runs averaged 216 completion tokens and $0.00058 per call. With it off, they averaged 61 completion tokens and $0.00017 per call. That is roughly 70% less, whether you count tokens or dollars.
Max and the 2.4T still accept the other effort values. Requests with effort set to max, high, and minimal returned HTTP 200 on all four models in the same test run, so you can control how much they reason. You cannot stop them on the current endpoints.
Provider pricing and context limits
Max has one provider in our catalog, Alibaba. Flash has two, Alibaba and Makora. The 2.4T has seven. The 27B has 14, and they do not charge the same. On 11 September 2026, 27B prompt prices ran from $0.15 to $0.45 per million tokens and completion prices from $2.00 to $3.20. The current list is on the model page, and it changes, so treat the figures here as a snapshot.
The 27B providers do not all serve the same context either. Two of the 14 offer the 1,000,000-token window, one stops at 65,536, and the rest stop at 262,144. A request carrying 400,000 tokens of input has to be routed to a provider that accepts it. Providers also quantize differently. Most of the 14 list fp8, one lists fp4, and several list their quantization as unknown. The weights behind one model ID are not identical from provider to provider, and the Artificial Analysis scores above describe the model as measured by Artificial Analysis, not any particular provider’s copy.
No single figure describes what the 27B costs. Pin a provider if the number matters to you. You can pick a provider with the order field and switch off fallbacks. Provider routing has the full set of controls.
Reasoning tokens are not itemized everywhere. In our test on 11 September 2026, Chutes returned the model’s reasoning text but reported zero reasoning tokens in usage, three runs out of three. Alibaba, CoreWeave, and Novita reported non-zero reasoning token counts on every run. You pay for those tokens either way, so a dashboard that sums reasoning_tokens across providers will undercount.
Going direct to the 27B providers would mean 14 accounts and 14 integrations. Through us it is one model ID in the request, and when a provider slows down or returns an error, we route the request to another one.
FAQ
Is Qwen 3.8 open source?
Partly. Qwen3.8 27B is published under Apache 2.0. Qwen3.8 2.4T A95B is published under the Qwen3.8-Max License, which Qwen wrote and which adds conditions at scale. Qwen3.8 Flash has no repository of its own, but the model it is based on, Qwen3.8-Flash-Next, is published under the Qwen Community License 1.0. Alibaba describes the -Next model as an experimental preview of the architecture planned for Qwen4, and Qwen3.8 Flash as its hosted version, so the suffix marks a preview release rather than a second Flash model. Qwen3.8 Max has no public weights.
Is Qwen 3.8 free to use?
You can download Qwen3.8 27B, Qwen3.8 2.4T A95B, and Qwen3.8-Flash-Next from Hugging Face and run them on your own hardware without a per-token fee. Qwen3.8 Max is available only through an API. Calling any of the four through OpenRouter is billed per token.
Can I use Qwen 3.8 commercially?
Qwen3.8 27B is Apache 2.0. Qwen3.8 2.4T A95B allows commercial use with two conditions. If your product exceeds 100,000,000 monthly active users or US$20,000,000 in monthly revenue, you must display the model name in its interface. If you or your affiliates run a Model as a Service or AI Work Assistant business with aggregate revenue over US$50,000,000 in any consecutive twelve months, you need a separate license from Qwen before commercial use. The Qwen Community License 1.0 on Qwen3.8-Flash-Next has the same attribution condition and requires the separate license for any Model as a Service or AI Work Assistant business, with no revenue threshold. The conditions are set out in the Qwen3.8-Max License and the Qwen Community License 1.0 files.
Do the open Qwen 3.8 weights accept images?
Qwen3.8 27B accepts text, images, and video. Qwen3.8-Flash-Next is a causal language model with a vision encoder. Qwen3.8 2.4T A95B is text only. On OpenRouter, an image request to qwen/qwen3.8-2.4t-a95b returns a 404 with the message No endpoints found that support image input.
Was Qwen 3 the release before Qwen 3.8?
No. Qwen 3 was released in April 2025. Qwen 3.5, Qwen 3.6, and Qwen 3.7 were released between Qwen 3 and Qwen 3.8. Qwen 3.8 follows Qwen 3.7.
Can I turn off reasoning on Qwen 3.8?
It depends on the model and endpoint. On OpenRouter, qwen/qwen3.8-max-0902 and qwen/qwen3.8-2.4t-a95b rejected reasoning: { effort: 'none' } with a 400 error stating that reasoning is mandatory for the endpoint. qwen/qwen3.8-27b and qwen/qwen3.8-flash accepted it on the routes we tested and returned zero reasoning tokens. In our test, disabling reasoning on the 27B cut the cost of the same prompt by roughly 70%.
References
- Qwen, Qwen3.8-Max announcement.
- Hugging Face, Qwen/Qwen3.8-2.4T-A95B and its LICENSE.
- Hugging Face, Qwen/Qwen3.8-27B.
- Hugging Face, Qwen/Qwen3.8-Flash-Next and its LICENSE.
- GitHub, QwenLM/Qwen3.8.
- OpenRouter, Qwen3.8 Max, Qwen3.8 2.4T A95B, Qwen3.8 27B, and Qwen3.8 Flash model pages.
- OpenRouter, Reasoning tokens, Provider routing, Prompt caching, and Image inputs.
- Artificial Analysis, independent model benchmarks, retrieved 11 September 2026.