← Blog
·4 min read

MiMo-V2.6-Pro, GLM-5.3 Prime, Qwen3.8 Max Prime: who is actually open-weight this week

Two launches sharing the same suffix — "Prime" — dominated this week's headlines, but neither is a new weight release. The real shift on the open-weight leaderboard comes from a name few expected: Xiaomi.

Open modelsNewsLicenses
MiMo-V2.6-Pro, GLM-5.3 Prime, Qwen3.8 Max Prime: who is actually open-weight this week

This week's headlines were all about two launches sharing the same suffix — "Prime" — one from Z.ai, one from Alibaba. Neither, though, is a release of new weights: both are faster API lanes for models that already existed. The real open-weight news of the week, which got far less attention, is that a smartphone maker just leapfrogged every well-known Chinese AI lab to the top of the open-weight leaderboard.

MiMo-V2.6-Pro: Xiaomi takes the top of the open-weight leaderboard

On September 21, 2026, Xiaomi published MiMo-V2.6-Pro on Hugging Face, a sparse Mixture-of-Experts model with 1.02 trillion total parameters, only 42 billion of which are active per token, under a plain MIT license — free commercial use, no thresholds, no hidden clauses. According to Artificial Analysis, the model scores 46 on the Intelligence Index, the highest score any open-weight model has ever reached: it beats both Z.ai's GLM-5.3 and Moonshot's Kimi K3, both stuck at 44, and ranks sixth overall once closed models are counted too. The detail that surprised observers most is the reported training cost — around $3 million — far lower than what's normally associated with a model of this size, with a quality-to-price ratio that Artificial Analysis places on the Pareto frontier between intelligence and cost per token. Xiaomi also released MiMo-V2.6-Flash alongside it, a cheaper variant built for higher volumes.

GLM-5.3 Prime and Qwen3.8 Max Prime: two launches that aren't new weights

On September 23, 2026, Z.ai launched GLM-5.3 Prime, a high-speed variant of GLM-5.3 that inherits all of the base model's capabilities but delivers 1.5-2x the throughput through inference acceleration, with context up to 1 million tokens and reasoning that is always on — it cannot be disabled, with three effort levels (low, high, max, the latter being the default). It isn't a new checkpoint: it's an API product, priced at $2.80 per million input tokens and $8.80 per million output tokens, aimed at coding and long-horizon agentic orchestration workloads.

The same day, following the September 22 announcement at the Yunqi Conference in Hangzhou, Alibaba activated Qwen3.8 Max Prime, a faster lane for the same Qwen3.8-Max — the 2.4-trillion-parameter sparse MoE model that on August 13 became the first "Max"-class model in the Qwen family to be released open-weight, under Apache 2.0. Qwen3.8 Max Prime uses the same weights, the same 1-million-token context window, and the same tooling as the base model: only speed and price change, with price roughly doubling versus the standard version ($4 vs. $2 per million input tokens, $12 vs. $6 output). Here too, no new weights to download.

GLM-5.3's license isn't MIT: the clause that matters if you make a lot of revenue

It's worth circling back to the license of the GLM-5.3 base model (753 billion parameters, published on Hugging Face on August 28), because it's a case that shows once again how risky it is to trust the label. Z.ai didn't use MIT or Apache 2.0, but a bespoke document — the "GLM-5.3 License" — that in substance grants the same rights as an MIT license (use, copy, modify, distribute, sublicense, sell, explicitly covering weights, parameters, configuration files, and training and inference code), with one added clause that makes the difference: Model-as-a-Service operators above a given revenue threshold must undergo a security review. Two days earlier, the same lab had released GLM-5.3-Flash instead — 320 billion total parameters, 18 billion active, the first natively multimodal model (text, image, video) in the GLM-5 series — under a plain, unmodified MIT license, without that clause. Same lab, same week, two different licenses for two models in the same family.

What actually matters if you're evaluating an AI architecture today

It's exactly the work we do every week with our clients: telling a genuine weights release apart from a new API lane carrying the same model name, reading the exact license of each release — not the previous one's — and building an architecture where the right model runs where it needs to run, with your data staying yours, regardless of which lab trained the model or which leaderboard it leads this week.

Sources

Want to talk about it applied to your case?

Book a call →More articles
Keep reading

Open models vs. closed models: what actually changes for your sensitive data

A closed model and an open-weight model aren't two variants of the same product: they change who sees your data and who controls the system. Here's how to decide, case by case.

FADP, GDPR, and artificial intelligence: the compliance checklist for business decision-makers

Adopting an AI tool without checking where the data ends up is the fastest way to turn a productivity gain into a compliance problem. Here's what to check first.