TLDR: Z.ai’s GLM-5.2 is the strongest open-weight, MIT-licensed model for coding and agentic work, pairing a frontier-class long-context engine with a US-hosted, zero-retention deployment path that suits regulated teams.
What Z.ai shipped
Z.ai, the lab built from Zhipu AI, released GLM-5.2 in mid-June 2026 as its flagship model for long-horizon tasks. The release is an open-weight Mixture-of-Experts (MoE) system carrying 753 billion parameters, and it activates a small share of them on each token, which keeps the running cost close to a far smaller model. Two design choices separate it from the open-weight field. The first is a solid one-million-token context window, five times the 200,000-token limit of GLM-5.1, engineered to stay reliable across long coding-agent sessions. The second is a permissive MIT licence that stays open across regions, which grants unrestricted commercial use, local deployment, and the freedom to fine-tune.
The model ships with selectable thinking effort. High mode answers everyday tasks quickly, and Max mode spends extra computation on harder problems, so an engineering team trades latency for capability per task. Z.ai placed the weights on Hugging Face and ModelScope, and the model serves through transformers, vLLM, SGLang, and other standard frameworks. The headline capability the lab claims is end-to-end delivery: project-level engineering context, long-running tasks that complete reliably, consistent adherence to engineering standards, and a full path from requirements to multi-platform deployment inside a single task.
Unlock Premium
RESERVED TO MEMBERS