Z.ai's 320B GLM-5.3-Flash goes open-weight, nearing Claude on coding benchmarks

Z.ai released GLM-5.3-Flash under an MIT licence — 320 billion parameters, 18 billion active, with a one-million-token context. The lab says it approaches Anthropic's Claude Opus 4.8 on coding and agentic tests, scoring 84.3 to Opus's 85.0 on Terminal-Bench 2.1. API pricing is $0.15 per million input tokens and $0.50 per million output, halved promotionally until September 9.
4 sources
- 01Hugging Face — zai-org/GLM-5.3-Flash model cardhuggingface.co
- 02The Decoder — GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidiathe-decoder.com
- 03Artificial Analysis — GLM-5.3-Flash intelligence, performance and price analysisartificialanalysis.ai
- 04MarkTechPost — Z.ai releases GLM-5.3-Flash: a 320B-A18B natively multimodal MoE with a 1M-token contextmarktechpost.com