GLM-5.3-Flashdocs.z.ai ↗
What it does
multimodal coding and agent workflows: long-context codebase work, visual UI inspection, tool calling, structured output, and document tasks
Strong fit
developers and coding agents that need a low-cost multimodal model, large context, tool use, and an auditable API integration
Poor fit
non-technical users wanting a polished consumer chat app; teams that cannot evaluate vendor benchmarks or review agentic code changes
Limitations
Ox Alpha's free stealth preview ended; launch API pricing is promotional; the 320B-weight model is costly to self-host; headline benchmarks are vendor-reported and need independent validation
Privacy & data handling
Z.ai's API DPA says customer inputs and outputs are processed in real time and not stored; consumer Z.ai services follow a separate privacy policy.
Verified facts
Z.ai says GLM-5.3-Flash was tested anonymously as ox-alpha on OpenCode and OpenRouter before its 26 August 2026 reveal.
Model code glm-5.3-flash; 1M-token context; native image, video and file input; function calling, caching, streaming, and structured output.
API list price: $0.15 input, $0.03 cached input, and $0.50 output per 1M tokens. A 50% launch discount ($0.075/$0.015/$0.25) ends 9 September 2026 at 24:00 UTC+8.
Official weights are published under the MIT license. The model card lists 320B total parameters and 18B active, with SGLang and vLLM self-hosting guidance.
The API DPA says customer inputs and generated content are processed in real time to provide the API service and are not stored on Z.ai servers.
Wondering whether GLM-5.3-Flash is the right pick for your task, or what to pair it with? Describe the outcome and get a full Blueprint.
Build my routeFacts above were checked on the dates shown and can change — confirm on the vendor’s own site before you rely on them.