G

GLM-5.3

Z.ai · Released Aug 2026

Conditional

GLM-5.3 is Z.ai's latest open-weights coding and cyber-defense model — the same base as GLM-5.2, every gain from post-training. But the headline numbers are still vendor-run, the open weights are held ~2 weeks over emergent security capability, and per-token pricing is unpublished. Reach for defensive security and agentic automation now; wait for the weights and independent audits before standardizing.

Is it right for you?

Good for

  • Defensive cybersecurity: CyberGym 84.5%, top published result ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%)
  • Long-horizon coding: DeepSWE v1.1 66.9% (from 46.2%), Terminal-Bench 3.0 28.3 (six-fold over 4.6)
  • Agentic automation: AutomationBench v1.0.6 jumps 26.2 to 48.2 (+84%); SAO reinforcement learning drives the long-horizon gains
  • Cost-sensitive frontier-adjacent coding: re-verified on Z.ai's official table at $1.40 input / $4.40 output per 1M, the same standing rate as GLM-5.2 and GLM-5.1 despite the post-training gains
  • Self-hosters: the open weights are now published on Hugging Face (zai-org/GLM-5.3, about 756 GB of safetensors) on the same GLM-5.2 base

Not good for

  • Deep offensive exploitation: ExploitBench 54.4% still trails Mythos 5 (78.0%) and GPT-5.6 Sol (76.5%)
  • Budgeting the Flash tier off the launch discount: the 50% promo is no longer on Z.ai's pricing table, which lists GLM-5.3-Flash at $0.15 input / $0.50 output per 1M
  • Modest self-hosting hardware: the published checkpoint is roughly 756 GB, a multi-GPU node deployment rather than a workstation one
  • Vision and multimodal workloads: no multimodal capability announced or benchmarked for GLM-5.3; Z.ai's vision line is a separate GLM-4.6V family
  • Buyers who need independently reproduced results: every headline benchmark remains vendor-run

How it performs by task

Defensive security (CyberGym)

Excellent

84.5% — best published result, ahead of the closed frontier; vendor-reported, not independently audited

Long-horizon software engineering (DeepSWE v1.1)

Very Good

66.9%, up from 46.2% — a 20-point jump, though the figure is Z.ai-run

Terminal/CLI coding (Terminal-Bench 3.0)

Very Good

28.3, six-fold over GLM-5.2's 4.6; level with the frontier on Terminal-Bench 2.1 (88.2 vs 88.8)

Deep exploitation (ExploitBench)

Fair

54.4% more than doubles GLM-5.2 (24.4%) but still trails the closed frontier by 20+ points

Vision / multimodal

Poor

No vision capability announced or benchmarked at launch

Pricing

Input

$1.40 / 1M tokens

Output

$4.40 / 1M tokens

Context

1M ctx; Flash $0.15/$0.50, launch promo ended

View full pricing

Benchmarks

BenchmarkScoreSource
CyberGym84.5% Source
DeepSWE v1.166.9% Source
Terminal-Bench 3.028.3 Source
Terminal-Bench 2.188.2 Source
ExploitBench54.4% Source
ExploitGym (2h / 6h tasks)105 / 130 Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

No verification checks yet

We haven't logged a verification check for this entry. Once a check runs, its history shows here.

How we evaluate