GLM 5.3 Flash: lower AI costs

At current list prices, GLM 5.3 Flash on OpenRouter costs clearly less than the other three recommended models for the same input and output tokens.

Prices checked on 2026-09-19. Providers may change prices at any time; the billing page is always authoritative.

Price per million tokens

To make the comparison easier, Qwen uses list prices for the China North 2 (Beijing) region, converted to USD at $1 ≈ ¥7.26. Your actual bill is charged in the platform's own currency.

Model Input Output Highlights vs. GLM
z-ai/glm-5.3-flashLowest now OpenRouter, currently 50% off
$0.075 $0.25 Great value, slightly stronger reasoning 1×
qwen3.8-flash Alibaba Cloud Model Studio, China North 2 (Beijing)
¥0.80, about $0.110 ¥2.70, about $0.372 Great value, faster responses Input about 1.5×, output about 1.5×
gpt-5.6-luna OpenAI official
$0.20 $1.20 Best quality, most reliable Input about 2.7×, output about 4.8×
qwen3.7-plus Alibaba Cloud Model Studio, input up to 256K
¥2.00, about $0.276 ¥8.00, about $1.101 Best quality, most reliable Input about 3.7×, output about 4.4×

What do 10,000 auto-replies cost?

The estimates below use exactly the same token counts and compare text costs only. Web search, images and other tool calls are not included.

GLM 5.3 FlashOpenRouter current discounted price (50% off)
$2.25
Qwen3.8 FlashAlibaba Cloud Model Studio, China North 2 (Beijing) list price
¥24.10about $3.32
GPT-5.6 LunaOpenAI official standard price
$7.60
Qwen3.7 PlusAlibaba Cloud Model Studio, China North 2 (Beijing) list price
¥64.00about $8.82

GLM is also available through a China-based API

Alibaba Cloud Model Studio offers GLM 5.3 Flash supplied directly by Zhipu. It is easier to reach from mainland China and bills in RMB, but it currently costs more than OpenRouter's limited-time discount.

Alibaba Cloud Model Studio

Billed in RMB

¥0.8 / ¥2.8 per million input / output tokens

The model name is ZHIPU/GLM-5.3-Flash. The API URL must include your own WorkspaceId.

See official Model Studio pricing
Why your actual cost may be higher: GLM 5.3 Flash uses thinking tokens, which are billed at the output price. Conversation history, long business documents, retries and image inputs also use more tokens.

Official pricing sources

These fees are not charged by TGEgg. All API usage is billed to your own model provider account.

Picked a model? Start configuring

Have an API key from that provider ready. TGEgg runs a text and image test before saving.

Open AI model settings