whichllm — Browse and compare AI model specs and pricing

Venice AI

GLM 5.3 Flash model ID, context window & pricing

glm

Quick facts

Model ID z-ai-glm-5-3-flash
Source Venice AI
Context Window 1048576
Pricing $0.15 input / $0.50 output per 1M tokens
Capabilities tool calling, reasoning, structured output, temperature control, open weights

Model overview

GLM 5.3 Flash is an AI model from Venice AI with 1048576 token context window and text, image, video input support.

Published pricing is $0.15 input and $0.50 output per 1M tokens.

  • Workloads that use text, image, video inputs with text outputs.
  • Agent and tool workflows that need function calling.
  • Reasoning-heavy prompts where stepwise problem solving matters.
Model ID z-ai-glm-5-3-flash
Provider Venice AI
Family glm
Status -
Knowledge Cutoff -
Release Date 2026-08-21
Input Modalities text, image, video
Output Modalities text
Context Window 1048576
Input Limit -
Output Limit 131072
Tool Calling Yes
Reasoning Yes
Structured Output Yes
Temperature Control Yes
Open Weights Yes
Input Cost / 1M tokens $0.15
Output Cost / 1M tokens $0.50
Reasoning Cost / 1M tokens -
Cache Read Cost / 1M tokens $0.03
Cache Write Cost / 1M tokens -