GLM 4.7 Flash Original Thinking model ID, context window & pricing
glm-flash
Quick facts
Model ID z-ai/glm-4.7-flash-original:thinking
Source NanoGPT
Context Window 200000
Pricing $0.07 input / $0.40 output per 1M tokens
Capabilities tool calling, reasoning, open weights
Model overview
GLM 4.7 Flash Original Thinking is an AI model from NanoGPT with 200000 token context window and text input support.
Published pricing is $0.07 input and $0.40 output per 1M tokens.
- Workloads that use text inputs with text outputs.
- Agent and tool workflows that need function calling.
- Reasoning-heavy prompts where stepwise problem solving matters.
Model ID z-ai/glm-4.7-flash-original:thinking
Provider NanoGPT
Family glm-flash
Status -
Knowledge Cutoff -
Release Date 2026-01-19
Input Modalities text
Output Modalities text
Context Window 200000
Input Limit 200000
Output Limit 128000
Tool Calling Yes
Reasoning Yes
Structured Output No
Temperature Control -
Open Weights Yes
Input Cost / 1M tokens $0.07
Output Cost / 1M tokens $0.40
Reasoning Cost / 1M tokens -
Cache Read Cost / 1M tokens $0.04
Cache Write Cost / 1M tokens -