GLM (text & vision-language)¶
Zhipu / Z.ai's GLM family in pure Keras 3: the text LLMs (GLM-4 through the GLM-5 MoE series) and the GLM-4V vision-language models: one implementation per family, runnable unchanged on TensorFlow / Torch / JAX and bit-close to the HuggingFace reference.
Papers / refs: ChatGLM (GLM-4) · GLM-4.5 · GLM-4.1V · GLM-5.2
| Family | Module | Kind | Decoder |
|---|---|---|---|
| GLM-4-9B | kerasformers.models.glm |
text | GLM (interleaved partial RoPE) |
| GLM-4-0414 / GLM-Z1 | kerasformers.models.glm4 |
text | GLM + sandwich norms |
| GLM-4.5 / GLM-4.6 | kerasformers.models.glm4_moe |
text (MoE) | grouped-top-k router + shared expert, NeoX partial RoPE |
| GLM-5 / GLM-5.1 / GLM-5.2 | kerasformers.models.glm5_moe |
text (MoE) | MLA + DSA (DeepSeek Sparse Attention) + DeepSeekMoE |
| GLM-4.1V | kerasformers.models.glm4v |
image+video+text | GLM + Qwen2-VL-class M-RoPE vision |
| GLM-4.5V | kerasformers.models.glm4v_moe |
image+video+text (MoE) | GLM-4.5 MoE + GLM-4V vision |
Each family exposes a *Model (features, call → last_hidden_state) and a
*Generate (adds the LM head + .generate(), call → logits), plus a
*Tokenizer (and a *Processor for the VL families).
Loading¶
Weights convert on the fly from the public Hugging Face checkpoints
(safetensors downloaded + mapped at load time; bf16 cast to float32, FP8 MoE
checkpoints dequantized). Use the friendly variant name, or a raw hf: id:
from kerasformers.models.glm import GlmGenerate
from kerasformers.models.glm4_moe import Glm4MoeGenerate
gen = GlmGenerate.from_weights("glm-4-9b-chat") # text
gen = Glm4MoeGenerate.from_weights("glm-4.5-air") # text MoE
# raw hf: ids work too
gen = GlmGenerate.from_weights("hf:THUDM/glm-4-9b-chat-hf")
Available variants¶
Text:
| Family | Variants (from_weights("…")) |
Hub |
|---|---|---|
| GLM-4-9B | glm-4-9b, glm-4-9b-chat |
THUDM/glm-4-9b{,-chat-hf} |
| GLM-4-0414 / GLM-Z1 | glm-4-9b-0414, glm-4-32b-0414, glm-z1-9b-0414, glm-z1-32b-0414 |
THUDM/GLM-4-*-0414, THUDM/GLM-Z1-*-0414 |
| GLM-4.5 / GLM-4.6 (MoE) | glm-4.5, glm-4.5-air, glm-4.6 |
zai-org/GLM-4.5{,-Air}, zai-org/GLM-4.6 |
| GLM-5 / GLM-5.1 / GLM-5.2 (MoE) | glm5, glm5_1, glm5_2 |
zai-org/GLM-5{,.1,.2} |
Vision-language:
| Family | Variants | Hub |
|---|---|---|
| GLM-4.1V | glm-4.1v-9b-thinking, glm-4.1v-9b-base |
zai-org/GLM-4.1V-9B-{Thinking,Base} |
| GLM-4.5V (MoE) | glm-4.5v |
zai-org/GLM-4.5V |
Generation¶
.generate() is greedy decoding with a KV cache. LLMs use the tokenizer
(text only); VLMs use the processor (tokenizer + image/video processor), with
images inline in the conversation. Load the tokenizer / processor with the same
identifier you give the model.
# text LLM
from kerasformers.models.glm import GlmGenerate, GlmTokenizer
model = GlmGenerate.from_weights("glm-4-9b-chat")
tokenizer = GlmTokenizer.from_weights("glm-4-9b-chat")
messages = [{"role": "user", "content": "Name three prime numbers."}]
inputs = tokenizer(messages)
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0]))
# vision-language (GLM-4.1V)
from kerasformers.models.glm4v import Glm4vGenerate, Glm4vProcessor
model = Glm4vGenerate.from_weights("glm-4.1v-9b-thinking")
processor = Glm4vProcessor.from_weights("glm-4.1v-9b-thinking")
conversation = [
{"role": "user", "content": [
{"type": "image", "path": "/path/to/image.jpg"},
{"type": "text", "text": "What is in the image?"},
]},
]
inputs = processor(conversation)
outputs = model.generate(**inputs, max_new_tokens=128)
print(processor.decode(outputs[0], skip_special_tokens=True))
Architecture notes¶
- GLM-4 / GLM-4-0414 (
glm,glm4): post-norm GLM block with interleaved partial RoPE; GLM-4-0414 adds sandwich (input/post) norms around attention and MLP. Theglm-z1-*checkpoints are reasoning fine-tunes on the same arch. - GLM-4.5 / GLM-4.6 (
glm4_moe): DeepSeek-V3-style grouped-top-k MoE router with a shared expert and fused expert einsums, NeoX partial RoPE (not interleaved); the auxiliary MTP head is dropped, and FP8-released weights are dequantized on load. - GLM-5 / 5.1 / 5.2 (
glm5_moe): Multi-head Latent Attention (MLA) + DeepSeek Sparse Attention (DSA) with a Lightning-Indexer, on top of DeepSeekMoE. 5 ≡ 5.1 in the forward; 5.2 only changesmax_position_embeddings(1M) andrope_theta. Cached decode skips the indexer for exact short-context output. - GLM-4V / GLM-4.5V (
glm4v,glm4v_moe): Qwen2-VL-class M-RoPE vision tower with a learned position embedding (bicubic-interpolated), a Conv3d patch embed + downsample conv, and a SwiGLU patch merger; the text side is GLM-4 (4V) or GLM-4.5 MoE (4.5V).
Parity vs HuggingFace Reference¶
Validated against transformers (cloned main, eager attention) on real forward
passes: greedy generation is token-identical; max|Δ logits| ≈ 1.8e-7 (GLM-4),
2.8e-7 (GLM-4-0414), 2e-7–3e-7 (GLM-4.5 / 4.1V / 4.5V), ~5e-8 (GLM-5 series,
fp32 tiny-config). Verified across the torch, jax, and tensorflow backends.