Main Classes¶
Every model in ZeroModels is assembled from the same small set of base classes in
zeromodels.base. You rarely instantiate them directly, but knowing what they provide
explains why every model page looks alike: the same from_weights, the same processor
call, the same generate.
from zeromodels.base import (
BaseConfig,
BaseModel,
BaseGeneration,
BaseSeq2SeqGeneration,
BaseTokenizer,
BaseImageProcessor,
BaseAudioFeatureExtractor,
BaseProcessor,
BaseQuantizer,
fused_attention,
)
Models¶
BaseModel¶
The single base every model builds on. A model assembles itself as a Keras functional
graph with super().__init__(inputs=..., outputs=...). Vision models (CLIP, ViT, the
detectors, segmenters, depth estimators) trace a fixed input shape at construction, which
is why they take an image_size and why changing it means rebuilding. Text and multimodal
models declare their sequence inputs with an undefined length, so a language model still
takes any sequence length; the imperative KV-cache decode that autoregressive generation
needs lives on the task side, in a BaseGeneration / BaseSeq2SeqGeneration mixin layered
over this backbone.
It carries the loading interface below.
Configuration¶
BaseConfig¶
Every model carries a typed config, a BaseConfig subclass whose annotated fields are the
architecture hyperparameters. Model constructors stay flat while the config serializes
nested (the text_config / vision_config blocks you see in kf_config.json), and
BaseConfig bridges the two: constructor_kwargs() flattens a config for the model, and
to_dict() / from_dict() handle serialization. You rarely build one by hand, since
from_weights reconstructs it from the repo, but the Configuration
page covers the field system, the composite sub_configs layout, and the sub-config
classes each model exports.
Loading Weights¶
Every model and preprocessor inherits these classmethods. What each source actually does
(Hub Keras repos, bare-variant on-the-fly conversion, hf: conversion, and caching) is
covered in Loading Weights.
from_weights¶
Model.from_weights(
identifier,
load_weights=True,
skip_mismatch=False,
attn_implementation=None,
quantization=None,
load_dtype=None,
cache_converted=False,
**kwargs,
)
The one entry point you normally use. It dispatches on identifier:
"org/repo"(for example"zeromodels/segformer_b0_ade_512") → Hub Keras viakf_config.json- a bare variant (for example
"qwen3-4b") → on-the-fly conversion from an upstreamhf_id "hf:org/repo"→ convert any compatible Hub checkpoint
Parameters
- identifier (
str): a Hub Keras repo ("zeromodels/segformer_b0_ade_512"), a bare LLM/VLM variant ("qwen3-4b"), or anhf:-prefixed Hub repo ("hf:nvidia/segformer-b0-finetuned-ade-512-512"). - load_weights (
bool, optional, defaults toTrue): setFalseto build the architecture with random initialization. - skip_mismatch (
bool, optional, defaults toFalse): skip weights whose shapes disagree instead of raising, for partially compatible fine-tunes. - attn_implementation (
str, optional): attention kernel to use, seefused_attention. - quantization (
str, optional): quantize while loading, for example"int8". The model builds atload_dtypeand quantizes after. See Quantization. - load_dtype (
str, optional): cast weights on load, typically"bfloat16". - cache_converted (
bool, optional, defaults toFalse): keep the converted Keras weights so the next conversion load skips work. - kwargs: forwarded to the constructor, so
image_size=448oras_backbone=Truego here.
model = SegFormerSemanticSegment.from_weights("zeromodels/segformer_b0_ade_512")
model = SegFormerSemanticSegment.from_weights("hf:<user>/my-finetune")
model = Qwen3TextGenerate.from_weights(
"qwen3-8b", load_dtype="bfloat16", quantization="int8"
)
from_hub_repo, from_variant, and from_hf¶
Model.from_hub_repo(
repo_id,
load_weights=True,
skip_mismatch=False,
**kwargs,
)
Model.from_variant(
variant,
load_weights=True,
skip_mismatch=False,
quantization=None,
**kwargs,
)
Model.from_hf(
hf_id,
load_weights=True,
variant=None,
skip_mismatch=False,
quantization=None,
**kwargs,
)
The three halves from_weights dispatches to. Call them directly only when you want to be
explicit about the source. from_hub_repo reads kf_config.json (and optionally
kf_preprocessor.json for processors). from_hf reads the repo's config.json, so a
fine-tune with a different class count or vocabulary needs no extra arguments.
Load the processor from the same source as the model. A fine-tune can ship a different tokenizer, label set, or normalization; mismatching them fails quietly with wrong output rather than loudly with an error.
quantize¶
Quantize an already-built model in place. Passing quantization= to from_weights is
usually better, since it avoids materializing float weights first. See
Quantization.
Generation¶
BaseGeneration¶
model.generate(
input_ids,
attention_mask=None,
max_new_tokens=None,
eos_token_id=None,
sampler=None,
seed=None,
**prefill_inputs,
)
Backend-agnostic autoregressive decoding for decoder-only models, the counterpart to
Hugging Face's GenerationMixin. Any extra tensors a multimodal model needs, pixel values
or audio features, ride along in **prefill_inputs, which is why a VLM call looks like
model.generate(**inputs, max_new_tokens=64).
Parameters
- input_ids: the prompt token ids from a tokenizer or processor.
- attention_mask (optional): padding mask for batched prompts.
- max_new_tokens (
int, optional): decode budget. - eos_token_id (
int, optional): stop token, defaulting to the model's own. - sampler (optional): a sampler from
zeromodels.samplers; greedy if omitted. - seed (
int, optional): seed for stochastic samplers.
BaseSeq2SeqGeneration¶
model.generate(
encoder_inputs,
decoder_input_ids,
max_new_tokens=None,
eos_token_id=None,
sampler=None,
seed=None,
)
model.encode(encoder_inputs)
The encoder-decoder flavor, used by Whisper,
Speech2Text, and Moonshine. encode runs the encoder
once so you can decode repeatedly against the same audio.
Those speech models wrap this in a friendlier generate(audio, processor, ...) that owns
the whole pipeline; see their pages.
Preprocessing¶
All preprocessors share PreprocessorMixin, so they also get from_weights,
from_hub_repo, from_variant, and from_hf.
BaseTokenizer¶
Text to token ids and back. Subclasses add encode, chat templating, and any
model-specific parsing (LocateAnything's parse_boxes, for example).
BaseImageProcessor¶
Images to pixel_values. Beyond the call, it carries the shared, backend-agnostic pixel
helpers every image model reuses: resize, center_crop, pad, rescale,
normalize_image, rescale_and_normalize, preprocess_image, and stack_images for
batching. Normalization constants live here too (IMAGENET_STANDARD_MEAN,
OPENAI_CLIP_MEAN, and friends).
Task-specific post-processing is added by subclasses:
post_process_object_detection, post_process_semantic_segmentation,
post_process_depth_estimation, post_process_masks.
BaseAudioFeatureExtractor¶
Waveform to model input. What that means depends on the model: a log-mel spectrogram for Whisper and Granite Speech, normalized filterbanks for Speech2Text, and the raw waveform itself for Moonshine, which has no spectrogram step.
sampling_rate tells the extractor what you are handing it; it does not resample.
Feed 44.1 kHz audio while claiming 16 kHz and you get a confident, wrong transcript.
BaseProcessor¶
The composite that bundles a tokenizer with an image processor or audio feature extractor
and renders chat templates. Multimodal models expose this as the single object you call.
Components are declared as class attributes, so processor.tokenizer and
processor.image_processor are always reachable.
Quantization¶
BaseQuantizer¶
Base for the weight-only (tensor-level) quantizers. Helpers normalize_axes(axis, ndim)
and single_axis(axis, ndim) resolve contraction axes. The quantized layers built on this
(QuantizedDense, QuantizedEinsumDense, QuantizedEmbedding, QuantizedExperts) and
the quantize_model / dequantize_model entry points are covered in
Quantization.
KfQuantizer¶
Model-level quantizer, the transformers HfQuantizer analog. from_weights reads a
repo's quantization_config and runs the matching KfQuantizer to swap in the packed /
int layers before the weights load, so the model stays quantization-agnostic. Dispatched
by quant_method: Mxfp4KfQuantizer (GPT-OSS native experts) / WeightOnlyKfQuantizer
(int8 / int4 / fp8). Covered in Quantization.
Attention¶
fused_attention¶
fused_attention(
query,
key,
value,
scale,
attention_mask=None,
soft_cap=None,
dropout=None,
training=None,
attn_implementation=None,
)
Scaled dot-product attention with a selectable backend kernel, used by every attention layer in the library so a single implementation choice applies everywhere.
- scale (
float): the1/sqrt(head_dim)factor, applied inside. - attention_mask (optional): additive mask broadcastable to
(B, heads, T_q, T_kv). - soft_cap (
float, optional): logit soft-capping, used by Gemma 2. - attn_implementation (
str, optional): pick the kernel; pass it throughfrom_weightsto set it model-wide.
See also Utilities for the image, video, visualization, and label helpers.