int8¶
Per-channel symmetric int8, roughly 4x smaller. The default scheme, and the only one besides int4 that runs on all three backends.
Usage¶
from kerasformers.models.qwen3 import Qwen3Generate
from kerasformers.quantization import quantize_model
# load and quantize in one call
model = Qwen3Generate.from_weights("qwen3-4b", quantization="int8")
# or quantize a model you already have, in place
quantize_model(model, "int8")
quantize_model swaps every built Dense, EinsumDense, Embedding, and fused MoE expert
bank for its quantized counterpart and frees the float weights. Functional models cannot be
mutated in place, so those are cloned: use the returned model.
Int8Config¶
The declarative recipe for int8, used when the bare "int8" string is not enough control.
It is QuantizationConfig with mode="int8" fixed and group_size dropped, since block
size means nothing for a per-channel scheme.
Parameters
- skip_modules (
tupleofstr, optional, defaults to("lm_head",)): name substrings; any layer whose path contains one is left in float. The output head is the most accuracy-sensitive layer in a decoder, so it is skipped by default. - quantize_embeddings (
bool, optional, defaults toTrue): quantizeEmbeddinglayers. SetFalseto keep the token table in float. - overrides (
dict, optional):{name_substring: mode}, per-layer precision, checked beforemode. Use it to raise or lower precision on specific layers.
Resolution order for any layer is skip_modules first, then overrides, then the mode. A
layer matched by skip_modules stays float even if overrides also names it.
Using the config¶
Pass it anywhere a scheme string is accepted, from_weights included:
from kerasformers.quantization import Int8Config, quantize_model
cfg = Int8Config(skip_modules=("lm_head", "q_proj"))
model = Qwen3Generate.from_weights("qwen3-4b", quantization=cfg)
quantize_model(model, cfg)
Keeping an extra layer in float, when one measures badly:
Keeping the token embedding table in float, worth trying when a model with a large vocabulary loses quality:
Mixing precisions, here dropping most of the model to int4 while holding the first decoder block at int8:
from kerasformers.quantization import QuantizationConfig
QuantizationConfig(mode="int4", group_size=128, overrides={"decoder_layer_0": "int8"})
The effect is visible on the layers themselves after quantizing:
Embeddings are quantized int8 with a per-row scale. That holds in int4 mode too: the 4-bit
savings live in the Dense weights, so QuantizedEmbedding carries an Int8Quantizer
either way.
See Quantization for the shared machinery, and int4 / fp8 for the other schemes.