Open source · LLM inference
OpenDynamicGGUF
Dynamic per-tensor mixed-precision GGUF quantisation — for any model.
Measures how sensitive every tensor is to quantisation, then picks the bit-width for each one under your size or VRAM budget — and ships a reproducible recipe, not just a smaller file.
- Every bit decision traced to a measured ΔKL-divergence probe.
- Recipes that rebuild the exact GGUF file, bit for bit.
- One pipeline for dense, MoE and SSM models, built on llama.cpp.
$ odg quantize --model functiongemma:latest \
--target-size 3.2GB
token_embd → Q8_0 # pinned
attn_v (all) → Q6_K # sensitive, small
attn_q/k (early) → Q5_K
ffn_gate (mid) → Q3_K # large, robust
ffn_down (all) → Q4_K
output → Q8_0 # pinned
- Python
- PyTorch
- llama.cpp
- GGUF
- Hugging Face
- KL-divergence


