Why Low-Bit QAT Needs to Happen In Training
Below roughly four bits per weight, post-hoc QAT on byte-level models is catastrophic, and quantization-aware training recovers nearly all of the lost quality
Thesis
If you want a language model small enough to run on a thirty-dollar single-board computer, you have to shrink the weights to a few bits each, or fewer. There are two ways to get there. Post-training quantization (PTQ) takes a finished full-precision model and rounds its weights down to low precision afterward. Quantization-aware training (QAT) makes the model experience the low-precision arithmetic during training, so it learns weights that survive the rounding. The Veritate experiments are unambiguous on which one works at the aggressive end of the precision range: below roughly four bits per weight, post-hoc quantization on our byte-level models is catastrophic, and quantization-aware training recovers nearly all of the lost quality. This article reports the measured cliff and the recoveries.
Ternary as the Destination
The most striking public result in this space is BitNet b1.58 (Ma et al., 2024), which constrains every weight to one of three values, minus one, zero, or plus one. Three levels encode log-base-two of three, about 1.58, bits per weight. The key claim, validated at the three-billion-parameter scale, is that a ternary model trained natively in that regime matches a full-precision model of the same size and token budget on perplexity and downstream accuracy, while using several times less memory and running faster. The operative word is natively. The model is born ternary; it is not rounded to ternary after the fact. Our results explain why that distinction is not a detail but the whole game.
The Post-Hoc Cliff
On our trained 85M byte model, naive data-free post-training quantization falls off a cliff as precision drops. Ternary PTQ applied to the feed-forward layers alone inflated cross-entropy by 266 percent; applied to all layers, by 344 percent. Stacking ternary with structured 2:4 sparsity post-hoc produced a 410 percent regression, which is to say, garbage. Sub-one-bit PTQ schemes were similarly destructive. The diagnosis is not that the dynamic range is wrong; it is that the available level set, three values, or two, cannot represent the weights a full-precision objective settled on. The information lives in distinctions the coarse grid cannot make.
A related and instructive failure: swapping the activation function on a trained model also breaks it. Replacing GELU with ReLU on the finished 85M, with no retraining, drove cross-entropy up by 302 percent even though it produced 81 percent zeros. The model's weights were baked around GELU's smooth negative regime; structurally changing the operator after training destroys the very distinctions the weights encode. The same lesson, in a different costume: aggressive structural change is a training-time decision, not a deployment-time one.
Quantization-Aware Training Recovers It
When the model trains with the quantization in the loop, the cliff bends. On the 85M, ternary that was garbage as PTQ (cross-entropy of 3.56 versus a 0.43 baseline) recovered to 0.56 after only one thousand fine-tuning steps with quantization in the forward pass, and the resulting model is 17 megabytes on disk versus 326 megabytes for full precision, a nineteen-times reduction and five-times smaller than INT8. A second QAT pass that fixes the layer-norm fold ordering dropped a per-column INT8 perplexity from 7.88 to 4.44, a 44 percent improvement.
Related publications
Why Data Quality Decides Small-Model Quality
On saturated consumer hardware, you do not buy small-model quality with a faster framework. You buy it with better data and a better training objective, with numbers to back it.
LLM Decoding Without Changing Output Bytes
Speculative decoding and multi-token prediction cut byte-level generation latency without touching output quality. Here is what we measured on Veritate's models
Making an LLM Efficient By Optimizing Thinking Time, Not Parameter Count
For a model that must fit on a small device, the parameter budget is fixed. The place to spend is inference-time compute.
What Building Infrastructure Taught Us About Trust
Years of building infrastructure taught me that trust is earned quietly, honesty about limits beats big promises, and locking people in is a slow death.
The Hidden Energy Cost of AI, and What Low-Power Inference Could Look Like
The AI energy cost is treated as an inconsequential while the power demand is at the size of some countries and why efficiency, not scale, is the future.
