HIGGS: Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Vladimir Malinovskii, Andrei Panferov, Ivan Ilin, Han Guo, Peter Richtárik, Dan Alistarh
November, 2024
Abstract
A theoretical link between layer reconstruction error and model perplexity, leading to data-free quantization and non-uniform bit allocation.

PhD Candidate in Computer Science
I work on efficient large language models, with a focus on low-precision training, quantization, and scaling laws.