HIGGS: Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Abstract

A theoretical link between layer reconstruction error and model perplexity, leading to data-free quantization and non-uniform bit allocation.

Publication
NAACL 2025
Andrei Panferov
Andrei Panferov
PhD Candidate in Computer Science

I work on efficient large language models, with a focus on low-precision training, quantization, and scaling laws.