Quartet: Native FP4 Training Can Be Optimal for Large Language Models
Roberto Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling, Jiale Chen, Mahdi Nikdan, Saleh Ashkboos, Dan Alistarh
May, 2025
Abstract
End-to-end FP4 training of language models, combining low-precision scaling laws with CUDA kernels for Blackwell GPUs.

PhD Candidate in Computer Science
I work on efficient large language models, with a focus on low-precision training, quantization, and scaling laws.