quantization 4
- Int8Tensor: A PyTorch-Native INT8 Quantization Subclass for torchao
- TwinQuant: Learnable Subspace Decomposition for W4A4 LLM Quantization
- SVDQuant: Absorbing Weight Outliers into a Low-Rank Branch for 4-Bit Diffusion Models
- Processing-in-Memory for Decode-Stage GEMV: Mixed-Signal, Digital, and the Cost of Noise