Archives
- 05 Aug Int8Tensor: A PyTorch-Native INT8 Quantization Subclass for torchao
- 19 Jul TwinQuant: Learnable Subspace Decomposition for W4A4 LLM Quantization
- 19 Jul SVDQuant: Absorbing Weight Outliers into a Low-Rank Branch for 4-Bit Diffusion Models
- 09 Jul Processing-in-Memory for Decode-Stage GEMV: Mixed-Signal, Digital, and the Cost of Noise
- 09 May Diary
- 27 Apr From CUDA Graph Crash to Tensor Core: Tracing a vLLM INT8 Bug
- 10 Apr 안뇽하세요....