CODEBOOK COMPRESSION FOR VECTOR-QUANTIZED NEURAL NETWORKS
ID202606473APending Publication Date: 2026-08-07QUALCOMM INC
0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- ID · ID
- Patent Type
- Applications
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-12-03
- Publication Date
- 2026-08-07
Abstract
Systems and techniques are described herein for quantizing codebooks used in the context of post-training parameter quantization (e.g., vectors of weights) of a pre-trained model. For example, the device may perform rank reduction on a tensor from the codebook corresponding to the parameters of a layer of the pre-trained machine learning model to produce a first tensor factor having the first form and a second tensor factor having the second form. The device may perform optimization techniques on the first tensor factor and the second tensor factor to minimize the output reconstruction error of the layer. The device may quantize the first tensor factor to produce a codebook with a reduced size.
Need to check novelty before this filing date? Find Prior Art