Large language model compression method and system considering sensitivity degrees of different linear layers to numerical value precision and application
By analyzing the numerical accuracy sensitivity of each linear layer in the large language model, low-rank decomposition and quantization are carried out for layers with lower sensitivity, the problem of degradation of performance after model compression in the prior art is solved, and efficient model compression and acceleration are achieved.
Patent Information
- Application Number
- CN202410807141.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-21
- Publication Date
- 2025-05-09
AI Technical Summary
Existing large language model compression methods fail to effectively consider the sensitivity of different linear layers to numerical accuracy, resulting in a degradation of model performance after compression and requiring fine-tuning or retraining to recover losses.
By analyzing the sensitivity of each linear layer to numerical accuracy in a large language model, using singular value decomposition and other methods for analysis, the linear layer with lower sensitivity is decomposed and/or quantized with lower quantization bit width to achieve model compression and acceleration.
A more fine-grained hybrid precision quantization and matrix low-rank decomposition enable greater model compression and acceleration without affecting model performance without fine-tuning.
Smart Images

Figure BDA0004905108570000031 
Figure BDA0004905108570000061 
Figure FDA0004905108560000011
Abstract
Citation Information
Cited By
Large model KV cache multi-dimensional compression method and system oriented to long text task
CN120745761A