Large language model compression method and system considering sensitivity degrees of different linear layers to numerical value precision and application

By analyzing the numerical accuracy sensitivity of each linear layer in the large language model, low-rank decomposition and quantization are carried out for layers with lower sensitivity, the problem of degradation of performance after model compression in the prior art is solved, and efficient model compression and acceleration are achieved.

CN119961669APending Publication Date: 2025-05-09SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410807141.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing large language model compression methods fail to effectively consider the sensitivity of different linear layers to numerical accuracy, resulting in a degradation of model performance after compression and requiring fine-tuning or retraining to recover losses.

Method used

By analyzing the sensitivity of each linear layer to numerical accuracy in a large language model, using singular value decomposition and other methods for analysis, the linear layer with lower sensitivity is decomposed and/or quantized with lower quantization bit width to achieve model compression and acceleration.

Benefits of technology

A more fine-grained hybrid precision quantization and matrix low-rank decomposition enable greater model compression and acceleration without affecting model performance without fine-tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004905108570000031
    Figure BDA0004905108570000031
  • Figure BDA0004905108570000061
    Figure BDA0004905108570000061
  • Figure FDA0004905108560000011
    Figure FDA0004905108560000011
Patent Text Reader

Abstract

The invention discloses a large language model compression method considering the sensitivity degree of different linear layers to numerical value precision. The compression method comprises the following steps: step 1, analyzing the sensitivity degree of each linear layer to numerical value precision in a large language model; step 2, according to an analysis result in the step 1, performing low-rank decomposition on the linear layer with low sensitivity and / or performing quantization by using a lower quantization bit width; and 3, reconstructing the large language model to obtain a compressed and reconstructed large language model. The invention further discloses a system and application for implementing the large language model compression method.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Cited By

  • Large model KV cache multi-dimensional compression method and system oriented to long text task

    CN120745761A