Iot large model pruning and compression method and system based on attention matrix optimization

By evaluating the importance of parameters and performing regional structured pruning on the Attention matrix of large IoT models, the problems of decreased model accuracy and high storage consumption caused by the lack of differentiation of the importance of Attention parameters in traditional methods are solved, achieving efficient inference and accurate compression ratio of lightweight models.

CN121981187BActive Publication Date: 2026-06-26XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN ยท China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD
Filing Date
2026-04-07
Publication Date
2026-06-26

Smart Images

  • Figure CN121981187B_ABST
    Figure CN121981187B_ABST
Patent Text Reader

Abstract

The application provides an Internet of Things large model pruning and compression method and system based on Attention matrix optimization, relates to the technical field of data processing, and comprises the following steps: obtaining a pre-trained Internet of Things large model, extracting parameters in an Attention matrix to obtain a parameter set to be optimized; based on the parameter set, performing importance evaluation on the parameters in the Attention matrix to generate parameter importance distribution data; taking the obtained parameter importance distribution data as a processing basis, establishing high importance calibration, medium importance calibration and low importance calibration, and constructing a ternary calibration matrix. The application effectively balances the compression ratio and inference accuracy.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Deep neural network compression method and device, computer equipment and storage medium

    CN118171697A

  • Parameter pruning method, device and equipment for large language model and readable storage medium

    CN119849579A