Iot large model pruning and compression method and system based on attention matrix optimization
By evaluating the importance of parameters and performing regional structured pruning on the Attention matrix of large IoT models, the problems of decreased model accuracy and high storage consumption caused by the lack of differentiation of the importance of Attention parameters in traditional methods are solved, achieving efficient inference and accurate compression ratio of lightweight models.
CN121981187BActive Publication Date: 2026-06-26XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD
Patent Information
- Authority / Receiving Office
- CN ยท China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD
- Filing Date
- 2026-04-07
- Publication Date
- 2026-06-26
Smart Images

Figure CN121981187B_ABST
Abstract
The application provides an Internet of Things large model pruning and compression method and system based on Attention matrix optimization, relates to the technical field of data processing, and comprises the following steps: obtaining a pre-trained Internet of Things large model, extracting parameters in an Attention matrix to obtain a parameter set to be optimized; based on the parameter set, performing importance evaluation on the parameters in the Attention matrix to generate parameter importance distribution data; taking the obtained parameter importance distribution data as a processing basis, establishing high importance calibration, medium importance calibration and low importance calibration, and constructing a ternary calibration matrix. The application effectively balances the compression ratio and inference accuracy.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Deep neural network compression method and device, computer equipment and storage medium
CN118171697A
Parameter pruning method, device and equipment for large language model and readable storage medium
CN119849579A