一种多重权重显著度驱动的大模型混合精度量化的词序列预测方法
By employing simulated annealing and a saliency classification strategy based on matrix sparsity distribution, the problem of inaccurate weight identification in word sequence prediction tasks of large language models is solved, achieving efficient resource utilization and dynamic performance adjustment, and improving the deployment efficiency and accuracy of the model in edge computing environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2025-06-04
- Publication Date
- 2026-07-17
AI Technical Summary
Existing mixed-precision quantization methods struggle to accurately identify key weights in large language models, leading to logical errors or semantic biases in word sequence prediction tasks. Furthermore, they cannot adapt to performance fluctuations under different task requirements and resource-constrained environments.
The simulated annealing algorithm is used to dynamically allocate bit width. Combined with matrix sparsity distribution and optimal split point strategy, the saliency classification is used to achieve accurate protection of key weights and efficient use of resources. A dynamic adjustment mechanism is designed to adapt to different data characteristics and model architectures.
It improves the inference speed and prediction accuracy of large language models in word sequence prediction tasks, especially in resource-constrained environments such as edge computing, thereby improving the deployment efficiency and performance of the models.
Smart Images

Figure CN120597871B_ABST