一种多重权重显著度驱动的大模型混合精度量化的词序列预测方法

By employing simulated annealing and a saliency classification strategy based on matrix sparsity distribution, the problem of inaccurate weight identification in word sequence prediction tasks of large language models is solved, achieving efficient resource utilization and dynamic performance adjustment, and improving the deployment efficiency and accuracy of the model in edge computing environments.

CN120597871BActive Publication Date: 2026-07-17HANGZHOU DIANZI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2025-06-04
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing mixed-precision quantization methods struggle to accurately identify key weights in large language models, leading to logical errors or semantic biases in word sequence prediction tasks. Furthermore, they cannot adapt to performance fluctuations under different task requirements and resource-constrained environments.

Method used

The simulated annealing algorithm is used to dynamically allocate bit width. Combined with matrix sparsity distribution and optimal split point strategy, the saliency classification is used to achieve accurate protection of key weights and efficient use of resources. A dynamic adjustment mechanism is designed to adapt to different data characteristics and model architectures.

Benefits of technology

It improves the inference speed and prediction accuracy of large language models in word sequence prediction tasks, especially in resource-constrained environments such as edge computing, thereby improving the deployment efficiency and performance of the models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597871B_ABST
    Figure CN120597871B_ABST
Patent Text Reader

Abstract

本发明公开了一种多重权重显著度驱动的大模型混合精度量化的词序列预测方法,获取问答数据集,将问答数据集中的文本数据转换为token ID序列;搭建一个加载了基准参数的大语言模型,对大语言模型进行量化处理得到目标函数;将token ID序列输入至目标函数,对于下一个词序列的进行预测,根据概率分布得到最优的词序列预测结果。该方法缓解传统方法因无法适应复杂权重分布而导致的性能下降问题,同时克服动态调整机制缺失引发的模型在不同数据特征和架构设计下的性能波动,最终提升大语言模型在实际部署中的推理速度与预测准确性。
Need to check novelty before this filing date? Find Prior Art