大模型推理的优化方法、装置、计算机设备和存储介质

By performing multi-round bit filtering on the self-attention mechanism module of the large language model, the key-value cache access strategy was optimized, which solved the problem of slow inference speed of the large language model and improved the overall inference efficiency.

CN120633851BActive Publication Date: 2026-07-17TSINGHUA UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2025-06-03
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

The inference speed of large language models is limited by the access speed of key-value cache, resulting in low inference efficiency.

Method used

A pre-defined iterative method is used to perform multiple rounds of bit filtering on the input parameters of the self-attention mechanism module. By predicting importance, unimportant words are judged and eliminated, reducing unnecessary key-value cache access.

Benefits of technology

It improves the speed and efficiency of large model inference, reduces memory access operations of key-value cache, and improves the output efficiency of the self-attention mechanism module.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633851B_ABST
    Figure CN120633851B_ABST
Patent Text Reader

Abstract

本申请涉及一种大模型推理的优化方法、装置、计算机设备和存储介质。所述方法包括:针对大模型的自注意力机制模块的预测过程,采用预设的迭代方法进行预测,得到所述自注意力机制模块的输出;所述迭代方法包括选取所述自注意力机制模块的输入参数中第一比特位的参数进行第一轮预测,并基于所述第一轮预测的预测信息从所述输入参数中选取第二比特位的参数进行第二轮预测,以此类推,直至迭代完所述自注意力机制模块的预测轮次;基于所述自注意力机制模块的输出进行大模型推理,获取推理结果。采用本方法能够降低大模型推理过程中不必要的键值缓存访问,提升大模型的推理速度。
Need to check novelty before this filing date? Find Prior Art