基于KV缓存和大语言模型的对话方法、装置、设备及介质

By constructing a two-layer storage structure and semantic retrieval mechanism in the large language model, intelligent management and cross-session reuse of KV cache are realized, solving the problems of memory occupation and computational redundancy of KV cache, and achieving efficient semantic state reuse and persistent memory.

CN122019737BActive Publication Date: 2026-07-17YUXIANG TECH (HANGZHOU) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUXIANG TECH (HANGZHOU) CO LTD
Filing Date
2026-04-15
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In the autoregressive inference process of large language models, the memory consumption and computational redundancy of KV cache become bottlenecks in existing technologies. They cannot effectively handle the problem of semantically similar but different texts, which leads to the reuse of computational states relying on the complete cache and failing to achieve intelligent long-term reuse.

Method used

By acquiring key-value (KV) caches in real time, using attention weights to filter and generate semantic summary vectors, and constructing a two-layer storage structure, we can achieve the scoring and storage of semantic unit blocks. Combined with a lightweight neural network, we can perform semantic retrieval and splicing, enabling efficient reuse across texts and sessions.

Benefits of technology

It reduces redundant computation by more than 70%, enables efficient reuse of semantically similar states, ensures high fidelity and security of states, and endows large language models with stable and reliable persistent memory capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019737B_ABST
    Figure CN122019737B_ABST
Patent Text Reader

Abstract

本申请公开了基于KV缓存和大语言模型的对话方法、装置、设备及介质,涉及大模型技术领域,包括:实时获取并存储KV缓存,根据KV缓存筛选结果生成语义摘要向量;根据语义单元块的出现记录确定目标评分,基于目标评分确定是否将语义单元块及语义摘要向量存储至第一数据库;将第一KV缓存存储至第二数据库,将用户新问题转换为查询向量,检索第一数据库中与查询向量的语义相似度满足目标阈值的目标语义摘要向量,基于目标语义摘要向量确定第二KV缓存;将第二KV缓存进行拼接,得到拼接数据,利用目标轻量级神经网络对拼接数据进行调整,基于调整后数据和大语言模型进行推理,以得到用户新问题对应的答复。实现了对话历史的高效、长期复用。
Need to check novelty before this filing date? Find Prior Art