基于KV缓存和大语言模型的对话方法、装置、设备及介质
By constructing a two-layer storage structure and semantic retrieval mechanism in the large language model, intelligent management and cross-session reuse of KV cache are realized, solving the problems of memory occupation and computational redundancy of KV cache, and achieving efficient semantic state reuse and persistent memory.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUXIANG TECH (HANGZHOU) CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-17
AI Technical Summary
In the autoregressive inference process of large language models, the memory consumption and computational redundancy of KV cache become bottlenecks in existing technologies. They cannot effectively handle the problem of semantically similar but different texts, which leads to the reuse of computational states relying on the complete cache and failing to achieve intelligent long-term reuse.
By acquiring key-value (KV) caches in real time, using attention weights to filter and generate semantic summary vectors, and constructing a two-layer storage structure, we can achieve the scoring and storage of semantic unit blocks. Combined with a lightweight neural network, we can perform semantic retrieval and splicing, enabling efficient reuse across texts and sessions.
It reduces redundant computation by more than 70%, enables efficient reuse of semantically similar states, ensures high fidelity and security of states, and endows large language models with stable and reliable persistent memory capabilities.
Smart Images

Figure CN122019737B_ABST