An Adaptive Granularity Management System for a Large Language Model KV Cache Based on Access Co-occurrence Awareness

By adopting an adaptive granularity management system, the problem of mismatch between storage granularity and access granularity in the KV cache of large language models is solved, achieving efficient IO bandwidth utilization and reduced inference latency.

CN122332302APending Publication Date: 2026-07-03BEIJING TREND TECHNOLOGY CO LTD
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING TREND TECHNOLOGY CO LTD
Filing Date
2026-04-20
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing KV cache management for large language models lacks awareness of attention mechanism access patterns, resulting in a mismatch between storage granularity and actual access granularity. This leads to read amplification and write amplification issues, wasting IO bandwidth and increasing inference latency.

Method used

An adaptive granularity management system based on the access co-occurrence awareness of the large language model KV Cache is adopted. It achieves adaptive granularity management through a pattern feature extraction module, a head group clustering construction module, an elastic boundary generation module, a semantic weight calculation module, and a hierarchical compression storage module. It includes techniques such as sparse structure analysis, attention pattern fingerprint data extraction, head tuple aggregation, dynamic mapping table generation, and data block merging.

Benefits of technology

It effectively decouples the model structure from the sequence length, eliminates boundary jumps and memory fragmentation, and improves IO bandwidth utilization and inference efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122332302A_ABST
    Figure CN122332302A_ABST
Patent Text Reader

Abstract

This invention discloses an adaptive granularity management system for a large language model KV cache based on access co-occurrence awareness, relating to the field of large language model technology. It includes a pattern feature extraction module, a head group clustering construction module, a flexible boundary generation module, a semantic weight calculation module, a hierarchical compression storage module, and a data merging and scheduling module. The pattern feature extraction module performs sparse structure analysis on the attention weight matrix of the large language model to extract attention pattern fingerprint data. The head group clustering construction module receives the attention pattern fingerprint data, aggregates multiple attention heads with similar pattern features into head tuples, and generates a parameterized co-occurrence template. The flexible boundary generation module obtains the current inference sequence length and, combined with the parameterized co-occurrence template, converts the relative position parameters within the template into a dynamic mapping table through proportional scaling. The semantic weight calculation module calculates the range of co-occurrence groups based on the dynamic mapping table.
Need to check novelty before this filing date? Find Prior Art