Block-based attention sparse processing method, device, equipment and medium
By dividing and filtering query vectors and key vectors into blocks, and selecting important block products, the problem of excessive computational cost and memory consumption when large models process long texts, high-resolution images, or multimodal data is solved, and more efficient attention computation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KUNWANG (SHANGHAI) TECH CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-03
AI Technical Summary
Modern large models suffer from excessive computational costs and memory consumption when processing long texts, high-resolution images, or multimodal data, and existing attention mechanisms have failed to effectively alleviate the problem of information overload.
A block-based attention sparsity processing method is adopted. By dividing the query vector and key vector into blocks, important tokens are filtered, the importance score of the initial block product is calculated, and the target block product is filtered to reduce the number of blocks involved in the calculation, thus achieving sparsity processing.
It significantly reduces computational complexity and memory consumption while preserving key semantic information, thereby improving the efficiency of attention computation and the processing speed of the model.
Smart Images

Figure CN122334348A_ABST