Block-based attention sparse processing method, device, equipment and medium

By dividing and filtering query vectors and key vectors into blocks, and selecting important block products, the problem of excessive computational cost and memory consumption when large models process long texts, high-resolution images, or multimodal data is solved, and more efficient attention computation is achieved.

CN122334348APending Publication Date: 2026-07-03KUNWANG (SHANGHAI) TECH CO LTD
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KUNWANG (SHANGHAI) TECH CO LTD
Filing Date
2026-05-08
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Modern large models suffer from excessive computational costs and memory consumption when processing long texts, high-resolution images, or multimodal data, and existing attention mechanisms have failed to effectively alleviate the problem of information overload.

Method used

A block-based attention sparsity processing method is adopted. By dividing the query vector and key vector into blocks, important tokens are filtered, the importance score of the initial block product is calculated, and the target block product is filtered to reduce the number of blocks involved in the calculation, thus achieving sparsity processing.

Benefits of technology

It significantly reduces computational complexity and memory consumption while preserving key semantic information, thereby improving the efficiency of attention computation and the processing speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122334348A_ABST
    Figure CN122334348A_ABST
Patent Text Reader

Abstract

This disclosure provides a block-based attention sparse processing method, apparatus, device, and medium, relating to the field of computer technology, particularly artificial intelligence, deep learning, and computer vision. The specific implementation scheme includes: dividing a query vector into blocks to obtain query blocks, and dividing a key vector into blocks to obtain key blocks; filtering the tokens included in each vector block and updating the vector block; performing block multiplication on the query blocks in the query vector and the key blocks in the key vector to obtain multiple initial block products; determining the importance score of each initial block product based on the elements in each initial block product; filtering each initial block product based on its importance score and the number of filters to obtain at least one target block product; and determining the attention result of the target input after sparse processing based on each target block product. Embodiments of this disclosure can improve the efficiency of attention computation.
Need to check novelty before this filing date? Find Prior Art