A sparse attention calculation method, device and medium for a GPU
By mapping the sparse attention method to a unit hypersphere and generating binary hash codes using a sphere hash function, combined with multi-probe bucket retrieval and a load-adaptive kernel, the problems of low parallelism and low hardware utilization of existing sparse attention methods on GPUs are solved, achieving efficient sparse attention computation.
CN121683877BActive Publication Date: 2026-06-02CENT SOUTH UNIV
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-02-10
- Publication Date
- 2026-06-02
Smart Images

Figure CN121683877B_ABST
Abstract
The present application relates to the technical field of GPU computing optimization, and in particular to a sparse attention computing method, device and medium for GPU, wherein the method realizes the high efficiency of long context reasoning through the geometric perception sparse attention framework of ball hashing, and combines a large-scale parallel hashing optimization algorithm and a load adaptive computing kernel. Compared with the existing sparse attention methods based on heuristics or gradient learning, the present application realizes higher retrieval recall rate, lower preprocessing overhead and efficient hardware adaptation to irregular sparse patterns.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Efficient sparse attention algorithm optimization method and device, equipment and medium
CN119476360A
Methods and devices for accelerating a transformer with a sparse attention pattern
US20230133305A1