A sparse attention calculation method, device and medium for a GPU

By mapping the sparse attention method to a unit hypersphere and generating binary hash codes using a sphere hash function, combined with multi-probe bucket retrieval and a load-adaptive kernel, the problems of low parallelism and low hardware utilization of existing sparse attention methods on GPUs are solved, achieving efficient sparse attention computation.

CN121683877BActive Publication Date: 2026-06-02CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-02-10
Publication Date
2026-06-02

Smart Images

  • Figure CN121683877B_ABST
    Figure CN121683877B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of GPU computing optimization, and in particular to a sparse attention computing method, device and medium for GPU, wherein the method realizes the high efficiency of long context reasoning through the geometric perception sparse attention framework of ball hashing, and combines a large-scale parallel hashing optimization algorithm and a load adaptive computing kernel. Compared with the existing sparse attention methods based on heuristics or gradient learning, the present application realizes higher retrieval recall rate, lower preprocessing overhead and efficient hardware adaptation to irregular sparse patterns.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Efficient sparse attention algorithm optimization method and device, equipment and medium

    CN119476360A

  • Methods and devices for accelerating a transformer with a sparse attention pattern

    US20230133305A1