Hardware device for sparse attention calculation in neural networks

DE202026002052U1Undetermined Publication Date: 2026-07-09NEUROVEXON UG (HAFTUNGSBESCHRÄNKT)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
NEUROVEXON UG (HAFTUNGSBESCHRÄNKT)
Filing Date
2026-05-04
Publication Date
2026-07-09

Smart Images

  • Figure 00000007_0000
    Figure 00000007_0000
Patent Text Reader

Abstract

Device for processing attention scores in a hardware neural network accelerator, comprising: a) a compute array with a plurality of compute units (1), each compute unit having zero-detection logic which, on each computation operation, generates a sparsity flag (2) indicating whether at least one of the operands has the value zero; b) an attention unit with a score input which, in addition to a score value, has a dedicated port for the sparsity flag (3), the attention unit having a score buffer (4) in which received score values ​​are stored together with the associated sparsity flags;c) a hardware filter stage (8) arranged as a separate state in a state machine of the attention unit between score reception and softmax computation, wherein the filter stage iteratively traverses the stored scores and transfers non-null scores to a compacted buffer (5) with a reduced length (6); and d) a softmax pipeline operating exclusively on the compacted buffer (5) with the reduced length (6), thereby reducing the number of computational operations in the scaling, maximum search, exponential computation, and normalization stages proportionally to the sparsity rate.
Need to check novelty before this filing date? Find Prior Art