Attention neural network accelerator and acceleration method
By employing multi-modal prediction and sparse workload balancing techniques in attention neural networks, the problem of wasted computational and hardware resources caused by the sparsity of prediction data in existing methods is solved, achieving efficient neural network acceleration and improving speed and energy efficiency.
Patent Information
- Application Number
- CN202610540631.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-17
AI Technical Summary
Existing methods consume additional computational and hardware resources due to the sparsity of the predicted data, hindering the efficient deployment of attention-based neural networks on resource-constrained devices.
A multi-mode prediction unit is used to perform bitwise AND operations on the most significant mantissa of the input feature matrix and weight matrix using AND gates and adders to generate estimates and a skip mask. Combined with a sparse workload balancing unit and a precise computation unit, a unified acceleration of the attention computation layer and the linear layer is achieved.
It reduces the workload of the precision computing unit, lowers computational complexity and hardware resource consumption, improves inference speed and energy efficiency, and maintains the accuracy of the model.
Smart Images

Figure CN122414263A_ABST