Full-integer hardware acceleration circuit of attention module in vision Transform
By designing a fully integrated hardware acceleration circuit, the problems of computing complexity and high power consumption of the visual Transformer model on edge devices are solved, and efficient edge-end deployment is achieved, reducing computing delay and power consumption and improving computing speed.
Patent Information
- Application Number
- CN202510317945.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to efficiently deploy visual Transformer models on resource-constrained edge devices, especially due to the computational complexity and high power consumption of its attention modules, which leads to inefficient deployment of models at the edge.
A fully integrated hardware acceleration circuit is designed, including a PE array module, a shift approximation Softmax module, an on-chip SRAM cache module and a controller. Through quantization and compression processing matrix operations, the shift approximation Softmax algorithm is used to realize high-precision full integer inference and reduce calculation delay and power consumption.
It realizes efficient calculation of the visual Transformer model on edge devices, reduces computing delay and power consumption, improves computing speed and energy efficiency, and meets the energy-efficient inference requirements of edge devices.
Smart Images

Figure CN120354898A_ABST
Abstract
Citation Information
Cited By
Data processing method based on fusion attention and quantization operation and accelerator
CN121351891A
Pure integer quantization method of visual Transform model and related device
CN121724076A