Full-integer hardware acceleration circuit of attention module in vision Transform

By designing a fully integrated hardware acceleration circuit, the problems of computing complexity and high power consumption of the visual Transformer model on edge devices are solved, and efficient edge-end deployment is achieved, reducing computing delay and power consumption and improving computing speed.

CN120354898APending Publication Date: 2025-07-22SOUTHEAST UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510317945.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently deploy visual Transformer models on resource-constrained edge devices, especially due to the computational complexity and high power consumption of its attention modules, which leads to inefficient deployment of models at the edge.

Method used

A fully integrated hardware acceleration circuit is designed, including a PE array module, a shift approximation Softmax module, an on-chip SRAM cache module and a controller. Through quantization and compression processing matrix operations, the shift approximation Softmax algorithm is used to realize high-precision full integer inference and reduce calculation delay and power consumption.

Benefits of technology

It realizes efficient calculation of the visual Transformer model on edge devices, reduces computing delay and power consumption, improves computing speed and energy efficiency, and meets the energy-efficient inference requirements of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354898A_ABST
    Figure CN120354898A_ABST
Patent Text Reader

Abstract

The invention discloses a full-integer hardware acceleration circuit of an attention module in a visual Transform, and belongs to the technical field of calculation, reckoning or counting. According to the accelerator, a hardware-friendly circuit design used for accelerating attention module calculation is provided for the full-integer reasoning requirement of a ViT model. The circuit comprises a PE array module used for matrix calculation, a Shiftmax module used for full-integer calculation of a Softmax function, an on-chip SRAM (Static Random Access Memory) cache module and a controller. The PE array module supports row or column combination output, the Shiftmax module uses a base number conversion and shift algorithm to replace traditional floating-point number calculation, and high reasoning precision can be kept. The circuit adopts a row data flow design, is matched with a multi-head self-attention mechanism, and can hide the calculation delay of a Softmax function. According to the method, the calculation delay and energy consumption of the attention module in the visual Transform are remarkably reduced, and the method is suitable for an efficient image recognition task of edge equipment.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Cited By

  • Data processing method based on fusion attention and quantization operation and accelerator

    CN121351891A

  • Pure integer quantization method of visual Transform model and related device

    CN121724076A