Efficient self-attention reasoning method on Versa ACAP architecture

By using ConSmax activation function and block calculation method on the Versal ACAP architecture, combining AIE processing unit and PL data engine, a self-attention hardware accelerator is designed, which solves the high complexity of the self-attention mechanism, and realizes efficient self-attention reasoning acceleration, improving system performance and resource utilization.

CN120430341APending Publication Date: 2025-08-05BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510628839.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing hardware platforms are difficult to effectively accelerate Transformer's self-attention mechanism, especially in long-sequence computing, the accelerator effect is not ideal and the development cost is high.

Method used

The ConSmax activation function is used to replace Softmax, and self-attention is calculated in blocks. Combined with the AIE processing unit and PL data engine of Versal ACAP architecture, a self-attention hardware accelerator is designed and multiple self-attention processing modules are deployed to achieve efficient computing.

Benefits of technology

It improves the throughput of self-attention reasoning, reduces inference latency, and has huge advantages in batch inference, achieving maximum utilization of hardware resources and system performance improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430341A_ABST
    Figure CN120430341A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient self-attention reasoning method on a Versa ACAP architecture, and belongs to the field of software and hardware collaborative acceleration numerical calculation. Firstly, a self-attention Softmax activation function is replaced with an activation function, self-attention block calculation is carried out, and the block calculation method is more suitable for hardware design. Secondly, designing a hardware accelerator for realizing self-attention on a Versa ACAP architecture, realizing an AIE processing unit by using an AIE array of the Versa ACAP, realizing a data engine for providing data and scheduling for the AIE processing unit by using a PL of the Versa ACAP, and storing source data and results by using an on-chip DDR (Double Data Rate), and combining the hardware accelerator, the AIE processing unit and the on-chip DDR into a self-attention module (ASA) for undertaking self-attention operation. Experiments prove that through the accelerator deployed by adopting the method, the throughput of the self-attention accelerator is effectively improved, the reasoning delay is reduced, meanwhile, the accelerator has huge advantages in the aspect of batch reasoning, the reasoning cost is reduced, and the reasoning speed is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a deployment method for an efficient self-attention inference accelerator on the Versal ACAP architecture, belonging to the field of software and hardware collaborative accelerated numerical computing. Background Art

[0002] The success of Transformer is attributed to its unique self-attention mechanism, which dynamically weights the degree of association between each element in the sequence and other elements, thereby modeling contextual information. This mechanism makes it perform well in expressing long-distance dependencies between data. However, due to its time and memory complexity of O(n 2 ), making the running time and memory overhead of Transformer long sequences a current challenge.

[0003] Currently, GPUs are widely used as the primary hardware acceleration platform in the industry. However, due to inherent hardware limitations, GPUs inevitably suffer from high power consumption and are difficult to customize. Other platforms, such as FPGAs and ASICs, while offering finer design granularity, offer a larger design space to explore, often resulting in longer design cycles and higher development costs.

[0004] Considering the increasing demand for acceleration of the self-attention mechanism, the accelerator implemented by traditional hardware platforms is not ideal. The present invention proposes a deployment method for an efficient self-attention reasoning accelerator on the Versal ACAP architecture. A block self-attention calculation method using the ConSmax activation function is adopted, which is more suitable for hardware design. Secondly, a self-attention hardware accelerator is designed and implemented on the Versal ACAP architecture. The AIE array of the Versal ACAP is used to implement the AIE processing unit, and the PL of the Versal ACAP is used to implement the data engine to provide data and scheduling for the AIE processing unit, ultimately achieving high-performance self-attention reasoning calculation. Summary of the Invention

[0005] The present invention is different from the existing self-attention accelerator deployment method. It utilizes the characteristics of Versal ACAP and proposes a block self-attention (SelfAttention) calculation method using the ConSmax activation function. First, the Softmax activation function is replaced by the activation function, and the self-attention is calculated in blocks. This block calculation method is more suitable for hardware design. Secondly, a self-attention hardware accelerator is designed and implemented on the Versal ACAP architecture. The AIE array of Versal ACAP is used to implement the AIE processing unit, and the PL of Versal ACAP is used to implement the data engine to provide data and scheduling for the AIE processing unit. The on-chip DDR is used to store source data and results. The three are combined into a self-attention module (ASA) to undertake self-attention operations. Finally, multiple ASA designs are deployed in the hardware, and the x86 CPU assigns tasks to them, ultimately achieving high-performance self-attention calculations, so that it can efficiently complete forward propagation (inference) on the Xilinx Versal VCK5000 board. This method can be divided into the following five steps:

[0006] (1) Replace Softmax with ConSmax

[0007] The Softmax activation function of the self-attention calculation is replaced by the ConSmax activation function, and its accuracy is very close to Softmax (<1%), effectively decoupling the row element dependency, which is more conducive to the parallel design of hardware.

[0008] (2) Specify the method for calculating self-attention in blocks

[0009] The block-wise parallel computing of self-attention can fully utilize the characteristics of the ACAP architecture, reduce memory and computing costs, and reduce memory complexity to O(n).

[0010] (3) Implementing AIE processing units in AIE arrays

[0011] The AIE processing unit based on the AIE array is used to undertake computationally intensive matrix multiplication and ConSmax calculations. It can adapt to customized Self-Attention data streams and has the characteristics of high performance and low power consumption.

[0012] (4) Implementing the PL data engine

[0013] The data engine on the PL side can fully guarantee the high-speed data supply of the AIE processing unit.

[0014] (5) Deploy multiple self-attention processing modules

[0015] The AIE processing unit is combined with the PL data engine to form a self-attention processing module. At the same time, multiple such modules can be deployed in the hardware. After being called on the HOST side, the self-attention reasoning acceleration is completed.

[0016] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:

[0017] First, compared with traditional hardware-implemented accelerators, this invention fully utilizes the hardware customizability of Versal ACAP to implement a self-attention inference accelerator deployment architecture, greatly improving the availability of Versal ACAP.

[0018] Secondly, the deployment method of the efficient self-attention inference accelerator on the Versal ACAP architecture effectively improves throughput and reduces inference latency. At the same time, this accelerator has huge advantages in batch inference.

[0019] Finally, the present invention makes full use of the idea of software and hardware collaboration to improve the compatibility between the algorithm and hardware, enhance the overall performance of the system, and maximize the utilization of hardware resources.

[0020] Experiments have confirmed that when accelerating the BERT-Base Int8 model, the present invention has an overall throughput of 57.3852TOPS, a running time of 0.014 milliseconds, a power of 100.48W, and an energy efficiency of 571.11GOPS / W under the conditions of AIE of 1.25GHZ and PL of 300MHZ frequency. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Schematic diagram of the AIE processing unit design.

[0022] Figure 2 Schematic diagram of the self-attention processing module.

[0023] Figure 3 This is a schematic diagram of the top-level architecture of the accelerator. DETAILED DESCRIPTION

[0024] Based on the above description, the following is a specific implementation process, but the scope of protection of this patent is not limited to this implementation process.

[0025] Step 1: Replace the original self-attention Softmax activation function with ConSmax

[0026] In the ConSmax-based self-attention mechanism, given an input sequence Its calculation formula can be defined as:

[0027]

[0028] Where, is the key vector dimension, Q is the query matrix, K is the key matrix, V is the value matrix, which are used as the incoming data, and T is the matrix transpose operator.

[0029] Step 2: Specify the method for calculating self-attention in blocks

[0030] The computation steps for block self-attention with the ConSmax activation function are as follows. Here, we assume that the total size of the three QKV matrices is [n,d], the output matrix O is of size [n,d], and its block sizes are all [d,d], where n is the sequence length and d is the number of columns in the QKV matrix:

[0031] 1) Calculate the transpose K of K T , the size becomes [d,n].

[0032] 2) Read the next Q block Q'.

[0033] 3) Read the next K T ,K block of V matrix T ′,V'.

[0034] 4) Initialize each element of the output matrix O' of size [d,d] to 0.

[0035] 5) Matrix multiplication: Calculate Q'×K T ′, and get the matrix S of size [d,d].

[0036] 6) Divide each element in S by

[0037] 7) Perform the ConSmax operation on each element in S.

[0038] 8) Matrix multiplication: Calculate S×V′ to obtain a matrix O” of size [d,d].

[0039] 9) Add each element of O” to O’.

[0040] 10) Repeat steps 3 to 9 times, and obtain a final result O' of [d,d].

[0041] 11) Concatenate the O' matrix to O.

[0042] 12) Reset K T ,V matrix block pointer.

[0043] 13) Repeat steps 2 to 12 times, and obtain the calculation result O of size [n,d].

[0044] Step 3: Implementing AIE processing units in the AIE array

[0045] The AIE processing unit is used to undertake the matrix multiplication operation and ConSmax operation in the self-attention. Each iteration can bear 4 accumulation loads of 8 degrees of parallelism, corresponding to the 8-way parallelism of step 13 and the 4-way parallelism of step 10 in claim 3. The design diagram of the AIE processing unit is shown in the attached figure. Figure 1 shown.

[0046] Step 3.1: Implementing an AIE single core

[0047] Each AIE core utilizes AIE Intrinsics instructions, using mac16 and mul16 instructions to perform d×d×d matrix multiplication operations within the core. The calculations are performed using the vector processor within the AIE core, forming an internal instruction pipeline with an iteration interval of 1. Matrix data is stored within a 4KB AIE window, which uses double buffering.

[0048] Step 3.2: Organize AIE multicore

[0049] Each AIE processing unit contains 64 AIE cores, 8 Q' input PLIO ports, 4 K T ′ input PLIO port, 4 V' input PLIO ports, and 8 O' output PLIO ports. The 64 cores of this AIE processing unit are divided into 8 groups, each with 8 AIE cores, where core [0,3] is responsible for parallel computing 4 Q'×K T ′ blocks and ConSmax output, core [4,7] is responsible for parallel calculation of S×V′ blocks and output.

[0050] Step 3.3: Data Flow Connection

[0051] For data stream connections, the 8 Q' input PLIO ports broadcast their data to the [4,7] (referring to cores 4 to 7) cores of the corresponding group AIE. T The ' input PLIO port and the four V' input PLIO ports are broadcast to the [0,3] cores of each AIE group. The output of the [0,3] cores of each AIE group is directly passed to the [4,7] cores. The [4,7] cores use a cascaded data flow to accumulate data in series. The output of core 7 (the last core) is connected to the O' output PLIO port of the corresponding group.

[0052] The operation process of the entire AIE processing unit is as follows:

[0053] 1) According to the above data flow connection method, the required matrix data is obtained from the input PLIO port connected to each AIE core.

[0054] 2) 8 groups of AIE [0,3] cores parallelly calculate Q'×K T ′, divide each element by And apply the ConSmax operation to obtain the S matrix.

[0055] 3) The [4,7] core receiving S matrix of 8 groups of AIE.

[0056] 4) 8 groups of AIEs’ [4,7] cores parallelly compute S×V′.

[0057] 5) Send the output of 8 groups of AIE7 cores to 8 O'PLIO ports.

[0058] Step 4: Implement the PL data engine

[0059] The PL data engine is responsible for providing efficient data supply to the AIE processing unit. The data engine consists of onboard DDR memory, DDR read module, data transmitter, data receiver, result accumulator, and DDR write module.

[0060] Step 4.1: Implement DDR Read Module

[0061] The data engine contains three DDR read modules, each of which occupies a dedicated DDR memory. The DDR read module accesses the DDR memory through the m_axi bus, stores the results in the on-chip BRAM, performs data multiplexing, and finally passes the data to the data transmitter.

[0062] The DDR read module of the Q matrix reads 8 [d,d]Q' blocks in parallel and in sequence. times, each read block will be reused times and pass it on to lower levels.

[0063] The DDR read module of the K matrix reads 4 [d,d]K' blocks in parallel and in sequence. The blocks do not use the multiplexing strategy and are read repeatedly. Next, pass it on to the lower level.

[0064] The DDR read module of the V matrix reads 4 [d,d]V' blocks in parallel and in sequence. The blocks do not adopt the multiplexing strategy and are read repeatedly. Next, pass it on to the lower level.

[0065] Step 4.2: Implement the data transmitter

[0066] The data transmitter can send [d,d] blocks in parallel to the 16 input ports of the AIE processing unit. The workflow is as follows:

[0067] 1) Streaming and parallel receiving of 16 Q'K'V' matrices from the DDR read module.

[0068] 2) Pack the data into the receiving format of the PLIO port.

[0069] 3) Send 16 Q'K'V' matrix data in parallel to the 16 input PLIOs of the AIE processing unit.

[0070] 4) Repeat steps 1 to 3 Second-rate.

[0071] Step 4.3: Implementing the Data Receiver

[0072] The data transmitter can receive [d,d] blocks in parallel from the eight output ports of the AIE processing unit. Its workflow is as follows:

[0073] 1) Streaming and parallel receiving of 8 O' matrices from the AIE processing unit output PLIO.

[0074] 2) Convert the data into PL internal data stream element format.

[0075] 3) Send 8 O' matrix data to the next level in parallel.

[0076] 4) Repeat steps 1 to 3 Second-rate.

[0077] Step 4.4: Implement the result accumulator

[0078] The result accumulator can accumulate 8 outputs in parallel, and its working process is as follows:

[0079] 1) Initialize the on-chip BRAM cache of size [8, d, d] to 0.

[0080] 2) Parallel reception The 8 O' matrices from the upper module are accumulated in the BRAM buffer each time they are received.

[0081] 3) Output the accumulated 8 [d,d] matrices.

[0082] 4) Repeat steps 1 to 3 Second-rate.

[0083] Step 4.5: Implementing the DDR Write Module

[0084] The data engine contains a DDR read module that occupies a dedicated DDR memory. The DDR write module accesses the DDR memory through the m_axi bus. The DDR write module of the O matrix reads 8 [d, d]O' blocks in parallel from the upper level and writes these blocks to the DDR memory. Second-rate.

[0085] Step 5: Deploy multiple self-attention processing modules

[0086] The self-attention module (ASA) can bear the complete self-attention load. The design diagram of the self-attention module is shown in the attached figure. Figure 2 As shown in the figure, it consists of an AIE processing unit and a data engine on the PL side. On the VCK5000 board, the hardware resources are sufficient to deploy 6 groups of such ASAs at the same time, so 6 groups of ASAs are deployed. The top-level architecture diagram of the accelerator is shown in the attached figure. Figure 3 shown.

[0087] The accelerator customized by the self-attention calculation method will be implemented as a whole in the Versal ACAP architecture, so that a single call from the host can complete six self-attention hardware operations in hardware.

Claims

1. An efficient self-attention inference method on the Versal ACAP architecture, characterized by: Firstly, a block-wise self-attention calculation method using ConSmax activation function is proposed, which replaces the Softmax activation function with the activation function and calculates the self-attention in blocks. Secondly, a self-attention hardware accelerator was designed and implemented on the Versal ACAP architecture. The AIE processing unit is implemented using the Versal ACAP's AIE array, and the data engine is implemented using the Versal ACAP's PL to provide data and scheduling for the AIE processing unit. On-chip DDR is used to store source data and results. The three together form the self-attention module (ASA), which is used to perform self-attention operations. Finally, multiple ASA designs are deployed into the hardware, and the CPU assigns tasks to them, ultimately realizing the calculation of self-attention.

2. The inference method according to claim 1, characterized in that The method of replacing the original self-attention Softmax activation function with ConSmax is as follows: Step 1: Replace the original self-attention Softmax activation function with ConSmax In the ConSmax-based self-attention mechanism, given an input sequence Its calculation formula is defined as: Where, is the key vector dimension, Q is the query matrix, K is the key matrix, V is the value matrix, which are used as the incoming data, and T is the matrix transpose operator.

3. The inference method according to claim 1, characterized in that The block-wise calculation method of self-attention is as follows: Step 2: Specify the method for calculating self-attention in blocks The calculation steps of block self-attention using ConSmax activation function are as follows. The total size of the three QKV matrices is set to [n, d], the size of the output matrix O is [n, d], and the block size is [d, d], where n is the sequence length and d is the number of columns in the QKV matrix: 1) Calculate the transpose K T , the size becomes [d,n]; 2) Read the next Q block Q'; 3) Read the next K T ,K block of V matrix T ′,V′; 4) Initialize each element of the output matrix O' of size [d, d] to 0; 5) Matrix multiplication: Calculate Q'×K T ′, and obtain the matrix S of size [d, d]; 6) Divide each element in S by 7) Perform ConSmax operation on each element in S; 8) Matrix multiplication: Calculate S×V′ to obtain the matrix O” of size [d, d]; 9) Add each element of O' to O'; 10) Repeat steps 3 to 9 times, and obtain a final result O' of [d,d]; 11) Splice O' matrix to O; 12) Reset K T ,V matrix block pointer; 13) Repeat steps 2) to 12) times, and obtain the calculation result O of size [n,d].

4. The inference method according to claim 1, characterized in that The AIE processing unit is implemented using the Versal ACAP's AIE array, specifically: Step 3: Implementing AIE processing units in the AIE array The AIE processing unit is used to undertake the matrix multiplication operation and the ConSmax operation in the self-attention, and each iteration can bear 4 accumulation loads of 8 degrees of parallelism, corresponding to the 8-way parallelism of step 13 and the 4-way parallelism of step 10 in claim 3; Step 3.1: Implementing an AIE single core For a single AIE core, each core uses AIE Intrinsics instructions, including mac16 and mul16 instructions, to perform d×d×d matrix multiplication operations within the core. The calculations are performed using the vector processor within the AIE core, forming an internal instruction pipeline with an iteration interval of 1. Matrix data is stored in a 4KB AIE window, which uses a double buffering mode. Step 3.2: Organize AIE multicore Each AIE processing unit contains 64 AIE cores, 8 Q' input PLIO ports, 4 K T ′ input PLIO port, 4 V' input PLIO ports, 8 O' output PLIO ports; the 64 cores of this AIE processing unit are divided into 8 groups, each with 8 AIE cores, where core [0,3] is responsible for parallel computing 4 Q'×K T ′ block and ConSmax output, core [4,7] is responsible for parallel calculation of S×V′ blocks and output; Step 3.3: Data Flow Connection For data stream connections, the 8 Q' input PLIO ports broadcast their data to the [4,7] (referring to cores 4 to 7) cores of the corresponding group AIE; the 4 K T The ' input PLIO port and the four V' input PLIO ports are broadcast to the [0,3] cores of each AIE group respectively; the output results of the [0,3] cores of each AIE group are directly passed to the [4,7] cores. The [4,7] cores use cascaded data flow to accumulate in series, and the output of core 7 (the last core) is connected to the O' output PLIO port of the corresponding group; The operation process of the entire AIE processing unit is as follows: 1) Obtain the required matrix data from the input PLIO port to which each AIE core is connected according to the above data flow connection method; 2) 8 groups of AIE [0,3] cores parallelly calculate Q'×K T ′, divide each element by And apply the ConSmax operation to obtain the S matrix; 3) [4,7] core receiving S matrix of 8 groups of AIE; 4) 8 groups of AIEs’ [4,7] cores parallelly compute S×V′; 5) Send the output of 8 groups of AIE7 cores to 8 O'PLIO ports.

5. The inference method according to claim 1, characterized in that The data engine based on the PL side is as follows: Step 4: Implement the PL Data Engine The PL data engine is responsible for providing efficient data supply to the AIE processing unit. The data engine consists of onboard DDR memory, DDR read module, data transmitter, data receiver, result accumulator, and DDR write module. Step 4.1: Implement DDR Read Module The data engine contains three DDR read modules, each of which occupies a dedicated DDR memory. The DDR read module accesses the DDR memory through the m_axi bus, stores the results in the on-chip BRAM, performs data multiplexing, and finally passes the data to the data transmitter. The DDR read module of the Q matrix reads 8 [d,d]Q' blocks in parallel and in sequence. times, each read block will be reused times, and pass it on to lower levels; The DDR read module of the K matrix reads 4 [d,d]K' blocks in parallel and in sequence, and reads repeatedly Second, pass it on to the lower level; The DDR read module of the V matrix reads 4 [d,d]V' blocks in parallel and in sequence, and reads repeatedly Second, pass it on to the lower level; Step 4.2: Implement the data transmitter The data transmitter sends the [d,d] blocks in parallel to the 16 input ports of the AIE processing unit. The workflow is as follows: 1) Streaming and parallel receiving of 16 Q'K'V' matrices from the DDR read module; 2) Pack the data into the receiving format of the PLIO port; 3) Send 16 Q'K'V' matrix data in parallel to the 16 input PLIOs of the AIE processing unit; 4) Repeat steps 1 to 3 Second-rate; Step 4.3: Implementing the Data Receiver The data transmitter receives the [d,d] blocks in parallel from the eight output ports of the AIE processing unit. The workflow is as follows: 1) Streaming and parallel receiving of 8 O' matrices from the AIE processing unit output PLIO; 2) Convert the data into PL internal data stream element format; 3) Send 8 O' matrix data to the next level in parallel; 4) Repeat steps 1 to 3 Second-rate; Step 4.4: Implement the result accumulator The result accumulator accumulates 8 outputs in parallel, and its working process is as follows: 1) Initialize the on-chip BRAM cache of size [8, d, d] to 0; 2) Parallel reception The 8 O' matrices from the upper module are accumulated in the BRAM buffer each time they are received; 3) Output the accumulated 8 [d,d] matrices; 4) Repeat steps 1 to 3 Second-rate; Step 4.5: Implementing the DDR Write Module The data engine contains a DDR read module that occupies a dedicated DDR memory. The DDR write module accesses the DDR memory through the m_axi bus. The DDR write module of the O matrix reads 8 [d, d]O' blocks in parallel from the upper level and writes these blocks to the DDR memory. Second-rate.

6. The inference method according to claim 1, characterized in that The self-attention module ASA design is as follows: Step 5: Deploy multiple self-attention processing modules The self-attention module (ASA) can bear the entire self-attention load and consists of an AIE processing unit and a data engine on the PL side. On the VCK5000 board, the hardware resources are sufficient to deploy six such ASAs at the same time, so six ASs are deployed.

7. The inference method according to claim 1, characterized in that The self-attention hardware accelerator designed and implemented on the Versal ACAP architecture is as follows: The accelerator customized by the self-attention calculation method will be implemented as a whole in the Versal ACAP architecture, so that a single call from the host can complete six self-attention hardware operations in hardware.