A high-energy-efficiency diversified hardware integrated computing architecture and method based on Transformer model

By adopting a computing allocation strategy of diversified hardware resources in the Transformer model, combining AI engine and programmable logic resources, the problem of fast inference and low power consumption when deploying Transformer-like intelligent models is solved, and an efficient, stable and reliable computing architecture is achieved.

CN118690788BActive Publication Date: 2025-05-13SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410770715.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2025-05-13
Estimated Expiration
2044-06-14

AI Technical Summary

Technical Problem

When deploying Transformer-like intelligent models, it is difficult to take into account the requirements of fast reasoning and low power consumption, especially in application scenarios where real-time response and resource-constrained.

Method used

We adopt a high-efficiency and diversified hardware integrated computing architecture based on the Transformer model, and combine AI engine (AIE) resources and programmable logic (PL) resources to deploy different computing modules separately to achieve efficient matrix multiplication operations and flexible logic operations.

Benefits of technology

It realizes significant performance advantages in Transformer model inference computing, reduces potential compatibility and interface problems, improves system stability and reliability, and meets the needs of real-time response and low power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118690788B_ABST
    Figure CN118690788B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-energy-efficiency diversified hardware integrated computing architecture and method based on a Transformer model, wherein the hardware integrated computing architecture includes two parts: an AI engine resource and a programmable logic resource; an AI engine dedicated data reading kernel, an AI engine multiplication module, an AIE-QK multiplication module, an AIE-SV multiplication module, an AIE-FC multiplication module, an AIE-FC2 multiplication module, and an AIE-FC3 multiplication module are deployed in the AI ​​engine resource; a PL pre-normalization kernel, a PL matrix information aggregation kernel, a PL-division kernel, a PL-SoftMax kernel, a PL pre-residual kernel, a PL post-normalization kernel, a PL post-residual kernel, and a PL data write operation kernel are deployed in the programmable logic resource. The computing architecture of the present invention shows significant performance advantages in Transformer model reasoning calculations, and reduces potential compatibility and interface problems, thereby improving the stability and reliability of the system, and can be widely used in the field of image processing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a high-energy-efficiency diversified hardware integrated computing architecture and method based on a Transformer model. Background Art

[0002] Transformer is a key technology for the development of intelligent large models in deep learning in recent years. The large-scale parallel training technology based on the Transformer architecture has rapidly increased the size of deep learning models, and has produced large models with significant effects such as the language processing model ChatGPT, LLaMA and the image processing model ViT. However, as the scale of the model grows, the inference speed has become a bottleneck in its application. The traditional deep learning models deployed on CPU and GPU platforms cannot take into account the requirements of fast inference and low power consumption in the face of real-time response and resource-constrained application scenarios.

[0003] As a programmable gate array device, FPGA takes into account parallel computing and low power consumption, and is the preferred device for deep learning model deployment. In recent years, there have been studies on the design methods of FPGA-based convolutional neural network (CNN) hardware acceleration systems. Some existing technical solutions add the convolution operation process to the pipeline to reduce the resource occupation of the operation data buffer, and design the pipeline storage structure to reduce the interaction time between the chip and the chip; however, traditional CNN is often limited by the computational complexity of convolution operations and local perception capabilities when processing image tasks. In contrast, Transformer, as a model based on the self-attention mechanism, has shown higher parallelism, global perception capabilities and adaptability in the image field, and can better capture long-range dependencies in images, while avoiding the limitations of parameter sharing and local perception in convolution operations. To this end, some other existing technical solutions analyze the calculation methods of various parts of the Vision Transformer encoder, design hardware circuits for each part, and work together with the processor to implement a dedicated hardware accelerator based on the Vision Transformer encoder, effectively solving the hardware deployment of Transformer-type intelligent models; however, this method mainly relies on a single type of programmable logic resources (PL) for deployment, and when processing data-intensive tasks such as Transformer-type intelligent models, it often cannot compete with high-frequency GPUs. In addition, the use of a single resource is not enough to meet the fast reasoning and low power consumption requirements of deep learning models in real-time response and resource-constrained application scenarios, and sometimes additional CPU support is required, which further limits the system's computing integration and convenience. Summary of the invention

[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the object of the present invention is to provide a high-energy-efficiency diversified hardware integrated computing architecture and method based on the Transformer model.

[0005] The technical solution adopted by the present invention is:

[0006] A high-efficiency and diversified hardware integrated computing architecture based on the Transformer model, including AI engine (AIE) resources and programmable logic (PL) resources;

[0007] The AI ​​engine resources are deployed with an AI engine dedicated data reading core, an AI engine multiplication module, an AIE-QK multiplication module, an AIE-SV multiplication module, an AIE-FC multiplication module, an AIE-FC2 multiplication module, and an AIE-FC3 multiplication module;

[0008] The programmable logic resources are deployed with a PL pre-normalization kernel, a PL matrix information aggregation kernel, a PL-division kernel, a PL-SoftMax kernel, a PL pre-residual kernel, a PL post-normalization kernel, a PL post-residual kernel, and a PL data write operation kernel.

[0009] Another technical solution adopted by the present invention is:

[0010] A design method for the above-mentioned high-energy-efficiency diversified hardware integrated computing architecture based on the Transformer model includes the following steps:

[0011] Read large-size raw matrix data and weight matrix data required by the AI ​​engine multiplication module through the AI ​​engine dedicated data reading kernel, and set the data flow to the AI ​​engine resources;

[0012] The PL pre-normalization kernel is used to normalize the input large-size raw matrix data.

[0013] Through the AI ​​engine multiplication module, matrix multiplication operations are performed on the incoming large-size raw matrix data and weight matrix data in the AI ​​engine resources of the FPGA;

[0014] Through the PL matrix information aggregation kernel, the Q matrix, K matrix and V matrix information are extracted and aggregated according to the position of the incoming large-size matrix data;

[0015] Interactively use AI engine resources and programmable logic resources to calculate the self-attention mechanism module on the Q matrix, K matrix, V matrix and weight matrix data to obtain the potential feature matrix data;

[0016] Through the PL pre-residual kernel, the incoming latent feature matrix data is added to the original large matrix data in the editable logic resource;

[0017] The PL post-normalization kernel is used to normalize the input large-size matrix data.

[0018] Through the AIE multi-layer perceptron module, two multiplication operations are performed in the AI ​​engine resources;

[0019] Through the PL post-residual kernel, the incoming pre-output feature matrix data and the latent feature matrix data are added;

[0020] Through the PL data write operation kernel, the data is saved to the hardware memory resources, waiting to be read or output.

[0021] Furthermore, the method of reading large-size original matrix data and weight matrix data required by the AI ​​engine multiplication module through the AI ​​engine dedicated data reading kernel and setting the data flow into the AI ​​engine resources includes:

[0022] Configure the memory access address generator: traverse the large-size original matrix data and the weight matrix data required by each AI engine multiplication module through a multi-dimensional iterative control structure, calculate the memory address of each element, and output these addresses through the AXI output stream interface; the address generator optimizes the efficiency of data access and provides the necessary address information for data loading operations, thereby supporting high-speed data processing;

[0023] Configure the data loader: Use the AXI input stream interface to receive the memory access address from the AXI output stream interface, and load the large-size original matrix data and the weight matrix data required by each AI engine multiplication module from the DDR memory according to the memory access address;

[0024] Configure the data outputter: Write the read large-size original matrix data and the weight matrix data required by each AI engine multiplication module into different AXI stream interfaces to achieve continuous data stream output.

[0025] Furthermore, the PL pre-normalization kernel is used to perform normalization calculation on the input large-size original matrix data, including:

[0026] Use the AXI input stream interface to receive the large-size original matrix data Ori_x and the normalized hyperparameter data NORM_w1 in the weight matrix data, and perform normalization calculations to obtain the normalized large-size original matrix data stream NORM_x and write it into the AXI stream output interface.

[0027] Furthermore, the AI ​​engine multiplication module performs matrix multiplication operations on the incoming large-size original matrix data and weight matrix data in the AI ​​engine resources, including:

[0028] Configure the matrix block kernel: Use the AXI input stream interface in the AI ​​engine resource to receive the normalized large-size raw matrix data stream NORM_x, traverse the raw matrix data stream through the multi-dimensional iteration control structure, divide the incoming large-size raw matrix data into blocks, obtain several small-size matrices, and the size of each small-size matrix does not exceed the computing resource limit of the AI ​​engine, and write them to the AXI stream output interface;

[0029] Configure the matrix multiplication kernel: Use the AXI input stream interface in the AI ​​engine resource to receive the small-size matrix data stream and the weight matrix data from the AXI output stream interface, use the AI ​​engine resource to perform the multiplication operation of the two matrices, and write to the AXI stream output interface;

[0030] Configure the matrix merge kernel: use the AXI input stream interface to receive the multiplication result data stream, traverse the multidimensional iteration control structure in reverse, merge the small-size matrix data streams into a large-size matrix data stream QKV_x, and write it to the AXI stream output interface.

[0031] Furthermore, the PL matrix information aggregation kernel extracts and aggregates Q matrix, K matrix and V matrix information of the incoming large-size matrix data by position, including:

[0032] Use the AXI input stream interface to receive the large-size matrix data stream QKV_x. According to the requirements of the attention mechanism, the Q matrix, K matrix and V matrix are accurately extracted by position. Then, according to the multi-head mechanism, the Q matrix, K matrix and V matrix are accurately extracted into n small-size matrices by position and written into the AXI stream output interface respectively.

[0033] Furthermore, the interaction uses AI engine resources and programmable logic resources to perform calculations of the self-attention mechanism module on the Q matrix, K matrix, V matrix and weight matrix data to obtain potential feature matrix data, including:

[0034] Configure the AIE-QK multiplication module, use the AXI input stream interface to receive the Q matrix data and the K matrix data, use the same structure as the AI ​​engine multiplication module, perform block division, multiplication and merging of the Q matrix and the K matrix in the AI ​​engine resources, obtain the first matrix data, and write the result to the AXI stream output interface;

[0035] Configure the PL-division kernel, use the AXI input stream interface to receive the first matrix data, and perform a division operation with the vector dimension parameter SCALE_dim in the weight matrix data in the programmable logic resource to obtain the second matrix data, and write the result to the AXI stream output interface;

[0036] Configure the PL-SoftMax kernel, use the AXI input stream interface to receive the second matrix data, perform exponential operations in the programmable logic resources, and calculate the sum of all exponential values, then each exponential value will be normalized by its sum, convert the input real number output into a probability distribution, and get the S matrix data and write it to the AXI stream output interface;

[0037] Configure the AIE-SV multiplication module, use the AXI input stream interface to receive the S matrix data and the V matrix data, use the same structure as the AI ​​engine multiplication module, perform block, multiplication and merging of the S matrix and the K matrix in the AI ​​engine resources, obtain the R matrix data, and write the resulting R matrix data to the AXI stream output interface;

[0038] Configure the AI ​​engine FC multiplication module, use the AXI input stream interface to receive the R matrix data and the PROJ_w weight matrix in the weight matrix data, use the same structure as the AI ​​engine multiplication module, perform block, multiplication and merging of the R matrix and the PROJ_w weight matrix in the AI ​​engine resources, and obtain the potential feature matrix data HID_x and write it to the AXI stream output interface.

[0039] Further, the step of performing an addition operation on the incoming potential feature matrix data and the original large matrix data in the editable logic resource by using the PL post-normalization kernel and the PL pre-residual kernel includes:

[0040] Configure the PL front residual core, use the AXI input stream interface to receive the potential feature matrix data HID_x and the large-size original matrix data Ori_x, perform addition operations in the programmable logic resources, and obtain the attention matrix data stream ATEN_x and write it into the AXI stream output interface;

[0041] The normalization calculation of the input large-size matrix data includes:

[0042] Use the AXI input stream interface to receive the attention matrix data ATEN_x and the normalized hyperparameter data NORM_w2 in the weight matrix data, and perform normalization calculations to obtain the perception matrix data stream MID_x and write it to the AXI stream output interface.

[0043] Furthermore, the AIE multilayer perceptron module is used to perform two multiplication operations in the AI ​​engine resources, including:

[0044] Configure the AIE-FC2 multiplication module, use the AXI input stream interface to receive the perception matrix data stream MID_x and the MLP_w1 weight matrix in the weight matrix data, use the same structure as the AI ​​engine multiplication module, perform block, multiplication and merging of the perception matrix and the weight matrix in the AI ​​engine resources, and perform GELU operations to obtain the FC2 matrix data FC2_x and write it to the AXI stream output interface;

[0045] Configure the AIE-FC4 multiplication module, use the AXI input stream interface to receive the FC2 matrix data FC2_x and the MLP_w2 weight matrix in the weight matrix data, use the same structure as the AI ​​engine multiplication module, perform block, multiplication and merging of the perception matrix and the weight matrix in the AI ​​engine resources, and obtain the FC3 matrix data FC3_x and write it to the AXI stream output interface.

[0046] Furthermore, the addition operation of the input pre-output feature matrix data and the potential feature matrix data through the PL post-residual kernel includes:

[0047] Configure the post-residual core, use the AXI input stream interface to receive the FC3 matrix data FC3_x and the potential feature matrix data HID_x, perform addition operations in the programmable logic resources, and write the final result TRAN_x to the AXI stream output interface.

[0048] The beneficial effect of the present invention is as follows: the present invention implements a computing allocation strategy based on the characteristics of diversified hardware resources for each computing module of the Transformer model. Each module is deployed to AI engine (AIE) resources or programmable logic (PL) resources according to its different computing complexity, forming a high-energy-efficiency computing architecture. This diversified hardware integrated computing architecture shows significant performance advantages in Transformer model reasoning calculations, and reduces potential compatibility and interface problems, thereby improving the stability and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1 is a schematic diagram of a high-energy-efficiency diversified hardware integrated computing architecture based on a Transformer model in an embodiment of the present invention;

[0051] Figure 2 is a design flow chart of a high-energy-efficiency diversified hardware integrated computing architecture based on the Transformer model in an embodiment of the present invention;

[0052] Figure 3 It is a schematic diagram of the definition of the AXI output stream interface of the AI ​​engine dedicated data reading core in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0054] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0055] In the description of the present invention, the meaning of "several" is one or more, the meaning of "more" is two or more, and the meanings of "greater than", "less than", "exceed" and the like are understood as not including the number itself, and the meanings of "above", "below", "within" and the like are understood as including the number itself. If there is a description of the first and the second, it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, "and / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0056] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0057] Terminology explanation:

[0058] DDR memory: The full name is DDR SDRAM, which stands for Double Data Rate SDRAM.

[0059] Existing methods mainly deploy Transformer models based on a single type of programmable logic (PL) resources. Although FPGA provides certain programmability and flexibility, its performance is often insufficient to meet the requirements of efficient processing when performing deep learning tasks, especially in a high-frequency GPU competition environment. In addition, in order to achieve the required computing power, additional CPU or other external hardware support is often required, which not only increases the complexity of the system, but also may affect the overall system stability and reliability. Based on this, the present invention deploys the Transformer model using a high-efficiency computing architecture that integrates AI engine (AIE) resources and programmable logic (PL) resources. The architecture integrates the processor, AI engine and programmable logic PL on a single chip, providing a one-stop hardware solution. In particular, for the attention mechanism (Attention) module with large parameter volume and computational complexity in the Transformer model, the new AI engine (AIE) hardware platform can realize efficient matrix multiplication operations. At the same time, the programmable logic (PL) hardware platform can flexibly process logical operations other than matrix multiplication. This diversified hardware integrated computing architecture shows significant performance advantages in Transformer model reasoning calculations, and reduces potential compatibility and interface problems, thereby improving the stability and reliability of the system.

[0060] like Figure 1 As shown, the embodiment of the present invention proposes a high-energy-efficiency diversified hardware integrated computing architecture based on the Transformer model. For each computing module of the Transformer model, we implement a computing allocation strategy based on the characteristics of diversified hardware resources. Each module is deployed to the AI ​​engine (AIE) resource or the programmable logic (PL) resource according to its different computing complexity to form a high-energy-efficiency computing architecture. The AI ​​engine (AIE) resource is respectively deployed with the AI ​​engine dedicated data reading kernel, the AI ​​engine multiplication module, the AIE-QK multiplication module, the AIE-SV multiplication module, the AIE-FC multiplication module, the AIE-FC2 multiplication module, and the AIE-FC3 multiplication module; in the programmable logic resource (PL), the PL pre-normalization kernel, the PL matrix information aggregation kernel, the PL-division kernel, the PL-SoftMax kernel, the PL pre-residual kernel, the PL post-normalization kernel, the PL post-residual kernel, and the PL data write operation kernel are respectively deployed.

[0061] For the above hardware integrated computing architecture, the specific design steps are as follows: Figure 2As shown. The first step is to read the original matrix data into the hardware resources of the FPGA through the AI ​​engine dedicated data reading kernel, and set the data stream to flow into the AI ​​engine resources. The second step is to design the PL pre-normalization kernel to perform normalization calculations on the input large-size original matrix data. The third step is to design the AI ​​engine multiplication module, which includes the matrix block kernel, the matrix multiplication kernel, and the matrix merge kernel. This module divides the incoming large-size original matrix data into blocks to obtain multiple small-size matrices, and then performs multiplication operations separately. Finally, the results of multiple small-size matrix multiplications are merged to obtain the original-size matrix data, and the data stream is set to flow into the matrix information aggregation kernel. The fourth step is to extract and aggregate the Q matrix, K matrix, and V matrix information of the incoming large-size matrix data by position in the PL matrix information aggregation kernel, and set the data stream to flow into the self-attention module. The fifth step is to design the AIE-PL-AIE self-attention module, which includes the AI ​​engine QK multiplication module, PL division kernel, PLSoftMax division kernel, AI engine SV multiplication module, and AI engine FC multiplication module. This module first performs matrix block, multiplication and matrix merging on the incoming Q matrix and K matrix in the AI ​​engine QK multiplication module, and then enters the PL division kernel through the data stream to perform division operation with the vector dimension parameter, and then enters the PL-SoftMax division kernel through the data stream for SoftMax operation, and then enters the AI ​​engine SV multiplication module and V matrix for block, multiplication and matrix merging, and finally enters the AI ​​engine FC multiplication module and weight matrix for block, multiplication and matrix merging, and sets the data stream to enter the pre-residual kernel and post-residual kernel respectively. The sixth step is to perform addition operation on the incoming potential feature matrix data and the original matrix data in the PL pre-residual kernel, and set the data stream to enter the normalization unit. In the seventh step, normalize the data including mean, variance calculation and update operation in the PL post-normalization unit. In the eighth step, design the AI ​​engine multilayer perceptron module, which includes two AI engine multiplication modules, perform two matrix block divisions, multiplication operations and matrix merging, and perform a GELU operation between the two modules, and finally pass it to the post-residual kernel. In the ninth step, add the incoming pre-output feature matrix data and the potential feature matrix data in the PL post-residual kernel, and set the data flow to flow into the data write operation kernel. In the tenth step, save the data to the hardware memory resources through the PL data write operation kernel, waiting to be read or output.

[0062] The following is a detailed explanation of each step in the above design method.

[0063] S1: Reads large-size raw matrix data and weight matrix data required by each AI engine multiplication module from the FPGA off-chip storage unit through the AI ​​engine dedicated data reading core, and sets up AXI data streams to connect to each AI engine multiplication module.

[0064] As an optional implementation, step S1 specifically includes the following steps:

[0065] S1-1: Configure the memory access address generator. Through the multi-dimensional iterative control structure, it traverses the large-size original matrix data and the weight matrix data required by each AI engine multiplication module, accurately calculates the memory address of each element, and outputs these addresses through the AXI stream interface. This address generator optimizes the efficiency of data access and provides the necessary address information for data loading operations, thereby supporting high-speed data processing.

[0066] S1-2: Configure the data loader. Use the AXI input stream interface to receive the memory access address from the AXI output stream interface in step S1-1, and load the large-size original matrix data and the weight matrix data required by each AI engine multiplication module from the DDR memory according to the address.

[0067] S1-3: Configure the data output device. It is used to write the large-size original matrix data read in step S1-2 and the weight matrix data required by each AI engine multiplication module into different AXI stream interfaces to achieve continuous data stream output. Among them, the AXI stream output interfaces corresponding to the large-size original matrix data stream Ori_x and each weight matrix data stream NORM_w1, QKV_w, SCALE_dim, PROJ_w, NORM_w2, MLP_w1, and MLP_w2 are as follows: Figure 3 shown.

[0068] S2: Through the PL pre-normalization kernel, the input large-size original matrix data is normalized including mean, variance calculation and update operations.

[0069] Specifically, step S2 includes the following steps:

[0070] S2-1: Use the AXI input stream interface to receive the large-size original matrix data Ori_x and normalized hyperparameter data NORM_w1 of step S1-3 and perform normalization calculations. The specific operations include calculating the mean and variance of the data, and updating the data values ​​according to these statistics, so that the data set is normalized to a state with zero mean and unit variance. The normalized matrix data stream NORM_x is written to the AXI stream output interface.

[0071] The operation process of LayerNorm is as follows: first, the mean and variance of the input tensor along the feature dimension are calculated; then the input tensor is normalized by subtracting the mean and dividing it by the standard deviation; then a learnable scaling factor and offset are introduced to scale and translate the normalized tensor; finally, the normalized, scaled and translated tensor is output.

[0072]

[0073] Where x is the input feature vector, μ is the mean of the feature vector, σ is the standard deviation of the feature vector, γ and β are the learned scaling and translation coefficients, and ε is a very small number for stable calculation.

[0074] S3: Through the AI ​​engine multiplication module, matrix multiplication operations are performed on the incoming large-size original matrix data and weight matrix data in the AI ​​engine resources of the FPGA.

[0075] As an optional implementation, step S3 specifically includes the following steps:

[0076] S3-1: Configure the matrix blocking kernel, use the AXI input stream interface in the AI ​​engine resources to receive the large-size original matrix data stream NORM_x normalized in step S2-1, and traverse the original matrix data stream again through the multi-dimensional iterative control structure to block the incoming large-size original matrix data into several small-size matrices, and the size of each small-size matrix does not exceed the computing resource limit of the AI ​​engine, and write it to the AXI stream output interface.

[0077] S3-2: Configure the matrix multiplication kernel, use the AXI input stream interface in the AI ​​engine resources to receive the small-size matrix data stream from the AXI output stream interface of step S3-1, and the QKV_w weight matrix data stream received from the AXI output stream interface of step S1-3, and make full use of the AI ​​engine resources to perform multiplication operations on the two matrices and write them into the AXI stream output interface.

[0078] S3-3: Configure the matrix merging kernel, use the AXI input stream interface to receive the multiplication result data stream of step S3-2, and reversely traverse the multidimensional iteration control structure in step S3-1, merge the small-size matrix data streams into a large-size matrix data stream QKV_x, and write it to the AXI stream output interface.

[0079] S4: Through the PL matrix information aggregation kernel, Q matrix, K matrix and V matrix information are extracted and aggregated by position.

[0080] Specifically, step S4 includes:

[0081] S4-1: Use the AXI input stream interface to receive the large-size matrix data stream QKV_x output by step S3-3, and according to the requirements of the attention mechanism, accurately extract the Q matrix, K matrix and V matrix by position, and then according to the multi-head mechanism, accurately extract n small-size matrices from the Q matrix, K matrix and V matrix by position, and write them into the AXI stream output interface respectively.

[0082] S5: Through the AIE-PL-AIE self-attention module, the AI ​​engine and PL resources of the FPGA are interactively used to calculate the self-attention mechanism module for the incoming QKV matrix and weight matrix data, including the AI ​​engine QK multiplication module, PL division kernel, PLSoftMax division kernel and AI engine SV multiplication module.

[0083] As an optional implementation, step S5 specifically includes the following steps:

[0084] S5-1: Configure the AI ​​engine QK multiplication module, use the AXI input stream interface to receive the Q matrix data and K matrix data of step S4-1, adopt the same structure as the AI ​​engine multiplication module of step S3, perform block, multiplication and merging of the Q matrix and the K matrix in the AI ​​engine resources, and write the results to the AXI stream output interface.

[0085] S5-2: Configure the PL division kernel, use the AXI input stream interface to receive the matrix data of step S5-1 and the vector dimension parameter SCALE_dim of step S1, perform division operations in the PL resources, and write the results to the AXI stream output interface.

[0086] S5-3: Configure the PL-SoftMax division kernel, use the AXI input stream interface to receive the data of step S5-2, perform exponential operation in the PL resource, and calculate the sum of all exponential values, then each exponential value will be normalized by its sum, convert the input real number output into a probability distribution, and get the S matrix data written to the AXI stream output interface.

[0087] The operation process of Softmax is as follows: first, given an input vector X with a length of N; calculate the value of the exponential function: perform an exponential operation on each element in the input vector X to obtain a new vector E; calculate the exponential sum: sum all elements in vector E to obtain the sum S; calculate the probability distribution: divide each element in vector E by the sum S to obtain a probability vector P; output the probability vector P, in which each element represents the probability of the corresponding category, and its calculation formula is as follows:

[0088]

[0089] Among them, x i is the i-th element of the input vector X.

[0090] S5-4: Configure the AI ​​engine SV multiplication module, use the AXI input stream interface to receive the S matrix data of step S5-3 and the V matrix data of step S4-1, adopt the same structure as the AI ​​engine multiplication module of step S3, perform block, multiplication and merging of the S matrix and the K matrix in the AI ​​engine resources, and write the resulting R matrix data to the AXI stream output interface.

[0091] S5-5: Configure the AI ​​engine FC multiplication module, use the AXI input stream interface to receive the R matrix data of step S5-4 and the PROJ_w weight matrix data stream received from the step S1-3 interface, use the same structure as the AI ​​engine multiplication module of step S3, perform block, multiplication and merging of the R matrix and the PROJ_w weight matrix in the AI ​​engine resources, and obtain the potential feature matrix data HID_x and write it to the AXI stream output interface.

[0092] S6: Through the PL front residual core, the incoming potential feature matrix data and the original large matrix data are added in the PL resources of the FPGA.

[0093] Specifically, step S6 includes:

[0094] S6-1: Configure the front residual kernel, use the AXI input stream interface to receive the potential feature matrix data HID_x of step S5-4 and the large-size original matrix data Ori_x of step S1, perform addition operations in the PL resources, and obtain the attention matrix data stream ATEN_x to write into the AXI stream output interface.

[0095] S7: Through the PL post-normalization kernel, the input large-size matrix data is normalized including mean, variance calculation and update operation.

[0096] Specifically, step S7 includes:

[0097] S7-1: Use the AXI input stream interface to receive the attention matrix data ATEN_x and normalized hyperparameter data NORM_w2 of step S6-1, and perform normalization calculations. The specific operations include calculating the mean and variance of the data, and updating the data values ​​according to these statistics, so that the data set is normalized to a state with zero mean and unit variance, and the perception matrix data stream MID_x is obtained and written to the AXI stream output interface.

[0098] S8: Perform two multiplication operations in the AI ​​engine resources of the FPGA through the AIE multi-layer perceptron module.

[0099] As an optional implementation, step S8 specifically includes the following steps:

[0100] S8-1: Configure the AIE-FC2 multiplication module, use the AXI input stream interface to receive the perception matrix data stream MID_x of step S7-1 and the weight matrix data MLP_w1 of step S1, adopt the same structure as the AI ​​engine multiplication module of step S3, perform block, multiplication and merging of the perception matrix and the weight matrix in the AI ​​engine resources, and perform GELU operations to obtain the FC2 matrix data FC2_x and write it to the AXI stream output interface.

[0101] S8-2: Configure the AIE-FC4 multiplication module, use the AXI input stream interface to receive the FC2 matrix data FC2_x of step S8-1 and the weight matrix data MLP_w2 of step S1, use the same structure as the AI ​​engine multiplication module of step S3, perform block, multiplication and merging of the perception matrix and the weight matrix in the AI ​​engine resources, and obtain the FC3 matrix data FC3_x and write it to the AXI stream output interface.

[0102] S9: Perform addition operation on the incoming pre-output feature matrix data and the latent feature matrix data through the PL post-residual kernel.

[0103] Specifically, step S9 includes:

[0104] S9-1: Configure the post-residual kernel, use the AXI input stream interface to receive the FC3 matrix data FC3_x of step S8-2 and the potential feature matrix data HID_x of step S5-5, perform addition operation in the PL resources, and write the final result TRAN_x to the AXI stream output interface.

[0105] S10: The PL data write operation kernel saves the data to the hardware memory resource, waiting to be read or output.

[0106] As an optional implementation, step S10 specifically includes the following steps:

[0107] S10-1: Configure the memory access address generator. Traverse the large-size matrix data stream through the multi-dimensional iteration control structure, accurately calculate the memory address of each element, and output these addresses through the AXI stream interface.

[0108] S10-2: Configure the data writer. Use the AXI stream interface to receive the matrix data stream TRAN_x output from step S9-1, use the AXI input stream interface to receive the memory access address from the AXI output stream interface of step S10-1, and write the large-size matrix data into the DDR memory according to the address.

[0109] In summary, compared with the prior art, the present invention has at least the following advantages and beneficial effects: the present invention makes full use of the efficient data processing capability of the AI ​​engine (AIE) to process tasks that require high parallel processing and data throughput matrix multiplication operations. The AI ​​engine (AIE) is a key component in devices such as Xilinx's Versal ACAP. It consists of multiple high-performance vector computing units, each of which includes a vector processor and a set of memories. It is highly programmable and can be used exclusively for high-throughput, low-latency matrix multiplication operations by writing specific instruction streams to define computing tasks and data flows. It is very suitable for the application requirements of deep learning models. Furthermore, each weight matrix data stream will be cached in the FPGA resource at the beginning of the operation, waiting for the reading and calling of each module, which can reduce the data loading time. Finally, the flexible programmability of PL is used to process logical operations other than matrix multiplication, and the two achieve low-latency collaborative processing to form a diversified hardware integrated computing architecture. This strategy ensures that each computing module runs in the most suitable hardware environment, thereby optimizing overall performance, reducing power consumption, and increasing processing speed, speeding up the reasoning of complex Transformer-type intelligent models, achieving fast reasoning in resource-constrained application scenarios, and meeting real-time response and low power consumption requirements.

[0110] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0111] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0112] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A design method for a high-energy-efficiency diversified hardware integrated computing architecture based on a Transformer model, characterized in that: The energy-efficient and diversified hardware integrated computing architecture includes two parts: AI engine resources and programmable logic resources; The AI ​​engine resources are deployed with an AI engine dedicated data reading core, an AI engine multiplication module, an AIE-QK multiplication module, an AIE-SV multiplication module, an AIE-FC multiplication module, an AIE-FC2 multiplication module, and an AIE-FC3 multiplication module; The programmable logic resources are deployed with a PL pre-normalization kernel, a PL matrix information aggregation kernel, a PL-division kernel, a PL-SoftMax kernel, a PL pre-residual kernel, a PL post-normalization kernel, a PL post-residual kernel, and a PL data write operation kernel; The design method comprises the following steps: Through the AI ​​engine dedicated data reading core, read the large-size raw matrix data and the weight matrix data required by the AI ​​engine multiplication module; The PL pre-normalization kernel is used to normalize the input large-size raw matrix data. Through the AI ​​engine multiplication module, matrix multiplication operations are performed on the incoming large-size raw matrix data and weight matrix data in the AI ​​engine resources; Through the PL matrix information aggregation kernel, the Q matrix, K matrix and V matrix information are extracted and aggregated according to the position of the incoming large-size matrix data; Interactively use AI engine resources and programmable logic resources to calculate the self-attention mechanism module on the Q matrix, K matrix, V matrix and weight matrix data to obtain the potential feature matrix data; Through the PL pre-residual kernel, the incoming latent feature matrix data is added to the original large matrix data in the editable logic resource; The PL post-normalization kernel is used to normalize the input large-size matrix data. Through the AIE multi-layer perceptron module, two multiplication operations are performed in the AI ​​engine resources; Through the PL post-residual kernel, the incoming pre-output feature matrix data and the latent feature matrix data are added; Through the PL data write operation kernel, the data is saved to the hardware memory resource, waiting to be read or output; The method of reading large-size raw matrix data and weight matrix data required by the AI ​​engine multiplication module through the AI ​​engine dedicated data reading kernel includes: Traverse the large-size original matrix data and the weight matrix data required by each AI engine multiplication module, calculate the memory address of each element, and output these addresses through the AXI output stream interface; Receive the memory access address from the AXI output stream interface, and load the large-size original matrix data and the weight matrix data required by each AI engine multiplication module from the memory according to the memory access address; Write the read large-size raw matrix data and the weight matrix data required by each AI engine multiplication module into different AXI stream interfaces to achieve continuous data stream output; The PL pre-normalization kernel is used to perform normalization calculation on the input large-size original matrix data, including: Receive the large-size original matrix data Ori_x and the normalized hyperparameter data NORM_w1 in the weight matrix data, and perform normalization calculation to obtain the normalized large-size original matrix data stream NORM_x; The AI ​​engine multiplication module performs matrix multiplication operations on the incoming large-size original matrix data and weight matrix data in the AI ​​engine resources, including: Receive the normalized large-size original matrix data stream NORM_x, traverse the original matrix data stream through the multi-dimensional iteration control structure, divide the incoming large-size original matrix data into blocks, and obtain several small-size matrices; Receive small-size matrix data stream and weight matrix data, and use AI engine resources to perform multiplication of the two matrices; Receive the multiplication result data stream, traverse the multidimensional iteration control structure in reverse, merge the small-size matrix data streams into a large-size matrix data stream QKV_x, and write it to the AXI stream output interface.

2. The design method according to claim 1, characterized in that: The PL matrix information aggregation kernel extracts and aggregates the Q matrix, K matrix and V matrix information of the incoming large-size matrix data by position, including: Receive the large-size matrix data stream QKV_x, and extract the Q matrix, K matrix, and V matrix accurately by position according to the requirements of the attention mechanism. Then, according to the multi-head mechanism, extract n small-size matrices from the Q matrix, K matrix, and V matrix accurately by position and write them into the AXI stream output interface respectively.

3. The design method according to claim 2, characterized in that: The interaction uses AI engine resources and programmable logic resources to calculate the self-attention mechanism module on the Q matrix, K matrix, V matrix and weight matrix data to obtain potential feature matrix data, including: Configure the AIE-QK multiplication module, receive the Q matrix data and the K matrix data, perform block division, multiplication operation, and merging of the Q matrix and the K matrix in the AI ​​engine resources, and obtain the first matrix data; Configure the PL-division kernel, receive the first matrix data, and perform a division operation with the vector dimension parameter SCALE_dim in the weight matrix data in the programmable logic resource to obtain the second matrix data; Configure the PL-SoftMax kernel, receive the second matrix data, perform exponential operation in the programmable logic resources, and calculate the sum of all exponential values, then each exponential value will be normalized by its sum, convert the input real number output into a probability distribution, and obtain the S matrix data; Configure the AIE-SV multiplication module, receive the S matrix data and the V matrix data, perform block division, multiplication and merging of the S matrix and the V matrix in the AI ​​engine resources, and obtain the R matrix data; Configure the AI ​​engine FC multiplication module, receive the R matrix data and the PROJ_w weight matrix in the weight matrix data, perform block division, multiplication and merging of the R matrix and the PROJ_w weight matrix in the AI ​​engine resources, and obtain the potential feature matrix data HID_x.

4. The design method according to claim 3, characterized in that: The step of performing an addition operation on the incoming potential feature matrix data and the original large matrix data in the editable logic resource by using the PL post-normalization kernel and the PL pre-residual kernel includes: Configure the PL front residual core, receive the latent feature matrix data HID_x and the large-size original matrix data Ori_x, perform addition operations in the programmable logic resources, and obtain the attention matrix data stream ATEN_x; The normalization calculation of the input large-size matrix data includes: Receive the attention matrix data ATEN_x and the normalized hyperparameter data NORM_w2 in the weight matrix data, and perform normalization calculations to obtain the perception matrix data stream MID_x.

5. The design method according to claim 4, characterized in that: The AIE multi-layer perceptron module performs two multiplication operations in the AI ​​engine resources, including: Configure the AIE-FC2 multiplication module, receive the perception matrix data stream MID_x and the MLP_w1 weight matrix in the weight matrix data, perform block division, multiplication and merging of the perception matrix and the weight matrix in the AI ​​engine resources, and perform GELU operations to obtain the FC2 matrix data FC2_x; Configure the AIE-FC4 multiplication module, receive the FC2 matrix data FC2_x and the MLP_w2 weight matrix in the weight matrix data, perform block division, multiplication, and merging of the perception matrix and the weight matrix in the AI ​​engine resources, and obtain the FC3 matrix data FC3_x.

6. The design method according to claim 5, characterized in that: The method of performing an addition operation on the incoming pre-output feature matrix data and the potential feature matrix data through the PL post-residual kernel includes: Configure the post-residual core, receive the FC3 matrix data FC3_x and the potential feature matrix data HID_x, perform addition operations in the programmable logic resources, and write the final result TRAN_x to the AXI stream output interface.