A visual Transformer accelerator implementation method and system

By adopting a ring-shaped reconfigurable architecture and pipeline architecture in hardware accelerators, the design of visual Transformer accelerators solves the problems of computing bottlenecks, high energy consumption and insufficient flexibility when dealing with visual Transformer models, and realizes efficient acceleration and low power consumption computer vision technology.

CN119692408BActive Publication Date: 2025-05-23SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510199558.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-23
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing hardware accelerators face problems such as computing bottlenecks, high energy consumption and insufficient flexibility when dealing with visual Transformer models, making it difficult to effectively accelerate the execution speed of the model and reduce power consumption.

Method used

The visual Transformer accelerator is designed using a ring-shaped reconfigurable architecture, and multiple Transformer acceleration cores, storage units and shared memory are formed into a ring topology through the AXI bus, realizing unified management and flexible execution of path configuration. At the same time, the pipeline architecture is used to build the Transformer calculation unit, multiplex matrix multiplication and layer normalization subunits to improve resource utilization and computing efficiency.

Benefits of technology

It significantly improves the execution speed of the visual Transformer model, reduces the overall power consumption of the accelerator, and provides developers with more innovative space to promote the widespread application of computer vision technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119692408B_ABST
    Figure CN119692408B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for realizing a visual Transformer accelerator, which belongs to the field of computer vision technology. The visual Transformer accelerator is realized based on a ring-shaped reconfigurable architecture, wherein a plurality of Transformer acceleration cores, storage units, and shared memories are combined into a ring topology through an AXI bus; a pipeline architecture is used to construct a Transformer computing unit, which includes a multi-head attention unit and a multi-layer perceptron unit; the multi-head attention unit and the multi-layer perceptron unit reuse a matrix multiplication subunit and a layer normalization unit to realize data normalization and multiplication and addition operations. The present invention can not only greatly improve the execution speed of the visual Transformer model, but also provide developers with more room for innovation, and promote the development of computer vision technology to a wider range of application fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for realizing a visual Transformer accelerator. Background Art

[0002] In recent years, a new neural network architecture called Transformer has attracted widespread attention in the field of deep learning. Initially, Transformer was designed for natural language processing (NLP) tasks, and quickly became a mainstream model in this field due to its excellent parallelization capabilities and effective processing capabilities for long sequence data. As research deepened, scientists discovered that Transformer is also suitable for visual tasks, and thus the Vision Transformer (ViT) was born, which has demonstrated unprecedented performance in computer vision tasks such as image classification, object detection, and semantic segmentation.

[0003] However, despite ViT's excellent performance in terms of accuracy, its computational complexity and resource consumption have also increased significantly. Currently, in order to accelerate the training and reasoning process of deep learning models, there are many types of hardware accelerators on the market, including graphics processing units (GPUs), tensor processing units (TPUs), etc. These accelerators have their own characteristics, but they face the following challenges when facing ViT:

[0004] (1) Computational bottleneck: Existing hardware accelerators are usually optimized for convolutional neural networks (CNNs) and lack support for the multi-head attention (MHA) operation in ViT, resulting in low computational efficiency.

[0005] (2) Energy consumption problem: Continuous high-load computing makes the power consumption of existing hardware accelerators high, which is not conducive to the application of mobile devices or edge computing scenarios.

[0006] (3) Lack of flexibility: Although some accelerators provide a certain degree of programmability, they have poor adaptability to rapidly changing algorithm structures and emerging optimization techniques. Summary of the invention

[0007] The technical task of the present invention is to address the above shortcomings and provide a visual Transformer accelerator implementation method and system, which can not only greatly improve the execution speed of the visual Transformer model, but also provide developers with more room for innovation and promote the development of computer vision technology into a wider range of application fields.

[0008] The technical solution adopted by the present invention to solve its technical problem is:

[0009] A method for implementing a visual Transformer accelerator is provided. The method implements the visual Transformer accelerator based on a ring reconfigurable architecture. The ring reconfigurable architecture forms a ring topology by combining multiple Transformer accelerator cores (TCs), storage units, and shared memories through an AXI bus, thereby achieving unified management of the Transformer accelerator cores and flexible execution path configuration.

[0010] A pipeline architecture is used to construct the Transformer Unit (TU), which includes a Multi-Header Attention Unit (MHAU) and a Multi-Layer Perceptron Unit (MLPU). The MHAU and MLPU reuse the matrix multiplication subunit and layer normalization unit to realize data normalization and multiplication and addition operations. The Softmax subunit of the MHAU and the GeLU subunit of the MLPU respectively realize standardization and activation calculations. The reuse and mutual collaboration between subunits greatly improves resource utilization and computing efficiency, making the accelerator deployment more flexible.

[0011] Furthermore, according to the amount of preprocessed data and the current status, the host can send control instructions to choose to shut down the Transformer Core (TC) of the non-execution path, realize flexible configuration of the execution path, flexibly manage TC computing resources, and reduce the overall power consumption of the accelerator while balancing performance.

[0012] Furthermore, the overall implementation of the visual Transformer accelerator is as follows:

[0013] The host loads the feature image of the visual Transformer neural network and performs a series of preprocessing operations on it, including slicing and pixel format conversion, and generates corresponding control instructions; then caches the preprocessed feature map data to the specified destination storage block;

[0014] The storage system includes a storage unit and shared memory. The storage unit consists of multiple storage blocks, which are interconnected through the AXI bus. As an important component of the accelerator, the shared memory is responsible for connecting the upper and lower level Transformer acceleration cores to form a ring topology. The intermediate iterative data generated by any level of Transformer acceleration core can be forwarded to the shared memory through the AXI bus routing and passed to the next level acceleration core.

[0015] Furthermore, the visual Transformer accelerator, by default, is configured with four Transformer acceleration cores as four acceleration nodes of a ring architecture, and the data to be processed flows through different acceleration nodes according to different control instructions.

[0016] Furthermore, the Transformer acceleration core specifically includes a Transformer computing unit, an AXI-MM bus routing, an AXI-ST bus routing, and a bus arbiter;

[0017] The Transformer computing unit is the core unit of the acceleration core, responsible for implementing multi-head attention and other multiplication and addition and nonlinear transformation operations;

[0018] The AXI-MM bus router is responsible for receiving pre-processed data, control instructions sent by the host, or intermediate data to be processed transmitted by the previous level Transformer computing unit;

[0019] The AXI-ST bus routing directly skips the current Transformer computing unit and passes the data to the next level Transformer computing unit according to the control instruction parameters;

[0020] The bus arbiter is responsible for accurately forwarding the data to be forwarded in the AXI-MM bus routing to the destination storage block specified by the control instruction parameters.

[0021] Furthermore, the Transformer computing unit specifically includes a multi-head attention unit, a multi-layer perceptron unit, and a buffer; the multi-head attention unit includes a layer normalization subunit, a key value generation subunit, a MatMul (MatrixMultiplication, MatMul) subunit, a Softmax subunit, a transposition subunit, and a buffer; the multi-layer perceptron unit includes a layer normalization subunit, a MatMul subunit, a Dropout subunit, a GeLU activation subunit, and a buffer, wherein the layer normalization subunit and the MatMul subunit serve as multiplexing units;

[0022] When the image is converted into a vector sequence X after preprocessing I The multi-head attention unit input to the Transformer computing unit has a layer normalization subunit that normalizes the input vector sequence along X I The channel dimension is normalized and the parameters are adjusted to make each activation value more evenly distributed, accelerate the convergence of the algorithm model and improve the overall performance of the algorithm model; the normalized sequence data X L Store in buffer;

[0023] The key value generation subunit takes the sequence data X from the buffer L, respectively, with the matrix vector Q that can be trained W , K W 、V W Perform matrix multiplication operations to obtain three feature matrices: Q (Query), K (Key), and V (Value), and cache them in the corresponding caches respectively;

[0024] Then the Q matrix and the K matrix are multiplied by the MatMul subunit, and the result is forwarded to the Softmax subunit for normalization to obtain the R matrix. At the same time, the V matrix is ​​input to the transpose subunit to perform the transposition operation to obtain the result V T matrix;

[0025] R matrix and V T The matrix is ​​then multiplied by the MatMul subunit to perform matrix multiplication and attention calculation, and finally the m attention calculation results are connected to form the matrix A; the matrix A and the weight matrix W O Then perform matrix multiplication to obtain the multi-head matrix Z. When there is a residual path, input the vector sequence X I Adding it to the multi-head matrix Z realizes the residual connection operation, effectively alleviating the gradient vanishing problem, and the final output vector sequence X O Contains all attention information, and the matrix is ​​input into the multi-layer perceptron unit;

[0026] The multi-layer perceptron unit contains a layer normalization subunit, a GeLU activation subunit, and multiple Dropout subunits. Multiple MLPUs are usually used to build the forward transmission path of the network to increase the nonlinear computing power of the model. O Input to the multi-layer perceptron unit, the layer normalization subunit along X O The channel dimension is normalized and the parameters are adjusted, and the results are stored in the buffer; then the regularization operation is performed through the Dropout sub-unit to prevent overfitting; then the nonlinear activation operation is performed through the GeLU activation sub-unit to make the input approximate a linear function when it is close to zero, which can better deal with the gradient disappearance problem; the output is passed through a Dropout sub-unit again to improve the overall expression ability of the model.

[0027] Furthermore, the implementation process of the visual Transformer task on the visual Transformer accelerator is as follows:

[0028] (1) The host uses image block embedding technology to convert the image into a vector sequence that can be processed by the Transformer model. First, the image is resized to a fixed size and the pixel values ​​are normalized to a value between 0 and 1 so that it can be converted into a floating-point tensor. Then it is cut into multiple small blocks of the same size and the small blocks are expanded into one-dimensional vectors. Finally, position encoding is added to the input sequence to provide the model with the position information of the vector in the sequence. After the vector sequence is prepared, the host generates control instruction parameters and selects the specified storage block and planned acceleration node to execute the Transformer task.

[0029] (2) The host forwards the preprocessed input sequence and control instruction parameters to the AXI-MM bus router in the first Transformer acceleration core in the ring architecture through PCIe DMA. According to the destination storage block provided by the control instruction parameters, the AXI-MM bus router forwards the data to be processed to the bus arbiter for storage. When other Transformer acceleration cores access the destination storage block at the same time, the data is stored in the storage block in turn according to the polling arbitration mechanism. When each Transformer acceleration core accesses different storage blocks, concurrent storage is supported. At this time, only the first Transformer acceleration core accesses the destination storage block, so there is no need to wait for the data to be cached to the destination storage block through the AXI bus.

[0030] (3) After the data is cached to the destination storage block, the execution path is configured according to the control instruction parameters;

[0031] (4) When the first Transformer acceleration core is activated, the vector sequence to be processed is read from the destination storage block and used as the input sequence X of the Transformer computation unit. I , then the multi-head attention unit in the Transformer computing unit pays attention to X I Normalization is performed, and three feature matrices Q, K, and V are generated through the key value generation unit. After a series of matrix multiplication operations, Softmax normalization operations, and residual connection operations, the output matrix Z is finally obtained. O ; Multilayer perceptron unit will Z O As the input matrix, continue the nonlinear calculation, including GeLU activation operation, Dropout regularization operation, etc., to obtain the final output result X O1 ;

[0032] (5) According to the configured execution path, the calculation result X O1The data is transmitted to the second Transformer acceleration core through the AXI bus. When the second Transformer acceleration core is online (activated), the Transformer computing unit integrated inside it receives the X O1 , repeat step (4) to finally obtain the intermediate sequence X O2 ; When forwarding to the next node, the third Transformer acceleration core, if the node is found to be offline (i.e. deactivated), X O2 Directly forwarded to the fourth Transformer accelerator core through AXI-ST routing, when the fourth Transformer accelerator core is online, it will continue to execute and O2 A series of normalization and nonlinear operations are performed as input, and the calculation results are finally cached in the result storage block;

[0033] (6) When the execution is completed, the host uses the PCIe DMA to access the result storage block based on the storage address information of the final result and uses the AXI-MM bus routing to retrieve the final inference result. At this point, the visual tranformer task is completed.

[0034] The present invention also claims a visual Transformer accelerator system, comprising a plurality of Transformer acceleration cores (TC), a storage unit and a shared memory.

[0035] The multiple Transformer Cores (TCs), storage units, and shared memories are connected via an AXI bus to form a ring topology, thereby achieving unified management of the Transformer Cores and flexible execution path configuration.

[0036] The Transformer acceleration core includes a Transformer computing unit, an AXI-MM bus routing, an AXI-ST bus routing, and a bus arbiter; the Transformer computing unit includes a multi-head attention unit (Multi-HeaderAttention Unit, MHAU) and a multi-layer perceptron unit (Multi-Layer Perceptron Unit, MLPU); the multi-head attention unit MHAU and the multi-layer perceptron unit MLPU reuse the matrix multiplication subunit and the layer normalization unit to realize data normalization and multiplication and addition operations; the Softmax subunit of the multi-head attention unit and the GeLU subunit of the multi-layer perceptron unit are used to realize standardization and activation calculations respectively;

[0037] The system can implement the above-mentioned visual Transformer accelerator implementation method.

[0038] The present invention also claims a visual Transformer accelerator implementation device, comprising: at least one memory and at least one processor;

[0039] The at least one memory is used to store a machine-readable program;

[0040] The at least one processor is used to call the machine-readable program to implement the above-mentioned visual Transformer accelerator implementation method.

[0041] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which, when executed by a processor, implement the above-mentioned visual Transformer accelerator implementation method.

[0042] Compared with the prior art, the visual Transformer accelerator implementation method and system of the present invention have the following beneficial effects:

[0043] The visual Transformer accelerator in the present invention adopts a unique ring-shaped reconfigurable architecture, and uses the AXI bus to form a ring topology of multiple Transformer acceleration cores, bus arbiters and storage units to achieve unified management of the Transformer acceleration cores and flexible execution path configuration; based on the ring topology structure, the data forwarding and computing efficiency between the acceleration cores are significantly improved, and compared with the GPU, the storage hierarchy and resource consumption are greatly reduced, the memory access delay is reduced, the execution efficiency of control instructions and data forwarding efficiency are improved, and the lightweight design is easier to deploy; by building such an accelerator, not only the execution speed of the visual Transformer model can be greatly improved, but also more room for innovation can be provided for developers, promoting the development of computer vision technology into a wider range of application fields.

[0044] In addition, according to the amount of preprocessed data and the current status, control instructions can be issued through the host, and the Transformer acceleration core of the non-execution path can be turned off to achieve flexible configuration of the execution path, flexibly manage the computing resources of the Transformer acceleration core, and reduce the overall power consumption of the accelerator while balancing performance.

[0045] The present invention also uses a pipeline architecture to construct a Transformer computing unit, which includes a multi-head attention unit and a multi-layer perceptron unit, and decomposes the Transformer operation into different sub-units. Through the reuse and mutual cooperation between sub-units, the resource utilization and computing efficiency are greatly improved, making the accelerator deployment more flexible. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1is a diagram showing the overall architecture of a visual Transformer accelerator provided by an embodiment of the present invention;

[0047] Figure 2 is a schematic diagram of the design architecture of a Transformer computing unit provided by an embodiment of the present invention;

[0048] Figure 3 This is an example diagram of a specific application method of a visual Transformer accelerator provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The present invention will be further described below in conjunction with specific embodiments.

[0050] The embodiment of the present invention provides a method for implementing a visual Transformer accelerator, which implements the visual Transformer accelerator based on a ring-shaped reconfigurable architecture. The ring-shaped reconfigurable architecture uses an AXI bus to form a ring topology with multiple Transformer accelerator cores (TC), storage units, and shared memories, so as to achieve unified management of the Transformer accelerator cores and flexible execution path configuration. The ring-shaped topology structure significantly improves the data forwarding and computing efficiency between the accelerator cores, greatly reduces the storage hierarchy and resource consumption compared to a GPU, reduces memory access latency, improves the execution efficiency of control instructions and data forwarding efficiency, and is lightweight and easier to deploy.

[0051] According to the amount of preprocessed data and the current status, the host can send control instructions to choose to shut down the Transformer acceleration core of the non-execution path, realize flexible configuration of the execution path, flexibly manage TC computing resources, and reduce the overall power consumption of the accelerator while balancing performance.

[0052] A pipeline architecture is used to construct the Transformer Unit (TU), which includes a Multi-Header Attention Unit (MHAU) and a Multi-Layer Perceptron Unit (MLPU). The MHAU and MLPU reuse the matrix multiplication subunit and layer normalization unit to realize data normalization and multiplication and addition operations. The Softmax subunit of the MHAU and the GeLU subunit of the MLPU respectively realize standardization and activation calculations. The reuse and mutual collaboration between subunits greatly improves resource utilization and computing efficiency, making the accelerator deployment more flexible.

[0053] The overall architecture of the visual Transformer accelerator is as follows: Figure 1 As shown,

[0054] The host is responsible for loading the feature image of the visual Transformer (ViT) neural network and performing a series of preprocessing operations such as slicing and pixel format conversion, while generating corresponding control instructions; then the preprocessed feature map data is cached to the specified destination storage block LR;

[0055] The storage system includes a storage unit and shared memory. The storage unit is composed of multiple storage blocks LR, which are interconnected through the AXI bus. As an important component of the accelerator, the shared memory is responsible for connecting the upper and lower level Transformer acceleration cores to form a ring topology structure. The intermediate iterative data generated by any level of Transformer acceleration core can be forwarded to the shared memory through the AXI bus routing and passed to the next level acceleration core.

[0056] The detailed Transformer acceleration core architecture is as follows Figure 1 As shown, it consists of a Transformer computing unit, an AXI-MM bus routing, an AXI-ST bus routing, and a bus arbiter.

[0057] The Transformer computing unit is the core unit of the acceleration core, responsible for implementing multi-head attention and other multiplication and addition and nonlinear transformation operations;

[0058] The AXI-MM bus router is responsible for receiving pre-processed data, control instructions sent by the host, or intermediate data to be processed transmitted by the previous level Transformer computing unit;

[0059] The AXI-ST bus routing directly skips the current Transformer computing unit and passes the data to the next level Transformer computing unit according to the control instruction parameters;

[0060] The bus arbiter is responsible for accurately forwarding the data to be forwarded in the AXI-MM bus routing to the destination storage block specified by the control instruction parameters.

[0061] like Figure 2 The following is a diagram showing the design architecture of the Transformer computing unit.

[0062] As the core module of the acceleration core, the Transformer computing unit adopts a pipeline architecture to realize all computing functions, including multi-head attention units, multi-layer perceptron units and buffers.

[0063] The multi-head attention unit consists of a layer normalization subunit, a key-value generation subunit, a MatMul (MatrixMultiplication, MatMul) subunit, a Softmax subunit, a transposition subunit, and a buffer; the multi-layer perceptron unit consists of a layer normalization subunit, a MatMul subunit, a Dropout subunit, a GeLU activation subunit, and a buffer, among which the layer normalization subunit and the MatMul subunit serve as reuse units.

[0064] When the image is converted into a vector sequence X after preprocessing I The multi-head attention unit input to the Transformer computing unit has a layer normalization subunit that normalizes the input vector sequence along X I The channel dimension is normalized and the parameters are adjusted to make each activation value more evenly distributed, accelerate the convergence of the algorithm model and improve the overall performance of the algorithm model; the normalized sequence data X L The key value generation subunit takes the sequence data X from the buffer. L , respectively, with the trainable matrix vector Q W , K W 、V W Perform matrix multiplication to obtain three feature matrices Q (Query), K (Key), and V (Value), which are cached in the corresponding caches. Then the Q matrix and the K matrix complete the matrix multiplication operation through the MatMul subunit, and then forward the result to the Softmax subunit for normalization operation to obtain the R matrix; at the same time, the V matrix is ​​input into the transpose subunit to perform the transposition operation to obtain the result V T Matrix. R matrix and V T The matrix is ​​then multiplied by the MatMul subunit to perform matrix multiplication and attention calculation, and finally the m attention calculation results are connected to form the matrix A; the matrix A and the weight matrix W O Then perform matrix multiplication to obtain the multi-head matrix Z. When there is a residual path, input the vector sequence X I Adding it to the multi-head matrix Z realizes the residual connection operation, effectively alleviating the gradient vanishing problem, and the final output vector sequence X O Contains all attention information, and this matrix is ​​input into the multi-layer perceptron unit.

[0065] The multi-layer perceptron unit contains a layer normalization subunit, a GeLU activation subunit, and multiple Dropout subunits. Multiple MLPUs are usually used to build the network's forward pass path to increase the model's nonlinear computing power. O Input to the multi-layer perceptron unit, the layer normalization subunit along X OThe channel dimension is normalized and the parameters are adjusted, and the results are stored in the buffer; then the regularization operation is performed through the Dropout sub-unit to prevent overfitting; then the nonlinear activation operation is performed through the GeLU activation sub-unit to make the input approximate a linear function when it is close to zero, which can better deal with the gradient disappearance problem; the output is passed through a Dropout sub-unit again to improve the overall expression ability of the model.

[0066] Combined with Figure 3 As shown, the specific application method of ViT accelerator is as follows:

[0067] The ViT accelerator designed by this method is based on a ring reconfigurable architecture. Several Transformer acceleration cores are distributed on the nodes of the ring architecture and share storage units and shared memory through the AXI bus. This architecture can flexibly configure the number of acceleration cores according to specific application scenarios. At the same time, it can flexibly select the execution path of data according to control instructions and shut down idle Transformer acceleration cores (TC) to reduce power consumption. By default, the accelerator configures 4 Transformer acceleration cores as 4 acceleration nodes of the ring architecture, such as Figure 3 The figure shows that by default, the data to be processed flows through different acceleration nodes according to different control instructions, where Node1 represents the first TC in the ring architecture, and so on.

[0068] The implementation process of the visual Transformer task on the visual Transformer accelerator is as follows:

[0069] 1. The host uses image block embedding technology to convert the image into a vector sequence that can be processed by the Transformer model. First, the image is resized to a fixed size and the pixel values ​​are normalized to a value between 0 and 1 so that it can be converted into a floating-point tensor; then it is cut into multiple small blocks of the same size and the small blocks are expanded into a one-dimensional vector. Finally, position encoding is added to the input sequence to provide the model with the position information of the vector in the sequence. After the vector sequence is prepared, the host generates control instruction parameters and selects the specified storage block and planned acceleration node to execute the Transformer task.

[0070] 2. Figure 3Taking the execution path of (a) as an example, the host forwards the pre-processed input sequence and control instruction parameters to the AXI-MM bus routing in Node1 in the ring architecture through PCIe DMA. According to the destination LR1 provided by the control instruction parameters, the AXI-MM bus routing forwards the data to be processed to the bus arbiter for storage. When other nodes access LR1 at the same time, the data is stored in LR in turn according to the polling arbitration mechanism. When each node accesses different LRs, concurrent storage is supported. At this time, only Node1 accesses LR1, so there is no need to wait for the data to be cached in LR1 through the AXI bus.

[0071] 3. After the data is cached in LR1, the execution path is configured according to the control instruction parameters. This control instruction deactivates Node3 and activates Node1, Node2, and Node4, forming a multi-node and multi-level Tranformer task data processing path.

[0072] 4. When Node1 is activated, the vector sequence to be processed is read from LR1 and used as the input sequence X of TU I , then MHAU in TU against X I Normalization is performed, and the three feature matrices Q, K, and V are generated through the key value generation unit (see the design part of the Transformer calculation unit for specific steps). After that, a series of matrix multiplication operations, Softmax normalization operations, and residual connection operations are performed to finally obtain the output matrix Z O MLPU will Z O As the input matrix, we continue to perform nonlinear calculations such as GeLU activation operation and Dropout regularization operation to obtain the final output result X O1 .

[0073] 5. According to the configured execution path, the calculation result X O1 It is transmitted to Node2 through the AXI bus. When Node2 is online (activated), its internal integrated TU computing unit receives X O1 Repeat step 4 to finally get the intermediate sequence X O2 When the next node Node3 is forwarded in sequence, it is found that Node3 is offline (i.e. deactivated), so X O2 Directly forwarded to Node4 via AXI-ST routing, when Node4 is online, it will continue to execute and O2 A series of normalization and nonlinear operations are performed as input and the calculation results are finally cached in LR4.

[0074] 6. When the execution is completed, the host uses the PCIe DMA to access LR4 based on the storage address information of the final result and uses the AXI-MM bus routing to retrieve the final inference result. At this point, the visual tranformer task is completed.

[0075] An embodiment of the present invention further provides a visual Transformer accelerator system, which can implement the visual Transformer accelerator implementation method described in the above embodiment.

[0076] The system includes multiple Transformer Cores (TCs), storage units, and shared memories, which are connected via an AXI bus to form a ring topology, thereby achieving unified management of the Transformer Cores and flexible execution path configuration.

[0077] The Transformer acceleration core includes a Transformer computing unit, an AXI-MM bus routing, an AXI-ST bus routing and a bus arbiter; the Transformer computing unit includes a multi-head attention unit (Multi-HeaderAttention Unit, MHAU) and a multi-layer perceptron unit (Multi-Layer Perceptron Unit, MLPU); the multi-head attention unit MHAU and the multi-layer perceptron unit MLPU reuse the matrix multiplication subunit and the layer normalization unit to realize data normalization and multiplication and addition operations; the standardization and activation calculations are respectively realized by the Softmax subunit of the multi-head attention unit and the GeLU subunit of the multi-layer perceptron unit.

[0078] The host is responsible for loading the feature image of the visual Transformer (ViT) neural network and performing a series of preprocessing operations such as slicing and pixel format conversion, while generating corresponding control instructions; then the preprocessed feature map data is cached to the specified destination storage block LR;

[0079] The storage system includes a storage unit and shared memory. The storage unit is composed of multiple storage blocks LR, which are interconnected through the AXI bus. As an important component of the accelerator, the shared memory is responsible for connecting the upper and lower level Transformer acceleration cores to form a ring topology structure. The intermediate iterative data generated by any level of Transformer acceleration core can be forwarded to the shared memory through the AXI bus routing and passed to the next level acceleration core.

[0080] According to the amount of preprocessed data and the current status, the host can send control instructions to choose to shut down the Transformer acceleration core of the non-execution path, realize flexible configuration of the execution path, flexibly manage TC computing resources, and reduce the overall power consumption of the accelerator while balancing performance.

[0081] The Transformer acceleration core is specifically composed of a Transformer computing unit, an AXI-MM bus routing, an AXI-ST bus routing, and a bus arbiter.

[0082] The Transformer computing unit is the core unit of the acceleration core, responsible for implementing multi-head attention and other multiplication and addition and nonlinear transformation operations;

[0083] The AXI-MM bus router is responsible for receiving pre-processed data, control instructions sent by the host, or intermediate data to be processed transmitted by the previous level Transformer computing unit;

[0084] The AXI-ST bus routing directly skips the current Transformer computing unit and passes the data to the next level Transformer computing unit according to the control instruction parameters;

[0085] The bus arbiter is responsible for accurately forwarding the data to be forwarded in the AXI-MM bus routing to the destination storage block specified by the control instruction parameters.

[0086] As the core module of the acceleration core, the Transformer computing unit adopts a pipeline architecture to realize all computing functions, including multi-head attention units, multi-layer perceptron units and buffers.

[0087] The multi-head attention unit consists of a layer normalization subunit, a key-value generation subunit, a MatMul (MatrixMultiplication, MatMul) subunit, a Softmax subunit, a transposition subunit, and a buffer; the multi-layer perceptron unit consists of a layer normalization subunit, a MatMul subunit, a Dropout subunit, a GeLU activation subunit, and a buffer, among which the layer normalization subunit and the MatMul subunit serve as reuse units.

[0088] When the image is converted into a vector sequence X after preprocessing I The multi-head attention unit input to the Transformer computing unit has a layer normalization subunit that normalizes the input vector sequence along X I The channel dimension is normalized and the parameters are adjusted to make each activation value more evenly distributed, accelerate the convergence of the algorithm model and improve the overall performance of the algorithm model; the normalized sequence data X LThe key value generation subunit takes the sequence data X from the buffer. L , respectively, with the matrix vector Q that can be trained W , K W 、V W Perform matrix multiplication to obtain three feature matrices Q (Query), K (Key), and V (Value), which are cached in the corresponding caches. Then the Q matrix and the K matrix complete the matrix multiplication operation through the MatMul subunit, and then forward the result to the Softmax subunit for normalization operation to obtain the R matrix; at the same time, the V matrix is ​​input into the transpose subunit to perform the transposition operation to obtain the result V T Matrix. R matrix and V T The matrix is ​​then multiplied by the MatMul subunit to perform matrix multiplication and attention calculation, and finally the m attention calculation results are connected to form the matrix A; the matrix A and the weight matrix W O Then perform matrix multiplication to obtain the multi-head matrix Z. When there is a residual path, input the vector sequence X I Adding it to the multi-head matrix Z realizes the residual connection operation, effectively alleviating the gradient vanishing problem, and the final output vector sequence X O Contains all attention information, and this matrix is ​​input into the multi-layer perceptron unit.

[0089] The multi-layer perceptron unit contains a layer normalization subunit, a GeLU activation subunit, and multiple Dropout subunits. Multiple MLPUs are usually used to build the network's forward pass path to increase the model's nonlinear computing power. O Input to the multi-layer perceptron unit, the layer normalization subunit along X O The channel dimension is normalized and the parameters are adjusted, and the results are stored in the buffer; then the regularization operation is performed through the Dropout sub-unit to prevent overfitting; then the nonlinear activation operation is performed through the GeLU activation sub-unit to make the input approximate a linear function when it is close to zero, which can better deal with the gradient disappearance problem; the output is passed through a Dropout sub-unit again to improve the overall expression ability of the model.

[0090] The specific application method of the visual Transformer accelerator is as follows:

[0091] The visual Transformer accelerator is based on a ring reconfigurable architecture. Several Transformer acceleration cores are distributed on the nodes of the ring architecture and share storage units and shared memory through the AXI bus. This architecture can flexibly configure the number of acceleration cores according to specific application scenarios. At the same time, it can flexibly select the execution path of data according to control instructions and shut down idle Transformer acceleration cores (TC) to reduce power consumption. By default, the accelerator configures 4 Transformer acceleration cores as 4 acceleration nodes of the ring architecture, such as Figure 3 The figure shows that by default, the data to be processed flows through different acceleration nodes according to different control instructions, where Node1 represents the first TC in the ring architecture, and so on.

[0092] The implementation process of the visual Transformer task on the visual Transformer accelerator is as follows:

[0093] 1. The host uses image block embedding technology to convert the image into a vector sequence that can be processed by the Transformer model. First, the image is resized to a fixed size and the pixel values ​​are normalized to a value between 0 and 1 so that it can be converted into a floating-point tensor; then it is cut into multiple small blocks of the same size and the small blocks are expanded into a one-dimensional vector. Finally, position encoding is added to the input sequence to provide the model with the position information of the vector in the sequence. After the vector sequence is prepared, the host generates control instruction parameters and selects the specified storage block and planned acceleration node to execute the Transformer task.

[0094] 2. Figure 3 Taking the execution path of (a) as an example, the host forwards the pre-processed input sequence and control instruction parameters to the AXI-MM bus routing in Node1 in the ring architecture through PCIe DMA. According to the destination LR1 provided by the control instruction parameters, the AXI-MM bus routing forwards the data to be processed to the bus arbiter for storage. When other nodes access LR1 at the same time, the data is stored in LR in turn according to the polling arbitration mechanism. When each node accesses different LRs, concurrent storage is supported. At this time, only Node1 accesses LR1, so there is no need to wait for the data to be cached in LR1 through the AXI bus.

[0095] 3. After the data is cached in LR1, the execution path is configured according to the control instruction parameters. This control instruction deactivates Node3 and activates Node1, Node2, and Node4, forming a multi-node and multi-level Tranformer task data processing path.

[0096] 4. When Node1 is activated, the vector sequence to be processed is read from LR1 and used as the input sequence X of TUI , then MHAU in TU against X I Normalization is performed, and the three feature matrices Q, K, and V are generated through the key value generation unit (see the design part of the Transformer calculation unit for specific steps). After that, a series of matrix multiplication operations, Softmax normalization operations, and residual connection operations are performed to finally obtain the output matrix Z O MLPU will Z O As the input matrix, we continue to perform nonlinear calculations such as GeLU activation operation and Dropout regularization operation to obtain the final output result X O1 .

[0097] 5. According to the configured execution path, the calculation result X O1 It is transmitted to Node2 through the AXI bus. When Node2 is online (activated), its internal integrated TU computing unit receives X O1 Repeat step 4 to finally get the intermediate sequence X O2 When the next node Node3 is forwarded in sequence, it is found that Node3 is offline (i.e. deactivated), so X O2 Directly forwarded to Node4 via AXI-ST routing, when Node4 is online, it will continue to execute and O2 A series of normalization and nonlinear operations are performed as input and the calculation results are finally cached in LR4.

[0098] 6. When the execution is completed, the host uses the PCIe DMA to access LR4 based on the storage address information of the final result and uses the AXI-MM bus routing to retrieve the final inference result. At this point, the visual tranformer task is completed.

[0099] An embodiment of the present invention further provides a visual Transformer accelerator implementation device, comprising: at least one memory and at least one processor;

[0100] The at least one memory is used to store a machine-readable program;

[0101] The at least one processor is used to call the machine-readable program to implement the visual Transformer accelerator implementation method described in the above embodiment.

[0102] The embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the processor executes the visual Transformer accelerator implementation method described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes implementing the functions of any of the above embodiments are stored, and a computer (or CPU or MPU) of the system or device is enabled to read and execute the program code stored in the storage medium.

[0103] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0104] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0105] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0106] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0107] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.

Claims

1. A method for implementing a visual Transformer accelerator, characterized in that: Implementing a visual Transformer accelerator based on a ring reconfigurable architecture. The ring reconfigurable architecture uses an AXI bus to combine multiple Transformer accelerator cores, storage units, and shared memory into a ring topology to achieve unified management of the Transformer accelerator cores and flexible execution path configuration. The Transformer computing unit is constructed using a pipeline architecture, which includes a multi-head attention unit and a multi-layer perceptron unit. The multi-head attention unit and the multi-layer perceptron unit reuse the matrix multiplication subunit and the layer normalization unit to realize data normalization and multiplication and addition operations. The Softmax subunit of the multi-head attention unit and the GeLU subunit of the multi-layer perceptron unit respectively realize standardization and activation calculations. The Transformer computing unit specifically includes a multi-head attention unit, a multi-layer perceptron unit, and a buffer; the multi-head attention unit includes a layer normalization subunit, a key value generation subunit, a MatMul subunit, a Softmax subunit, a transposition subunit, and a buffer; The multi-layer perceptron unit includes a layer normalization subunit, a MatMul subunit, a Dropout subunit, a GeLU activation subunit, and a buffer, wherein the layer normalization subunit and the MatMul subunit serve as multiplexing units; When the image is converted into a vector sequence X after preprocessing I The multi-head attention unit input to the Transformer computing unit has a layer normalization subunit that normalizes the input vector sequence along X I The channel dimension is normalized and the parameters are adjusted to make each activation value more evenly distributed; the normalized sequence data X L Store in buffer; The key value generation subunit takes the sequence data X from the buffer L , respectively, with the trainable matrix vector Q W , K W 、V W Perform matrix multiplication operations to obtain three feature matrices Q, K, and V, and cache them in corresponding caches respectively; Then the Q matrix and the K matrix complete the matrix multiplication operation through the MatMul sub-unit, and then forward the result to the Softmax sub-unit for normalization operation to obtain the R matrix; at the same time, the V matrix is ​​input into the transpose sub-unit to perform the transposition operation to obtain the result V T matrix; R matrix and V T The matrix is ​​then multiplied by the MatMul subunit to perform matrix multiplication and attention calculation, and finally the m attention calculation results are connected to form the matrix A; Matrix A and weight matrix W O Then perform matrix multiplication to obtain the multi-head matrix Z. When there is a residual path, input the vector sequence X I Adding it to the multi-head matrix Z realizes the residual connection operation, and the final output vector sequence X O Contains all attention information, and the matrix is ​​input into the multi-layer perceptron unit; The multi-layer perceptron unit contains a layer normalization subunit, a GeLU activation subunit, and multiple Dropout subunits. O Input to the multi-layer perceptron unit, the layer normalization subunit along X O The channel dimension is normalized and the parameters are adjusted, and the results are stored in the buffer; It then passes through a Dropout subunit for regularization to prevent overfitting; it then passes through a GeLU activation subunit for nonlinear activation to approximate a linear function when its input is close to zero; and it passes through a Dropout subunit again at output to improve the overall expressiveness of the model.

2. A method for implementing a visual Transformer accelerator according to claim 1, characterized in that: According to the amount of preprocessed data and the current status, the host can send control instructions to choose to shut down the Transformer acceleration core of the non-execution path to achieve flexible configuration of the execution path.

3. A method for implementing a visual Transformer accelerator according to claim 1 or 2, characterized in that: The overall implementation of the visual Transformer accelerator is as follows: Load the feature image of the visual Transformer neural network through the host and perform preprocessing operations including slicing and pixel format conversion on it, and generate corresponding control instructions; Then cache the preprocessed feature map data into the specified destination storage block; The storage system includes a storage unit and a shared memory. The storage unit consists of multiple storage blocks, which are interconnected through an AXI bus. The shared memory is responsible for connecting the upper and lower level Transformer acceleration cores to form a ring topology. The intermediate iterative data generated by any level of Transformer acceleration core can be forwarded to the shared memory through the AXI bus routing and passed to the next level acceleration core.

4. A method for implementing a visual Transformer accelerator according to claim 2, characterized in that: The visual Transformer accelerator is configured with four Transformer acceleration cores as four acceleration nodes of a ring architecture by default.

5. The method for implementing a visual Transformer accelerator according to claim 1, characterized in that: The Transformer acceleration core specifically includes a Transformer computing unit, an AXI-MM bus routing, an AXI-ST bus routing, and a bus arbiter; The Transformer computing unit is responsible for implementing multiplication and addition and nonlinear transformation operations; The AXI-MM bus router is responsible for receiving pre-processed data, control instructions sent by the host, or intermediate data to be processed transmitted by the previous level Transformer computing unit; The AXI-ST bus routing directly skips the current Transformer computing unit and passes the data to the next level Transformer computing unit according to the control instruction parameters; The bus arbiter is responsible for accurately forwarding the data to be forwarded in the AXI-MM bus routing to the destination storage block specified by the control instruction parameters.

6. A method for implementing a visual Transformer accelerator according to claim 1, characterized in that: The implementation process of the visual Transformer task on the visual Transformer accelerator is as follows: (1) The host uses image block embedding technology to convert the image into a vector sequence that can be processed by the Transformer model. First, the image is resized to a fixed size and the pixel values ​​are normalized to a value between 0 and 1 so that it can be converted into a floating-point tensor. Then it is cut into multiple small blocks of the same size and the small blocks are expanded into one-dimensional vectors. Finally, position encoding is added to the input sequence to provide the model with the position information of the vector in the sequence. After the vector sequence is prepared, the host generates control instruction parameters and selects the specified storage block and planned acceleration node to execute the Transformer task. (2) The host forwards the preprocessed input sequence and control instruction parameters to the AXI-MM bus router in the first Transformer acceleration core in the ring architecture through PCIe DMA. According to the destination storage block provided by the control instruction parameters, the AXI-MM bus router forwards the data to be processed to the bus arbiter for storage. When other Transformer acceleration cores access the destination storage block at the same time, the data is stored in the storage block in turn according to the polling arbitration mechanism. When each Transformer acceleration core accesses different storage blocks, concurrent storage is supported. At this time, only the first Transformer acceleration core accesses the destination storage block, and there is no need to wait for the data to be cached to the destination storage block through the AXI bus. (3) After the data is cached to the destination storage block, the execution path is configured according to the control instruction parameters; (4) When the first Transformer acceleration core is activated, the vector sequence to be processed is read from the destination storage block and used as the input sequence X of the Transformer computation unit. I , then the multi-head attention unit in the Transformer computing unit pays attention to X I Normalization is performed, and three feature matrices Q, K, and V are generated through the key value generation unit. After matrix multiplication, Softmax normalization, and residual connection operations, the output matrix Z is finally obtained. O ; Multilayer perceptron unit will Z O As the input matrix, continue the nonlinear calculation, including GeLU activation operation and Dropout regularization operation, and get the final output result X O1 ; (5) According to the configured execution path, the calculation result X O1 The data is transmitted to the second Transformer acceleration core through the AXI bus. When the second Transformer acceleration core is online, the Transformer computing unit integrated inside it receives the X O1 , repeat step (4) to finally obtain the intermediate sequence X O2 ; When forwarding to the next node, the third Transformer acceleration core, if the node is found to be offline, X O2 Directly forwarded to the fourth Transformer accelerator core through AXI-ST routing, when the fourth Transformer accelerator core is online, it will continue to execute and O2 Normalization and nonlinear operations are performed as input and the calculation results are finally cached in the result storage block; (6) When the execution is completed, the host uses the PCIe DMA to access the result storage block based on the storage address information of the final result and uses the AXI-MM bus routing to retrieve the final inference result. At this point, the visual tranformer task is completed.

7. A visual Transformer accelerator system, characterized in that: Including multiple Transformer acceleration cores, storage units and shared memory, The multiple Transformer acceleration cores, storage units and shared memories are connected via an AXI bus to form a ring topology, thereby achieving unified management of the Transformer acceleration cores and flexible execution path configuration; The Transformer acceleration core includes a Transformer computing unit, an AXI-MM bus routing, an AXI-ST bus routing, and a bus arbiter; the Transformer computing unit includes a multi-head attention unit and a multi-layer perceptron unit; the multi-head attention unit and the multi-layer perceptron unit reuse the matrix multiplication subunit and the layer normalization unit to realize data normalization and multiplication and addition operations; the Softmax subunit of the multi-head attention unit and the GeLU subunit of the multi-layer perceptron unit are used to realize standardization and activation calculations respectively; The system implements the method described in any one of claims 1 to 6.

8. A visual Transformer accelerator implementation device, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 6.

9. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Transform neural network system and operation method thereof

    CN116502675A

  • Transform model accelerator and construction and data processing method and device thereof

    CN116861966A