Layer Normalization Processing Hardware Accelerator and Method Applied to Transformer Neural Network
By designing a layer normalization processing hardware accelerator for Transformer neural networks, the problem of large delay in the execution layer normalization processing is solved, and the effect of improving computing speed and efficiency is achieved.
Patent Information
- Application Number
- CN202010898001.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-08-31
AI Technical Summary
Transformer neural network has a large delay when performing layer normalization processing, which affects its computing speed and efficiency.
A layer normalization processing hardware accelerator applied to Transformer neural network is designed, including an intermediate matrix storage unit, an average calculation unit, a square calculation unit and a square root reciprocal calculation unit. Through the combination of these units, the calculation process of layer normalization processing is optimized.
It effectively reduces the delay of layer normalization processing and improves the computing speed and efficiency of Transformer neural network.
Smart Images

Figure CN114118343B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of neural networks, and particularly to a layer normalization processing hardware accelerator and method applied to Transformer neural networks. Background Art
[0002] The Transformer network is a neural network model for solving natural language processing problems. Its model architecture is as Figure 1 shown, mainly including an encoder stack and a decoder stack. The encoder stack and the decoder stack each contain N encoder layers and multiple decoder layers. During the calculation process of the Transformer neural network, the input sequence first undergoes word vector embedding layer processing and position encoding superposition processing to obtain an input matrix. This input matrix is input into the encoder stack and successively undergoes operations of multiple encoder layers to obtain the output matrix of the encoder stack. After the encoding stage is the decoding stage. In each step of the decoding stage, an element of the target sentence is output to achieve natural language processing.
[0003] Each encoder layer and decoder layer is composed of a multi-head attention layer and a feed-forward layer. The multi-head attention layer includes three input matrices of the same size, namely the first input matrix, the second input matrix, and the third input matrix. The feed-forward layer only includes one input matrix. In the multi-head attention layer, after the three input matrices undergo a series of processes (including linear processing and processing by the Softmax layer), a first intermediate matrix is obtained. Then, layer normalization processing is performed on this intermediate matrix to obtain the final output matrix of the multi-head attention layer. Similarly, in the feed-forward layer, the input matrix undergoes a series of processes to obtain a second intermediate matrix, and then layer normalization processing is performed on this intermediate matrix to obtain the output matrix of the feed-forward layer.
[0004] Currently, the above calculation process is run on general computing platforms such as CPUs or GPUs. During the process of performing layer normalization processing, in order to obtain the variance value of each row element of the intermediate matrix, it is necessary to first calculate the average value of each row element in the intermediate matrix, then respectively obtain the difference between each element and the average value, and perform an accumulation operation after squaring the difference. Such processing steps are relatively cumbersome and there is a large delay. In order to improve the operation speed and efficiency of the Transformer neural network, it is urgent to design a hardware accelerator dedicated to layer normalization processing. Summary of the Invention
[0005] In order to reduce the delay in the process of layer normalization processing and improve the operation speed and efficiency of the Transformer neural network, this application discloses a layer normalization processing hardware accelerator and method applied to Transformer neural networks through the following embodiments.
[0006] The first aspect of the present application discloses a layer normalization processing hardware accelerator applied to a Transformer neural network. The layer normalization processing hardware accelerator includes:
[0007] An intermediate matrix storage unit, a first mean calculation unit, a second mean calculation unit, a first square calculation unit, a second square calculation unit, a reciprocal square root calculation unit, and an output matrix calculation unit;
[0008] The output end of the intermediate matrix storage unit is connected to the output matrix calculation unit;
[0009] The output end of the first mean calculation unit is respectively connected to the first square calculation unit and the output matrix calculation unit;
[0010] The output end of the first square calculation unit is connected to the reciprocal square root calculation unit;
[0011] The output end of the second square calculation unit is connected to the second mean calculation unit;
[0012] The output end of the second mean calculation unit is connected to the reciprocal square root calculation unit;
[0013] The output end of the reciprocal square root calculation unit is connected to the output matrix calculation unit.
[0014] Optionally, the intermediate matrix storage unit is used to obtain and store the intermediate matrix, and the intermediate matrix is the first intermediate matrix in the processing process of the multi-head attention layer or the second intermediate matrix in the processing process of the feed-forward layer;
[0015] The first mean calculation unit is used to calculate the mean value of each row of elements in the intermediate matrix and input the calculation result into the first square calculation unit;
[0016] The first square calculation unit is used to perform a square operation on the value input by the first mean calculation unit to obtain the square of the mean value of each row of elements in the intermediate matrix;
[0017] The second square calculation unit is used to perform a square operation on each element in the intermediate matrix to obtain a square matrix;
[0018] The second mean calculation unit is used to calculate the mean value of each row of elements in the square matrix;
[0019] The reciprocal square root calculation unit is used to obtain the reciprocal square root of the variance of each row of elements in the intermediate matrix according to the square of the mean value of each row of elements in the intermediate matrix and the mean value of each row of elements in the square matrix;
[0020] The output matrix calculation unit is used to perform layer normalization on each element of the intermediate matrix, the mean of each row of elements in the intermediate matrix, and the reciprocal of the square root of the variance of each row of elements in the intermediate matrix, to obtain the final output matrix of the multi-head attention layer or the feed-forward layer.
[0021] Optionally, in the process of obtaining the reciprocal of the square root of the variance of each row of elements in the intermediate matrix according to the square of the mean of each row of elements in the intermediate matrix and the mean of each row of elements in the square matrix, the square root reciprocal calculation unit is used to obtain the variance of each row of elements in the intermediate matrix according to the following formula:
[0022] ;
[0023] ;
[0024] where represents the variance of the elements in the i-th row of the intermediate matrix G, represents the mean of the elements in the i-th row of the intermediate matrix G, represents the mean of the squares of the elements in the i-th row of the square matrix, represents the element in the i-th row and k-th column of the intermediate matrix, represents the total number of columns of the intermediate matrix.
[0025] Optionally, the output matrix calculation unit is used to perform layer normalization on each element of the intermediate matrix, the mean of each row of elements in the intermediate matrix, and the reciprocal of the square root of the variance of each row of elements in the intermediate matrix according to the following formula, to obtain the final output matrix of the multi-head attention layer or the feed-forward layer:
[0026] ;
[0027] where represents the element in the i-th row and j-th column of the output matrix, represents the variance of the elements in the i-th row of the intermediate matrix G, represents the element in the i-th row and j-th column of the intermediate matrix G, represents the mean of the elements in the i-th row of the intermediate matrix, ε is the first parameter, represents the second parameter, represents the third parameter.
[0028] Optionally, the first mean calculation unit includes a plurality of first mean calculation sub-units, the second mean calculation unit includes a plurality of second mean calculation sub-units, the first square calculation unit includes a plurality of first square calculation sub-units, the second square calculation unit includes a plurality of second square calculation sub-units, the reciprocal square root calculation unit includes a plurality of reciprocal square root calculation sub-units, and the output matrix calculation unit includes a plurality of output matrix calculation sub-units;
[0029] The number of the first mean calculation sub-units, the second mean calculation sub-units, the first square calculation sub-units, the second square calculation sub-units, the reciprocal square root calculation sub-units, and the output matrix calculation sub-units is the same as the number of rows of any input matrix in the multi-head attention layer.
[0030] A second aspect of the present application discloses a layer normalization processing method applied to a Transformer neural network. The layer normalization processing method is applied to the layer normalization processing hardware accelerator for a Transformer neural network according to the first aspect of the present application. The layer normalization processing method includes:
[0031] All elements of the intermediate matrix are sequentially input into the intermediate matrix storage unit in column order. Among them, if the current operation belongs to the multi-head attention layer, the intermediate matrix is the first intermediate matrix; if the current operation is the feed-forward layer, the intermediate matrix is the second intermediate matrix;
[0032] Each row element of the intermediate matrix is respectively input into a plurality of first mean calculation sub-units to calculate the mean value of each row element in the intermediate matrix; and each row element of the intermediate matrix is respectively input into a plurality of second square calculation sub-units to obtain a square matrix;
[0033] The mean value of each row element in the intermediate matrix is respectively input into a plurality of first square calculation sub-units to obtain the square of the mean value of each row element in the intermediate matrix;
[0034] Each row element of the square matrix is respectively input into a plurality of second mean calculation sub-units to calculate the mean value of each row element in the square matrix;
[0035] The square of the mean value of each row element in the intermediate matrix and the mean value of each row element in the square matrix are respectively input into a plurality of reciprocal square root calculation sub-units to obtain the reciprocal of the square root of the variance of each row element in the intermediate matrix;
[0036] Each element of the intermediate matrix, the mean value of each row element in the intermediate matrix, and the reciprocal of the square root of the variance of each row element in the intermediate matrix are respectively input into a plurality of output matrix calculation sub-units to obtain the final output matrix of the multi-head attention layer or the feed-forward layer.
[0037] The third aspect of the present application discloses a computer device, including:
[0038] A memory for storing a computer program;
[0039] A processor for implementing the steps of the layer normalization processing method applied to the Transformer neural network as described in the second aspect of the present application when executing the computer program.
[0040] The fourth aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored, and the computer program realizes the steps of the layer normalization processing method applied to the Transformer neural network as described in the second aspect of the present application when being processed and executed.
[0041] The present application discloses a layer normalization processing hardware accelerator and method applied to a Transformer neural network. The hardware accelerator includes an intermediate matrix storage unit, a first mean calculation unit, a second mean calculation unit, a first square calculation unit, a second square calculation unit, a reciprocal square root calculation unit, and an output matrix calculation unit. The output ends of the intermediate matrix storage unit, the first mean calculation unit, and the reciprocal square root calculation unit are all connected to the output matrix calculation unit, and the output end of the first mean calculation unit is connected to the first square calculation unit. The output end of the first square calculation unit is connected to the reciprocal square root calculation unit. The output end of the second square calculation unit is connected to the second mean calculation unit. The output end of the second mean calculation unit is connected to the reciprocal square root calculation unit. By performing layer normalization processing through the above hardware accelerator, the delay can be effectively reduced, and the operation speed and efficiency of the Transformer neural network can be improved. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is a schematic diagram of the model architecture of the Transformer neural network;
[0044] Figure 2 It is a schematic diagram of the hardware architecture of the layer normalization processing hardware accelerator applied to the Transformer neural network disclosed in the embodiments of the present application;
[0045] Figure 3 It is a schematic diagram of the working process of the layer normalization processing method applied to the Transformer neural network disclosed in the embodiments of the present application. Specific Embodiments
[0046] In order to reduce the latency in the layer normalization process and improve the operation speed and efficiency of the Transformer neural network, the present application discloses a layer normalization processing hardware accelerator and method applied to the Transformer neural network through the following embodiments.
[0047] In the present application, the three inputs of the multi-head attention layer are respectively defined as Q, K, and V, and the input of the feed-forward layer is defined as X. The size of the input tensor X is the same as the sizes of the input tensors Q, K, and V, and they are all equal to [batch_size, s, dmodel] (batch_size represents how many input sequences there are at a time, s represents how many words there are in an input sequence, and the size of dmodel represents the size of the neural network model). Considering the case where batch_size is 1, the input tensor can be considered to degenerate into a matrix, so all operations can be considered as operations on the input matrix (even if batch_size is greater than 1, it can be considered as multiple input matrices with the same size but different elements, and the same and non-interfering operations are performed on them and then combined together).
[0048] The first embodiment of the present application discloses a layer normalization processing hardware accelerator applied to the Transformer neural network. Refer to Figure 2 the structural schematic diagram shown, the layer normalization processing hardware accelerator includes:
[0049] an intermediate matrix storage unit, a first mean calculation unit, a second mean calculation unit, a first square calculation unit, a second square calculation unit, a square root reciprocal calculation unit, and an output matrix calculation unit.
[0050] The output end of the intermediate matrix storage unit is connected to the output matrix calculation unit.
[0051] The output end of the first mean calculation unit is respectively connected to the first square calculation unit and the output matrix calculation unit.
[0052] The output end of the first square calculation unit is connected to the square root reciprocal calculation unit.
[0053] The output end of the second square calculation unit is connected to the second mean calculation unit.
[0054] The output end of the second mean calculation unit is connected to the square root reciprocal calculation unit.
[0055] The output end of the square root reciprocal calculation unit is connected to the output matrix calculation unit.
[0056] Further, the intermediate matrix storage unit is used to obtain and store the intermediate matrix, where the intermediate matrix is the first intermediate matrix in the processing of the multi-head attention layer or the second intermediate matrix in the processing of the feed-forward layer.
[0057] The first mean calculation unit is used to calculate the mean of each row of elements in the intermediate matrix and input the calculation result into the first square calculation unit.
[0058] The first square calculation unit is used to perform a square operation on the value input by the first mean calculation unit to obtain the square of the mean of each row of elements in the intermediate matrix.
[0059] The second square calculation unit is used to perform a square operation on each element in the intermediate matrix to obtain a square matrix.
[0060] The second mean calculation unit is used to calculate the mean of each row of elements in the square matrix.
[0061] The square root reciprocal calculation unit is used to obtain the reciprocal of the square root of the variance of each row of elements in the intermediate matrix according to the square of the mean of each row of elements in the intermediate matrix and the mean of each row of elements in the square matrix.
[0062] The output matrix calculation unit is used to perform layer normalization processing on each element of the intermediate matrix, the mean of each row of elements in the intermediate matrix, and the reciprocal of the square root of the variance of each row of elements in the intermediate matrix to obtain the final output matrix of the multi-head attention layer or the feed-forward layer.
[0063] Further, the first mean calculation unit includes a plurality of first mean calculation sub-units, the second mean calculation unit includes a plurality of second mean calculation sub-units, the first square calculation unit includes a plurality of first square calculation sub-units, the second square calculation unit includes a plurality of second square calculation sub-units, the square root reciprocal calculation unit includes a plurality of square root reciprocal calculation sub-units, and the output matrix calculation unit includes a plurality of output matrix calculation sub-units.
[0064] The number of the first mean calculation sub-units, the second mean calculation sub-units, the first square calculation sub-units, the second square calculation sub-units, the square root reciprocal calculation sub-units, and the output matrix calculation sub-units is the same as the number of rows of any input matrix in the multi-head attention layer.
[0065] In the embodiment of the present application, the input of the layer normalization function operation module is an intermediate matrix G of size s×dmodel, and the output is also a matrix of the same size (named Output).
[0066] Currently, the following formula is usually used to calculate the mean of the i-th row elements of the intermediate matrix:
[0067] 。
[0068] The variance of the elements in the i-th row of the intermediate matrix is usually calculated using the following formula:
[0069] 。
[0070] In the process of performing layer normalization through the above formula, in order to obtain the variance value of each element in each row of the intermediate matrix, it is necessary to first calculate the average value of each element in each row of the intermediate matrix, then separately obtain the difference between each element and the average value, square the differences and then perform an accumulation operation. Such a calculation process has relatively cumbersome steps. In the actual processing process, it will cause a large delay, increase the operation time of the Transformer neural network, and reduce the operation efficiency of the Transformer neural network.
[0071] In the embodiments of the present application, in order to reduce the delay, an optimization method is proposed to calculate the variance of the elements in the i-th row of the intermediate matrix using another method. The calculation formula is as follows:
[0072] 。
[0073] Based on the above optimization method, in the process of obtaining the reciprocal of the square root of the variance of each element in each row of the intermediate matrix according to the square of the average value of each element in each row of the intermediate matrix and the average value of each element in each row of the square matrix, the square root reciprocal calculation unit is used to obtain the variance of each element in each row of the intermediate matrix according to the following formula:
[0074] 。
[0075] 。
[0076] Wherein, represents the variance of the elements in the i-th row of the intermediate matrix G, represents the average value of the elements in the i-th row of the intermediate matrix G, represents the average value of the squares of the elements in the i-th row of the square matrix, represents the element in the i-th row and k-th column of the intermediate matrix, represents the total number of columns of the intermediate matrix.
[0077] Further, the output matrix calculation unit is used to perform layer normalization on each element of the intermediate matrix, the average value of each element in each row of the intermediate matrix, and the reciprocal of the square root of the variance of each element in each row of the intermediate matrix according to the following formula to obtain the final output matrix of the multi-head attention layer or the feed-forward layer:
[0078] 。
[0079] Among them, represents the element in the i-th row and j-th column of the output matrix, represents the variance of the elements in the i-th row of the intermediate matrix G, represents the element in the i-th row and j-th column of the intermediate matrix G, represents the mean value of the elements in the i-th row of the intermediate matrix, and ε is the first parameter, represents the second parameter, represents the third parameter. ε is used to prevent the denominator from being equal to zero so that the operation result becomes infinite, and its value is 10 -8 . The second parameter includes ones ( , , …, , …, ), which are respectively used to calculate the elements of different columns of the output matrix. The third parameter includes ones ( , , …, , …, ), which are respectively used to calculate the elements of different columns of the output matrix. Both the second parameter and the third parameter are preset values.
[0080] The layer normalization processing hardware accelerator for the Transformer neural network disclosed in the above embodiment includes an intermediate matrix storage unit, a first mean value calculation unit, a second mean value calculation unit, a first square calculation unit, a second square calculation unit, a square root reciprocal calculation unit, and an output matrix calculation unit. By performing layer normalization processing through this hardware accelerator, the delay can be effectively reduced, and the operation speed and efficiency of the Transformer neural network can be improved.
[0081] The second embodiment of the present application discloses a layer normalization processing method for a Transformer neural network. The layer normalization processing method is applied to the layer normalization processing hardware accelerator for the Transformer neural network described in the first embodiment of the present application. Refer to Figure 3 the schematic diagram of the working process shown, and the layer normalization processing method includes:
[0082] Step S11, input all the elements of the intermediate matrix into the intermediate matrix storage unit in column order. Among them, if the current operation belongs to the multi-head attention layer, the intermediate matrix is the first intermediate matrix; if the current operation is the feed-forward layer, the intermediate matrix is the second intermediate matrix.
[0083] Step S12: Input each row element of the intermediate matrix into multiple first mean calculation sub-units respectively to calculate the mean value of each row element in the intermediate matrix, and input each row element of the intermediate matrix into multiple second square calculation sub-units respectively to obtain a square matrix.
[0084] Step S13: Input the mean value of each row element in the intermediate matrix into multiple first square calculation sub-units respectively to obtain the square of the mean value of each row element in the intermediate matrix.
[0085] Step S14: Input each row element of the square matrix into multiple second mean calculation sub-units respectively to calculate the mean value of each row element in the square matrix.
[0086] Step S15: Input the square of the mean value of each row element in the intermediate matrix and the mean value of each row element in the square matrix into multiple square root reciprocal calculation sub-units respectively to obtain the reciprocal of the square root of the variance of each row element in the intermediate matrix.
[0087] Step S16: Input each element of the intermediate matrix, the mean value of each row element in the intermediate matrix, and the reciprocal of the square root of the variance of each row element in the intermediate matrix into multiple output matrix calculation sub-units respectively to obtain the final output matrix of the multi-head attention layer or the feed-forward layer.
[0088] In one implementation, in combination with Figure 2 the disclosed structural diagram, the specific implementation process of the layer normalization processing method disclosed in the above embodiments is as follows:
[0089] Input the intermediate matrix G into the layer normalization processing hardware accelerator. Each time, input a column of elements of this matrix, that is, input G(1,1)-G(s,1) at the first moment, input G(1,j)-G(s,j) at the jth moment, and so on, until input G(1,dmodel)-G(s,dmodel) at the last moment. At the same time, the intermediate matrix storage unit, the first mean calculation unit, the second square calculation unit, and the second mean calculation unit in the layer normalization processing hardware accelerator perform the following operations: Store the intermediate matrix G in the "intermediate matrix storage unit"; accumulate and calculate 、 、……、 ; accumulate and calculate 、 、……、 ; After the input of the intermediate matrix G is completed, E(G,1)= 、E(G,2)= 、……、E(G,s)= , and E(G.*G,1)= 、E(G.*G,2)= …, E(G.*G, s) = ; Using the first square calculation unit, calculate E(G, 1) 2 , E(G, 2) 2 and E(G, s) 2 .
[0090] According to the operation results of the first mean calculation unit, the second mean calculation unit and the first square calculation unit, use the adder in the reciprocal square root calculation unit to calculate var(G, 1) = E(G, 1) 2 - E(G.*G, 1), var(G, 2) = E(G, 2) 2 - E(G.*G, 2), …, var(G, s) = E(G, s) 2 - E(G.*G, s), and then use the "x^(-0.5)" operation unit to find r1 = (var(G, 1) + ε)^(-0.5), r2 = (var(G, 2) + ε)^(-0.5), …, r s = (var(G, s) + ε)^(-0.5).
[0091] According to the operation results of the intermediate matrix storage unit, the first mean calculation unit and the reciprocal square root calculation unit, the output matrix calculation unit calculates the final output matrix according to the formula Output(i, j) = . The output at the first moment is Output(1, 1), Output(2, 1), …, Output(s, 1), the output at the second moment is Output(1, 2), Output(2, 2), …, Output(s, 2), until the dmodel-th moment, the output is Output(1, dmodel), Output(2, dmodel), …, Output(s, dmodel), and the final output matrix of the layer normalization processing hardware accelerator is obtained.
[0092] The third embodiment of the present application discloses a computer device, including:
[0093] A memory for storing a computer program.
[0094] A processor for implementing the steps of the layer normalization processing method applied to the Transformer neural network as described in the second embodiment of the present application when executing the computer program.
[0095] The fourth embodiment of the present application discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is processed and executed, it implements the steps of the layer normalization processing method applied to the Transformer neural network as described in the second embodiment of the present application.
[0096] The present application has been described in detail above in conjunction with specific embodiments and exemplary examples, but these descriptions should not be construed as limiting the present application. Those skilled in the art understand that, without departing from the spirit and scope of the present application, various equivalent replacements, modifications or improvements can be made to the technical solutions and implementation manners of the present application, and these all fall within the scope of the present application. The protection scope of the present application is subject to the appended claims.
Claims
1. A layer normalization processing hardware accelerator applied to a Transformer neural network, characterized in that, The layer normalization processing hardware accelerator includes: an intermediate matrix storage unit, a first mean calculation unit, a second mean calculation unit, a first square calculation unit, a second square calculation unit, a reciprocal square root calculation unit, and an output matrix calculation unit; The output end of the intermediate matrix storage unit is connected to the output matrix calculation unit; The output ends of the first mean calculation unit are respectively connected to the first square calculation unit and the output matrix calculation unit; The output end of the first square calculation unit is connected to the reciprocal square root calculation unit; The output end of the second square calculation unit is connected to the second mean calculation unit; The output end of the second mean calculation unit is connected to the reciprocal square root calculation unit; The output end of the reciprocal square root calculation unit is connected to the output matrix calculation unit; The intermediate matrix storage unit is used to obtain and store the intermediate matrix, where the intermediate matrix is the first intermediate matrix in the multi-head attention layer processing or the second intermediate matrix in the feed-forward layer processing; The first mean calculation unit is used to calculate the mean of each row of elements in the intermediate matrix and input the calculation result into the first square calculation unit; The first square calculation unit is used to perform a square operation on the value input by the first mean calculation unit to obtain the square of the mean of each row of elements in the intermediate matrix; The second square calculation unit is used to perform a square operation on each element in the intermediate matrix to obtain a square matrix; The second mean calculation unit is used to calculate the mean of each row of elements in the square matrix; The reciprocal square root calculation unit is used to obtain the reciprocal of the square root of the variance of each row of elements in the intermediate matrix according to the square of the mean of each row of elements in the intermediate matrix and the mean of each row of elements in the square matrix, and is used to obtain the variance of each row of elements in the intermediate matrix according to the following formula: ; ; Among them, represents the variance of the elements in the i-th row of the intermediate matrix G, represents the mean value of the elements in the i-th row of the intermediate matrix G, represents the mean value of the squares of the elements in the i-th row of the square matrix, represents the element in the i-th row and k-th column of the intermediate matrix, represents the total number of columns of the intermediate matrix; The output matrix calculation unit is used to perform layer normalization processing on each element of the intermediate matrix, the mean of each row of elements in the intermediate matrix, and the reciprocal of the square root of the variance of each row of elements in the intermediate matrix according to the following formula to obtain the final output matrix of the multi-head attention layer or the feed-forward layer: ; Among them, represents the element in the \(i\)-th row and \(j\)-th column of the said output matrix, represents the variance of the elements in the \(i\)-th row of the said intermediate matrix \(G\), represents the element in the \(i\)-th row and \(j\)-th column of the said intermediate matrix \(G\), represents the mean value of the elements in the \(i\)-th row of the said intermediate matrix, and \(\varepsilon\) is the first parameter, represents the second parameter, represents the third parameter.
2. The layer normalization processing hardware accelerator applied to the Transformer neural network according to claim 1, wherein The first mean calculation unit includes a plurality of first mean calculation sub-units, the second mean calculation unit includes a plurality of second mean calculation sub-units, the first square calculation unit includes a plurality of first square calculation sub-units, the second square calculation unit includes a plurality of second square calculation sub-units, the reciprocal square root calculation unit includes a plurality of reciprocal square root calculation sub-units, and the output matrix calculation unit includes a plurality of output matrix calculation sub-units; The number of the first mean calculation sub-units, the second mean calculation sub-units, the first square calculation sub-units, the second square calculation sub-units, the reciprocal square root calculation sub-units, and the output matrix calculation sub-units is the same as the number of rows of any input matrix in the multi-head attention layer.
3. A layer normalization processing method applied to a Transformer neural network, characterized in that, The layer normalization processing method is applied to the layer normalization processing hardware accelerator for Transformer neural network according to claim 1 or 2, and the layer normalization processing method includes: All elements of the intermediate matrix are sequentially input into the intermediate matrix storage unit column by column. Among them, if the current operation belongs to the multi-head attention layer, the intermediate matrix is the first intermediate matrix; if the current operation is the feed-forward layer, the intermediate matrix is the second intermediate matrix. Each row of elements of the intermediate matrix is respectively input into a plurality of first mean calculation sub-units to calculate the mean of each row of elements in the intermediate matrix; and each row of elements of the intermediate matrix is respectively input into a plurality of second square calculation sub-units to obtain a square matrix. The mean of each row of elements in the intermediate matrix is respectively input into a plurality of first square calculation sub-units to obtain the square of the mean of each row of elements in the intermediate matrix. Each row of elements in the square matrix is respectively input into a plurality of second mean calculation sub-units to calculate the mean of each row of elements in the square matrix. The square of the mean of each row of elements in the intermediate matrix and the mean of each row of elements in the square matrix are respectively input into a plurality of square root reciprocal calculation sub-units to obtain the reciprocal of the square root of the variance of each row of elements in the intermediate matrix. Each element of the intermediate matrix, the mean of each row of elements in the intermediate matrix, and the reciprocal of the square root of the variance of each row of elements in the intermediate matrix are respectively input into a plurality of output matrix calculation sub-units to obtain the final output matrix of the multi-head attention layer or the feed-forward layer.
4. A computer device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the layer normalization processing method applied to the Transformer neural network as described in claim 3 when executing the computer program.
5. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is processed and executed, it implements the steps of the layer normalization processing method applied to the Transformer neural network as described in claim 3.
Citation Information
Patent Citations
Information processing apparatus, neural network program, and processing method for neural network
CN111353578A
Independent component analysis processor
US20120158367A1