Calculation method for neural network based on conversion model and electronic device thereof
By pre-calculating and storing the merge weights of the weight matrix, combined with low-rank decomposition and merging technology, the problem of high computing burden on the neural network of the transformation model is solved, and efficient computing performance and accuracy are achieved.
Patent Information
- Application Number
- CN202510111185.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-01
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-29
AI Technical Summary
The neural networks of the existing transformation model have too much computing burden and memory access on devices with limited resources, making it difficult to realize instant application.
By pre-calculating and storing the combined weights of the weight matrix, reducing online computing time and memory access, using weight merge technology to generate attention scores, and perform low-rank decomposition and merge calculations when crossing the attention layer and the multi-layer perceptron layer.
While reducing the number of parameters and the number of operations, it improves computing efficiency, reduces hardware resource consumption, and ensures calculation accuracy.
Smart Images

Figure CN120386975A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a calculation method for a neural network based on a transformation model and an electronic device thereof. Background Art
[0002] A transformation model is a deep learning architecture that has revolutionized natural language processing in various tasks and achieved state-of-the-art results. The attention mechanism is one of the cores of the transformation model, which allows the deep learning model to focus on relevant parts of the input data. In a general transformation model, the attention weights are calculated by the dot product of the weighted query input and the weighted key input, followed by scaling calculation and softmax function calculation, and the attention scores are obtained by multiplying the attention weights with the matrix of the weighted value input, followed by weight calculation. However, the computational burden and memory access amount caused by this architecture make it difficult to achieve real-time applications on devices with limited resources. Summary of the Invention
[0003] The present invention proposes a calculation method for a neural network based on a transformation model and an electronic device thereof.
[0004] In an embodiment of the present invention, the above method includes receiving a query input, a key input, and a value input from the previous layer of the attention layer, obtaining a first combined weight pre-calculated based on the weight matrix for the query and the weight matrix for the key, obtaining a second combined weight pre-calculated based on the weight matrix for the value and the weight matrix for the output score, and performing calculations based on the query input, the key input, the value input, the first combined weight, and the second combined weight to generate the attention score of the attention layer.
[0005] In an embodiment of the present invention, the above electronic device includes a processor, which is used to receive a query input, a key input, and a value input from the previous layer of the attention layer, obtain a first combined weight pre-calculated based on the weight matrix for the query and the weight matrix for the key, obtain a second combined weight pre-calculated based on the weight matrix for the value and the weight matrix for the output score, and perform calculations based on the query input, the key input, the value input, the first combined weight, and the second combined weight to generate the attention score of the attention layer.
[0006] The present invention proposes a calculation method for crossing the attention layer and the multi-layer perceptron layer of a neural network based on a transformation model and an electronic device thereof.
[0007] In one embodiment of the present invention, the above method includes receiving the residual of the attention layer and the attention matrix, obtaining the low-order weight matrix of the first output score and the low-order weight matrix of the second output score generated by the low-order decomposition of the weight matrix of the output score, and obtaining the third combined weight, and performing calculations based on the residual, the attention matrix, the low-order weight matrix of the first output score, and the third combined weight to generate the output of the specified path of the multi-layer perceptron layer, where the third combined weight is pre-calculated based on the low-order weight matrix of the second output score and the weight matrix of the multi-layer perceptron layer.
[0008] In one embodiment of the present invention, the above electronic device includes a processor, which is used to receive the residual of the attention layer and the attention matrix, obtain the low-order weight matrix of the first output score and the low-order weight matrix of the second output score generated by the low-order decomposition of the weight matrix of the output score, and obtain the third combined weight, and perform calculations based on the residual, the attention matrix, the low-order weight matrix of the first output score, and the third combined weight to generate the output of the specified path of the multi-layer perceptron layer, where the third combined weight is pre-calculated based on the low-order weight matrix of the second output score and the weight matrix of the multi-layer perceptron layer. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a schematic diagram of an electronic device illustrated according to an embodiment of the present invention.
[0010] Figure 2 is a flowchart of a calculation method for a neural network based on a transformation model according to an embodiment of the present invention.
[0011] Figure 3 is a schematic diagram of a calculation method for an attention layer of a neural network based on a general transformation model in the prior art.
[0012] Figure 4 is a schematic diagram of a calculation method for a neural network based on a transformation model according to an embodiment of the present invention.
[0013] Figure 5 is a flowchart of a calculation method for crossing the attention layer and the multi-layer perceptron layer of a neural network based on a transformation model according to an embodiment of the present invention.
[0014] Figure 6 is a schematic diagram of a calculation method for the attention layer and the multi-layer perceptron layer of a neural network based on a general transformation model in the prior art.
[0015] Figure 7 is a schematic diagram of a calculation method for crossing the attention layer and the multi-layer perceptron layer of a neural network based on a transformation model according to an embodiment of the present invention.
[0016] Figure 8 It is a schematic diagram of a calculation method for crossing the attention layer and the multi-layer perceptron layer of a neural network based on a transformation model according to another embodiment of the present invention. Detailed implementation manners
[0017] Now, reference will be made in detail to the exemplary embodiments of the present invention. Examples of the exemplary embodiments are illustrated in the accompanying drawings. Whenever possible, the same element symbols are used in the drawings and the description to represent the same or similar parts.
[0018] Figure 1 It is a schematic diagram of an electronic device according to an embodiment of the present invention. Figure 1 First, each component and the configuration relationship in the electronic device will be introduced. The detailed functions will be disclosed in conjunction with the subsequent embodiments' procedures. Figure 1 and disclosed.
[0019] Please refer to Figure 1 , the electronic device 100 of this embodiment at least includes a processor 110 and a memory 120. The electronic device 100 may be an electronic system or a computer system. The processor 110 is used to perform calculations on the neural network based on the transformation model. It may be a central processing unit (CPU), a graphic processing unit (GPU), an application processor (AP), a programmable general or special-purpose microprocessor, a digital signal processor (DSP), a field programmable array (FPGA), an application-specific integrated circuit (ASIC), other similar devices, integrated circuits, or a combination thereof. The memory 120 is used to store data. It may be various forms of random access memory (RAM), such as dynamic random access memory (DRAM), or other similar devices, integrated circuits, or a combination thereof.
[0020] Figure 2 It is a flowchart of a calculation method for a neural network based on a transformation model according to an embodiment of the present invention, where Figure 2 the steps of Figure 1 can be implemented by the electronic device 100 of
[0021] Please also refer to Figure 2and Figure 1 The processor 110 receives a query input, a key input, and a value input from the layer preceding the attention layer (step S202). Here, the query input, the key input, and the value input may be feature maps output from a convolutional layer of the neural network.
[0022] Next, the processor 110 obtains a first combined weight pre-computed based on the weight matrix of the query and the weight matrix of the key (step S204), and obtains a second combined weight pre-computed based on the weight matrix of the value and the weight matrix of the output score (step S206). In this embodiment, the first combined weight and the second combined weight may be pre-computed offline and pre-stored in the memory 120 to reduce the number of parameters and the number of operations, thereby reducing the online calculation time, the memory access amount, the quantization error, and the hardware resources. Then, the processor 110 performs calculations based on the query input, the key input, the value input, the first combined weight, and the second combined weight to generate the attention score of the attention layer (step S208), thereby solving the problems in the prior art. More details will be described in full later.
[0023] Figure 3 is a schematic diagram of a calculation method of an attention layer for a neural network based on a general transformation model in the prior art.
[0024] Please refer to Figure 3 , the attention weight XS is calculated by the inner product of the weighted query input Q (the query input X Q given the linear weight W Q ) and the weighted key input K (the key input X K given the linear weight W K ), and then a scaling calculation and the softmax function are applied. The attention score X O is calculated by the matrix multiplication of the attention matrix X h and the weighted value input V (the value input X V given the linear weight W V ), and then the linear weight W O is given.
[0025] The goal of the method proposed in this embodiment is to improve the calculations of 310 and 320 in Figure 3 through weight combination techniques (hereinafter referred to as "QK combination" and "VO combination") respectively.
[0026] The inner product calculation in 310 can be rewritten as , where and 。The QK merge here will involve substitution, expansion, and merging, where the inner product can be further rewritten as the following equation (1):
[0027]
[0028] It should be noted that W QK represents the aforementioned first merge weight, which is the weight matrix W Q for queries and the transpose of the weight matrix W K T for keys.
[0029] The multiplication calculation in 320 can be rewritten as , where . The proposed VO merge will involve substitution and merging, where the multiplication calculation can be further rewritten as the following equation (2):
[0030]
[0031] It should be noted that W VO represents the aforementioned second merge weight, which is the weight matrix W V for values and the weight matrix W O for output scores.
[0032] Figure 4 is a schematic diagram of a calculation method for a neural network based on a conversion model according to an embodiment of the present invention, where Figure 4 the steps of Figure 1 can be implemented by the electronic device 100 of
[0033] Please also refer to Figure 4 and Figure 1 . In this embodiment, the processor 110 will obtain the first merge weight W QK pre-computed and pre-stored in the memory 120 VO and the second merge weight W Q . It should be noted that the calculation in 410 corresponds to the QK merge derived in equation (1). The processor 110 will perform a multiplication calculation on the query input X VO , the first merge weight W KT and the transpose of the key input X S to generate a first multiplication result. Next, the processor 110 will calculate the first multiplication result to generate the attention weight XFor example, the processor 110 may perform a scaling calculation on the first multiplication result to generate a scaled multiplication result, and may further apply the softmax function to the scaled multiplication result to generate the attention weight X S In addition, the calculation in 420 corresponds to the VO merging derived in Equation (2). The processor 110 will perform an operation on the attention weight X S the numerical input X V and the second merging weight W VO to perform a multiplication calculation to generate the attention score X of the attention layer O .
[0034] In terms of performance, assuming that the dimensions of X Q and X K are both N×D, the dimensions of W Q and W K are both D×d, and the dimensions of Q and K are both N×d. Compared with the prior art, the technology proposed in this embodiment only requires the number of operations (i.e., only requires fewer instructions). Intuitively, assume N = 1024, D = 384, and d = 384. Compared with the prior art, the technology proposed in this embodiment can reduce the number of parameters by 62% (i.e., reduce latency). In addition, since the technology proposed in this embodiment does not use approximations, the accuracy can be guaranteed
[0035] In another embodiment, in order to further reduce the number of parameters, the first merging weight may be pre-computed based on the low-rank decomposition of the weight matrix for the query and the low-rank decomposition of the weight matrix for the key. Specifically, Equation (1) can be further decomposed into low rank as shown in the following Equation (3):
[0036]
[0037] That is, the low-rank decomposition of the weight matrix W Q for the query will produce the first low-rank weight matrix U Q for the query and the second low-rank weight matrix S Q , where the low-rank decomposition of the weight matrix W K for the key will produce the first low-rank weight matrix U K for the key and the second low-rank weight matrix S K . In this case, the first merging weight will be the product of the first merged low-rank weight U Q ' and the transpose of the first low-rank weight matrix S K T for the key, where the first low-rank merged weight UQ ' is the first low-rank weight matrix U for querying Q , the second low-rank weight matrix S for querying Q and the first low-rank weight matrix U for key K T is the product of the transpose of
[0038] The weight merging mechanism can also be applied to other layers in the neural network based on the transformation model other than the attention layer. For example, Figure 5 is a flowchart of a calculation method for spanning the attention layer and the multi-layer perceptron layer of the neural network based on the transformation model according to an embodiment of the present invention, where Figure 5 The step of Figure 1 can also be implemented by the electronic device 100 of
[0039] Please refer to Figure 5 and Figure 1 at the same time. The processor 110 will receive the residue of the attention layer and the attention matrix (step S502). The processor 110 will obtain the first low-rank weight matrix for outputting scores and the second low-rank weight matrix for outputting scores generated by the low-rank decomposition of the weight matrix for outputting scores, and obtain the third merging weight, where the third merging weight is pre-calculated based on the second low-rank weight matrix for outputting scores and the weight matrix of the multi-layer perceptron layer (step S504). In this embodiment, the third merging weight can be pre-calculated offline and pre-stored in v120 to reduce the number of parameters and the number of operations. The processor 110 will perform calculations based on the residue, the attention matrix, the first low-rank weight matrix for outputting scores, and the third merging weight, so as to generate the output of the specified path of the multi-layer perceptron layer (step S506). Here, the specified path is the path spanning the attention layer and the multi-layer perceptron layer. More details will be described in full later.
[0040] Figure 6 is a schematic diagram of a method for calculating the attention layer and the multi-layer perceptron layer of a general neural network based on the transformation model in the prior art.
[0041] Please refer to Figure 6 , the residue X r will be added to the attention score of the attention layer 601 to produce an addition result, where the attention score is the product of the attention weight X h and the weight matrix W for outputting scores O . The addition result will be divided into two paths: the high-frequency signal of the addition result will be input to the multi-layer perceptron layer 602 for layer normalization calculation and given the linear weight W G, and the low-frequency signal of the addition result (i.e., X m ) will be directly output.
[0042] The goal of the method proposed in this embodiment is to improve the operations performed in the path spanning two layers through a weight merging technique (referred to as "SOG merging"). In addition, the layer normalization calculation is a linear transformation, so it can be postponed (i.e., calculated after the weight calculation).
[0043] The SOG merging proposed in this embodiment will involve the inverse matrix, low-rank decomposition, and merging, where the calculation spanning two layers can be written as the following equation (4):
[0044]
[0045] Here, X g represents the output of a specified path of the multi-layer perceptron layer, and W SOG represents the aforementioned third merging weight matrix, which is the product of the second low-rank weight matrix S O for outputting scores and the weight matrix W G of the multi-layer perceptron layer.
[0046] Figure 7 is a schematic diagram of a calculation method for spanning the attention layer and the multi-layer perceptron layer of a neural network based on a transformation model according to an embodiment of the present invention, where Figure 7 the method can be implemented by the Figure 1 electronic device 100.
[0047] Please also refer to Figure 7 and Figure 1 , the processor 110 will obtain from the memory 120 the first low-rank weight matrix U O for outputting scores, the second low-rank weight matrix S O for outputting scores, the inverse matrix of the product of the first low-rank weight matrix for outputting scores and the second low-rank weight matrix for outputting scores (U O S O ), and obtain the third merging weight W -1 . In addition, based on double-precision calculation, (U SOG S O ) O can also be replaced by (W -1 ) O . -1 .
[0048] In the path 710 without passing through the multi-layer perceptron layer, the processor 110 will calculate the residual X r and the attention matrix X h , the first low-rank weight matrix U for outputting scoresO and the second lowest-rank weight matrix S for output scores O The products are summed to generate the output X of path 710 of the multi-layer perceptron layer m . These operations can also be represented by the following equation (5):
[0049]
[0050] In the second path 720 spanning the attention layer and the multi-layer perceptron layer, the processor 110 combines the attention matrix X h and the residual X r with the inverse matrix of the product of the first low-rank weight matrix for output scores and the second low-rank weight matrix for output scores (U O S O ) -1 The products are summed to generate a first intermediate result. Then, the processor 110 multiplies the first intermediate result, the first low-rank weight matrix U for output scores O and the third combination weight W SOG to generate a second intermediate result. Then, the processor 110 performs layer normalization on the second intermediate result to generate the output X of path 720 of the multi-layer perceptron layer g .
[0051] In terms of performance, assume that the dimensions of U O and S O are D×r and r×D respectively, and the dimension of W O -1 is 2(D×D), and the dimension of WSOG is r×8D. The method proposed in this embodiment only requires the number of operations compared to the prior art. Intuitively, assuming the rank r = D / 4, the method proposed in this embodiment can reduce the number of parameters by 62% compared to the prior art.
[0052] As another more aggressive combination method, equation (4) can do additional combination as presented in the following equation (6):
[0053]
[0054] Here, W OU is the fourth combination weight, which is the product of the inverse matrix of the product of the first low-rank weight matrix for output scores and the second low-rank weight matrix for output scores (U O S O ) -1 and the first low-rank weight matrix U for output scores O .
[0055] Figure 8 It is a schematic diagram of a calculation method for crossing the attention layer and the multi-layer perceptron layer of a neural network based on a transformation model according to another embodiment of the present invention, where Figure 8 The method can be implemented by Figure 1 the electronic device 100.
[0056] Please also refer to Figure 8 and Figure 1 , the processor 110 will obtain the first low-rank weight matrix U for outputting scores from the memory 120 O , the second low-rank weight matrix S for outputting scores O , the third merging weight W SOG and the fourth merging weight W OU .
[0057] In the path 810 that does not pass through the multi-layer perceptron layer, the processor 110 will add and calculate the product of the residual X r and the attention matrix X h , the first low-rank weight matrix U for outputting scores O and the second low-rank weight matrix S for outputting scores O to generate the output X m of the path 810 of the multi-layer perceptron layer.
[0058] In the path 820 that crosses the attention layer and the multi-layer perceptron layer, the processor 110 will add and calculate the product of the residual X r and the fourth merging weight W OU and the product of the attention matrix X h and the first low-rank weight matrix U of the output scores O to generate a first intermediate result.
[0059] Next, the processor 110 will multiply the first intermediate result and the third merging weight W SOG to generate a second intermediate result. The processor 110 will perform layer normalization calculation on the second intermediate result to generate the output of the first path of the multi-layer perceptron layer, and further generate the output X g of the path 820 of the multi-layer perceptron layer.
[0060] In terms of performance, assuming that the dimensions of UO and SO are D×r and r×D respectively, the dimension of W OU is 2(D×r), and the dimension of W SOG is r×8D. The method proposed in this embodiment only requires The number of operations. Intuitively, assuming a rank r = D / 4, the method proposed in this embodiment can reduce the number of parameters by 62% compared with the prior art.
[0061] In summary, the present invention proposes various effective ways to perform calculations on neural networks based on transformation models, so as to reduce online calculation time, memory access volume, quantization error, and resource consumption while ensuring accuracy.
[0062] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A calculation method for an attention layer of a neural network based on a transformation model, comprising: Receiving a query input, a key input, and a value input from the previous layer of the attention layer; Obtaining a first combined weight, wherein the first combined weight is pre-calculated based on a weight matrix for the query and a weight matrix for the key; Obtaining a second combined weight, wherein the second combined weight is pre-calculated based on a weight matrix for the value and a weight matrix for the output score; And Performing calculations based on the query input, the key input, the value input, the first combined weight, and the second combined weight to generate an attention score of the attention layer.
2. The method according to claim 1, wherein the first combined weight is a product of the weight matrix for the query and the transpose of the weight matrix for the key.
3. The method according to claim 1, wherein the second combined weight is a product of the weight matrix for the value and the weight matrix for the output score.
4. The method according to claim 1, wherein the step of performing calculations based on the query input, the key input, the value input, the first combined weight, and the second combined weight to generate the attention score of the attention layer comprises: Performing a multiplication calculation on the query input, the first combined weight, and the transpose of the key input to generate a first multiplication result; Performing calculations on the first multiplication result to generate an attention weight; And Performing a matrix multiplication calculation on the attention weight, the value input, and the second combined weight to generate the attention score of the attention layer.
5. The method according to claim 4, wherein the step of performing calculations on the first multiplication result to generate the attention weight comprises: Performing a scaling calculation on the first multiplication result to generate a scaled multiplication result; And Applying a softmax function to the scaled multiplication result to generate the attention weight.
6. The method according to claim 1, wherein the first combined weight is pre-calculated based on a low-rank decomposition of the weight matrix for the query and a low-rank decomposition of the weight matrix for the key.
7. The method according to claim 6, wherein the low-rank decomposition of the weight matrix for the query generates a first low-rank weight matrix for the query and a second low-rank weight matrix for the query, wherein the transpose of the low-rank decomposition of the weight matrix for the key generates the transpose of a first low-rank weight matrix for the key and the transpose of a second low-rank weight matrix for the key, wherein the first combined weight is a product of a first combined low-rank weight and the transpose of the first low-rank weight matrix for the key, wherein the first combined low-rank weight is a product of the first low-rank weight matrix for the query, the second low-rank weight matrix for the query, and the transpose of the first low-rank weight matrix for the key.
8. A calculation method for spanning an attention layer and a multi-layer perceptron layer of a neural network based on a transformation model, comprising: Receive the residual of the attention layer and the attention matrix; Obtain a first low-rank weight matrix for outputting scores and a second low-rank weight matrix for outputting scores generated by low-rank decomposition of a weight matrix for outputting scores, and obtain a third merging weight, where the third merging weight is pre-calculated based on the second low-rank weight matrix for outputting scores and the weight matrix of the multi-layer perceptron layer; And Perform calculations based on the residual, the attention matrix, the first low-rank weight matrix for outputting scores, and the third merging weight to generate the output of a specified path of the multi-layer perceptron layer.
9. The method according to claim 8, wherein the step of performing calculations based on the residual, the attention matrix, the first low-rank weight matrix for outputting scores, and the third merging weight to generate the output of the specified path of the multi-layer perceptron layer includes: Obtain the inverse matrix of the product of the first low-rank weight matrix for outputting scores and the second low-rank weight matrix for outputting scores; Perform a summation calculation on the attention matrix, the residual, and the product of the inverse matrix of the product of the first low-rank weight matrix for outputting scores and the second low-rank weight matrix for outputting scores to generate a first intermediate result; Perform a multiplication calculation on the first intermediate result, the first low-rank weight matrix for outputting scores, and the third merging weight to generate a second intermediate result; And Perform layer normalization calculation on the second intermediate result to generate the output of the specified path of the multi-layer perceptron layer.
10. The method according to claim 8, wherein the step of performing calculations based on the residual, the attention matrix, the first low-rank weight matrix for outputting scores, and the third merging weight to generate the output of the specified path of the multi-layer perceptron layer includes: Obtain a fourth merging weight, where the fourth merging weight is pre-calculated based on the first low-rank weight matrix for outputting scores and the inverse matrix of the product of the first low-rank weight matrix for outputting scores and the second low-rank weight matrix for outputting scores; Perform a summation calculation on the product of the residual and the fourth merging weight and the product of the attention matrix and the first low-rank weight matrix for outputting scores to generate a first intermediate result; Perform a multiplication calculation on the first intermediate result and the third merging weight to generate a second intermediate result; And Perform layer normalization calculation on the second intermediate result to generate the output of the specified path of the multi-layer perceptron layer.
11. The method according to claim 8, further comprising: Perform a summation calculation on the residual, the attention matrix, and the product of the first low-rank weight matrix for outputting scores and the second low-rank weight matrix for outputting scores to generate the output of another specified path of the multi-layer perceptron layer.
12. An electronic device for a neural network based on a transformation model, comprising: A processor for: Receiving a query input, a key input, and a value input from a previous layer of the attention layer; Obtaining a first combination weight, where the first combination weight is pre-computed based on a weight matrix for the query and a weight matrix for the key; Obtaining a second combination weight, where the second combination weight is pre-computed based on a weight matrix for the value and a weight matrix for the output score; And Performing calculations based on the query input, the key input, the value input, the first combination weight, and the second combination weight to generate an attention score of the attention layer.
13. The electronic device according to claim 12, further comprising: A memory for pre-storing the first combination weight.
14. An electronic device for spanning an attention layer and a multi-layer perceptron layer of a neural network based on a transformation model, comprising: A processor for: Receiving a residual of the attention layer and an attention matrix; Obtaining a first low-rank weight matrix for the output score and a second low-rank weight matrix for the output score generated by a low-rank factorization of a weight matrix for the output score, and obtaining a third combination weight, where the third combination weight is pre-computed based on the second low-rank weight matrix for the output score and a weight matrix of the multi-layer perceptron layer; And Performing calculations based on the residual, the attention matrix, the first low-rank weight matrix for the output score, and the third combination weight to generate an output of a specified path of the multi-layer perceptron layer.
15. The electronic device according to claim 14, further comprising: A memory for pre-storing the third combination weight and the first low-rank matrix for the output score and the second low-rank matrix for the output score generated by the low-rank factorization of the weight matrix for the output score.