A Method, Device, Equipment and Medium for Optimizing Self-Attention of Large Models
By chunking large matrices and small matrix weighted calculations that sort similarity, the problem of high computational complexity in the self-attention mechanism of large models is solved, and the accuracy and performance of the model are improved, especially the ability to capture long-distance context information.
Patent Information
- Application Number
- CN202510025947.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Existing large models have high computational complexity and time-consuming in self-attention mechanisms, which affects the accuracy and performance of the model, and are difficult to capture long-distance context information.
Divide the large matrix into several small matrices for local attention calculation, and select small matrices with high similarity sort for weighted calculations to reduce the computational complexity and capture long-distance context information.
The computational complexity of the large model is optimized, and the accuracy of attention output results and the overall performance of the model are improved.
Smart Images

Figure CN119443183B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and medium for optimizing self-attention of a large model. Background Art
[0002] Large models have extensive applications in the field of deep learning. Large models have a huge parameter scale and a complex structure, and a large amount of computing resources and storage space are required during the training process. For example, in the standard self-attention mechanism, for each element (token) of the input sequence, it is necessary to interact with all other elements, so there are n×n computing units. In actual deployment, large models usually rely on NPUs (neural processing units) for inference. As a kind of hardware specifically optimized for deep learning inference tasks, NPU can accelerate the forward calculation of the model and improve the inference efficiency. However, due to the large parameter scale of large models, even during the inference process, the computing power of NPU may still be insufficient to meet the requirements of efficient inference. This may lead to slow embedding calculation speed, long inference time, and large consumption of hardware resources, thus increasing the deployment cost of actual applications.
[0003] Existing large models usually have a multi-head self-attention structure. In order to reduce the computational complexity of the self-attention mechanism, most of them adopt a local attention mechanism, that is, only consider elements within a fixed range or select elements based on content; or perform singular value decomposition (SVD) on the model parameter matrix to achieve data dimensionality reduction and compression, so as to optimize the large model. However, the above methods still have a high computational complexity. Especially when dealing with large-scale datasets, it will increase the training time and computational cost of the model; or it is only limited to the local area and lacks the ability to capture long-distance context information; or select to retain some singular values and ignore other smaller singular values, resulting in local information loss, thus affecting the accuracy and performance of the model.
[0004] In view of this, the applicant has specifically proposed this application after studying the existing technologies. Summary of the Invention
[0005] The present invention aims to provide a method, device, equipment and medium for optimizing self-attention of a large model, so as to solve the disadvantages of high computational complexity and time-consuming in the existing methods, which affect the accuracy and performance of the model.
[0006] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0007] A method for optimizing self-attention of a large model, comprising:
[0008] S1, obtaining the KQ large matrix obtained after the input sequence passes through the self-attention structure of the large model;
[0009] S2. Divide the large KQ matrix into several small KQ matrices;
[0010] S3. Perform local attention calculation on each small KQ matrix; and perform a descending order sorting of the similarity with other small KQ matrices to obtain a sorted list;
[0011] S4. For each element in the input sequence, according to the calculated sorted list, select the local attention corresponding to the top r small KQ matrices with the highest similarity for weighted calculation, as the attention representation of the current element, until all elements of the input sequence are completed, to obtain the optimized attention result of the input sequence.
[0012] Preferably, the large KQ matrix is a similarity matrix of the key matrix K and the query matrix Q generated by the self-attention structure of the large model.
[0013] Preferably, the similarity is obtained by calculating the dot product of elements.
[0014] Preferably, the similarity is obtained by calculating the cosine similarity of elements.
[0015] Preferably, the S4 is specifically:
[0016] For each element in the input sequence, according to the calculated sorted list, select the top r small KQ matrices with the highest similarity to obtain a high-similarity attention set;
[0017] For each small KQ matrix in the high-similarity attention set, perform weighted calculation of attention according to a preset weight, and then use the softmax function for normalization, so as to obtain the attention weight vector of the current element;
[0018] Perform weighted summation of the attention weight vector of the current element on the value matrix V generated by the self-attention structure of the large model to obtain the global attention representation of the current element.
[0019] The present invention also provides a large model self-attention optimization device, including:
[0020] An acquisition unit, configured to acquire the large KQ matrix obtained after the input sequence passes through the self-attention structure of the large model;
[0021] A segmentation unit, configured to divide the large KQ matrix into several small KQ matrices;
[0022] A similarity sorting unit, configured to perform local attention calculation on each small KQ matrix; and perform a descending order sorting of the similarity with other small KQ matrices to obtain a sorted list;
[0023] An attention calculation unit is used to calculate the attention representation of each element; for each element in the input sequence, according to the calculated sorted list, the local attention corresponding to the top r KQ small matrices with the highest similarity is selected for weighted calculation as the attention representation of the current element until all elements of the input sequence are completed, and the optimized attention result of the input sequence is obtained.
[0024] The present invention also provides a large model self-attention optimization device, including a processor and a memory. A computer program is stored in the memory and can be executed by the processor to implement a large model self-attention optimization method as described above.
[0025] The present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, a large model self-attention optimization method as described above is implemented.
[0026] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0027] When calculating the attention of the input sequence, the present invention reduces the computational complexity when calculating the global attention of each element by performing block processing of the global large matrix into small matrices, reduces the operation between attention matrices, and optimizes the computational complexity of the large model. At the same time, the present invention selectively performs weighted calculation with distant matrices through the similarity sorted list to capture distant context information, improves the accuracy of the attention output result, and thus improves the accuracy and performance of the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is a schematic diagram of a large model self-attention optimization method provided for Embodiment 1.
[0030] Figure 2 It is a schematic structural diagram of a large model self-attention optimization device provided for Embodiment 2.
[0031] The following will further elaborate on the present invention in conjunction with the drawings and specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0033] Embodiment 1
[0034] Embodiment 1 of the present invention provides a method for optimizing the self-attention of a large model, which can be implemented by a large model self-attention optimization device (hereinafter referred to as the optimization device), and particularly, is executed by one or more processors in the optimization device.
[0035] In this embodiment, the optimization device can be an electronic device equipped with a processor, and the processor has a computer program for this large model self-attention optimization method and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited herein.
[0036] In this embodiment, the large model can be used to process input sequences, such as different types of data including text data, image data, audio-video data, etc., to obtain processing results.
[0037] Q (Query): Represents the query matrix, which is the vector representation of the word or data point we want to understand.
[0038] K (Key): Represents the key matrix, which is the vector representation of all words or data points in the input sequence and is used to compare with Q.
[0039] V (Value): Represents the value matrix, which is the vector representation of all words or data points in the input sequence, but after calculating the similarity between Q and K, it will be used for weighted summation to generate the final output.
[0040] In the self-attention mechanism, K, Q, and V are vector matrices obtained by linearly transforming the input sequence (such as the embedding vector of text data). Specifically, three weight matrices are first defined , and , these weight matrices are parameters learned during the training process. For each element in the input sequence (such as the word embedding vector of each word), its products with the weight matrices , and are calculated respectively to obtain , and for each element i.
[0041] The working principle of the standard self-attention mechanism is as follows: The attention weights are dynamically allocated by comparing the similarity between Q and K; the weighted sum of V is calculated according to the attention weights to obtain the attention output of the current element. This output is a weighted combination based on the most relevant information in the input sequence, thus improving the flexibility and efficiency in processing complex sequence data.
[0042] The expression of the standard self-attention mechanism is:
[0043] ;
[0044] where, represents the result of self-attention, T is the transpose, is the normalization function, represents the dimension of the key vector K.
[0045] The calculations of K, Q, and V enable the model to capture the dependencies between any two elements within the input sequence, regardless of how far apart they are in the sequence. This is very important for tasks such as understanding context-based meanings and dealing with long-range dependencies.
[0046] In the standard self-attention mechanism, for each element (token) in the input sequence, it needs to interact with all other elements. Its computational complexity includes the following aspects:
[0047] (1) Calculate Q (query), K (key), V (value): Assume the length of the input sequence is n and the representation dimension of each token is d. Then, the complexity of generating Q, K, and V is ( );
[0048] (2) Calculate the dot product of Q and K to form an n×n similarity matrix. The complexity of this step is ( );
[0049] (3) Next, usually, a softmax operation is also needed to calculate the weights. The complexity of this operation is ( );
[0050] (4) Calculate the weighted V: Apply the attention weights to V to obtain the output. The complexity of this step is ( )。
[0051] Therefore, the computational complexity of the standard self-attention mechanism is ( ), where n is the length of the input sequence and d is the dimension of the token.
[0052] In the standard self-attention mechanism, since each sequence element must interact with all other elements, there are n×n computational units. Although this global interaction can capture rich context information, it will cause a significant increase in computational and memory overhead when dealing with long sequences.
[0053] To reduce the computational complexity of the self-attention mechanism, as Figure 1 shown, this embodiment provides a large model self-attention optimization method, which includes steps S1 to S4.
[0054] S1. Obtain the KQ large matrix obtained after the input sequence passes through the large model self-attention structure.
[0055] In this embodiment, the large model is a trained deep learning model. According to different application scenarios, the input sequence can be different types of data such as text data, image data, audio-visual data, etc.
[0056] The KQ large matrix is the similarity matrix of the key matrix K and the query matrix Q generated by the large model self-attention structure. The similarity can be calculated by computing the dot product of K and Q, or cosine similarity, or other measurement methods, which are not limited here.
[0057] S2. Divide the KQ large matrix into several KQ small matrices.
[0058] In this embodiment, block processing can significantly reduce the amount of computation. For example, an n×n (where n > 3) KQ large matrix can be divided into several KQ small matrices, and a total of | | KQ small matrices can be divided.
[0059] S3. Perform local attention calculation on each KQ small matrix, and perform a descending order sorting of similarities with other KQ small matrices to obtain a sorted list.
[0060] In this embodiment, for each KQ small matrix, calculate its corresponding local attention, that is, calculate the similarity between the K and Q vectors in the KQ small matrix to obtain attention scores; then calculate the similarity with other KQ small matrices based on the attention scores, and arrange them in descending order of similarity to obtain a sorted list.
[0061] For example, currently there are 64 3×3 KQ small matrices. After calculating the attention scores of each 3×3 KQ small matrix, the similarity between each KQ small matrix and other KQ small matrices is calculated in sequence. For example, when calculating the 1st KQ small matrix, the similarity with the remaining 63 KQ small matrices is calculated; the similarity between the 2nd KQ small matrix and the other 62 KQ small matrices, and so on, thereby obtaining a sorted list S arranged from largest to smallest similarity.
[0062] S4. For each element in the input sequence, according to the calculated sorted list, select the local attention corresponding to the top r KQ small matrices with the highest similarity for weighted calculation as the attention representation of the current element until all elements of the input sequence are completed, and obtain the optimized attention result of the input sequence.
[0063] For each element in the input sequence, only calculate the product of the corresponding K and Q of the current element. Specifically:
[0064] According to the calculated sorted list S, select the top r (r is less than the number of elements in the input sequence) KQ small matrices with the highest similarity to obtain a high-similarity attention set; perform weighted calculation on the corresponding KQ small matrices in the high-similarity attention set as the attention representation of the current element.
[0065] For example, assume that the top 2 KQ small matrices with the highest similarity are selected. Suppose the top 2 KQ small matrices are ( , ) and ( , ), that is, the high-similarity attention set. Then the attention weight of the current element is ( × + × + × + × );
[0066] Among them, 、 、 、 respectively represent the attention of the a-th, b-th, c-th, and d-th KQ small matrices; ( , ) represents the similarity between the a-th and b-th KQ small matrices; ( , ) represents the similarity between the c-th and d-th KQ small matrices; 、 、 、 respectively represent the set weights.
[0067] After the attention weight vector of the current element is normalized by the softmax function, the value matrix V is weighted and summed to obtain the global attention representation of the current element, and then the attention representation of each element is obtained.
[0068] Before normalization, the result of the dot product operation can be scaled as needed, such as dividing by the square root of the key vector dimension.
[0069] In summary, when calculating the attention of the input sequence, the present invention reduces the computational complexity when calculating the global attention of each element, reduces the operation between attention matrices, and optimizes the computational complexity of the large model by performing block processing of the large matrix into small matrices. At the same time, the present invention selects small matrices with higher similarity for weighting through the similarity sorted list, can perform weighted calculation with distant matrices, thereby capturing distant context information, improving the accuracy of the attention output result, and thus improving the accuracy and performance of the large model.
[0070] Embodiment 2
[0071] As Figure 2 shown, the second embodiment of the present invention also provides a large model self-attention optimization device, including:
[0072] An acquisition unit for acquiring the KQ large matrix obtained after the input sequence passes through the large model self-attention structure;
[0073] A segmentation unit for dividing the KQ large matrix into several KQ small matrices;
[0074] A similarity sorting unit for performing local attention calculation on each KQ small matrix; and performing a descending order sorting of similarities with other KQ small matrices to obtain a sorted list;
[0075] An attention calculation unit for calculating the attention representation of each element; for each element in the input sequence, according to the calculated sorted list, selecting the local attention corresponding to the top r KQ small matrices with the highest similarity for weighted calculation as the attention representation of the current element until all elements of the input sequence are completed to obtain the optimized attention result of the input sequence.
[0076] Embodiment 3
[0077] The third embodiment of the present invention also provides a large model self-attention optimization device, which includes a memory and a processor, and a computer program is stored in the memory, and the computer program can be executed by the processor to implement the large model self-attention optimization method as described above.
[0078] Example 4
[0079] The fourth embodiment of the present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, the above-mentioned large model self-attention optimization method is implemented.
[0080] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are only illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0081] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0082] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs. It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0083] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0084] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0085] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0086] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0087] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for optimizing the self-attention of a large model, characterized in that, Including: S1. Obtain the large KQ matrix obtained after the input sequence passes through the self-attention structure of the large model; wherein, the large KQ matrix is the similarity matrix of the key matrix K and the query matrix Q generated by the self-attention structure of the large model; the large model is used to process input text data, image data or audio-visual data to obtain a processing result; S2. Divide the KQ large matrix into several KQ small matrices; S3. Perform local attention calculation on each small KQ matrix; and perform a descending order sorting of similarities with other small KQ matrices to obtain a sorted list. Specifically: calculate the similarity of each small KQ matrix to obtain an attention score; then calculate the similarity with other small KQ matrices based on the attention score, and arrange them in descending order of similarity to obtain a sorted list; S4. For each element in the input sequence, according to the calculated sorted list, select the local attention corresponding to the top r small KQ matrices with the highest similarity for weighted calculation as the attention representation of the current element until all elements of the input sequence are completed to obtain the optimized attention result of the input sequence; wherein, the attention representation of the current element is specifically obtained through the following steps: For each element in the input sequence, according to the calculated sorted list, select the top r small KQ matrices with the highest similarity to obtain a high-similarity attention set; For each small KQ matrix in the high-similarity attention set, perform weighted calculation of attention according to a preset weight, and then use the softmax function for normalization to obtain the attention weight vector of the current element; Perform weighted summation of the attention weight vector of the current element on the value matrix V generated by the self-attention structure of the large model to obtain the global attention representation of the current element.
2. The self-attention optimization method for large models according to claim 1, wherein , The similarity is obtained by calculating the dot product of elements.
3. The self-attention optimization method for large models according to claim 1, characterized in that , The similarity is obtained by calculating the cosine similarity of elements.
4. An optimization device for self-attention of large models, characterized in that, Including: An acquisition unit, configured to obtain the large KQ matrix obtained after the input sequence passes through the self-attention structure of the large model; wherein, the large KQ matrix is the similarity matrix of the key matrix K and the query matrix Q generated by the self-attention structure of the large model; the large model is used to process input text data, image data or audio-visual data to obtain a processing result; The splitting unit is used to divide the KQ large matrix into several KQ small matrices; A similarity sorting unit, configured to perform local attention calculation on each small KQ matrix; and perform a descending order sorting of similarities with other small KQ matrices to obtain a sorted list. Specifically: calculate the similarity of each small KQ matrix to obtain an attention score; then calculate the similarity with other small KQ matrices based on the attention score, and arrange them in descending order of similarity to obtain a sorted list; An attention calculation unit, configured to calculate the attention representation of each element; for each element in the input sequence, according to the calculated sorted list, select the local attention corresponding to the top r small KQ matrices with the highest similarity for weighted calculation as the attention representation of the current element until all elements of the input sequence are completed to obtain the optimized attention result of the input sequence; wherein, the attention representation of the current element is specifically obtained through the following steps: For each element in the input sequence, according to the calculated sorted list, select the top r KQ small matrices with the highest similarity to obtain a high-similarity attention set; For each KQ small matrix in the high-similarity attention set, perform weighted calculation of attention according to the preset weight, and then use the softmax function for normalization to obtain the attention weight vector of the current element; Perform weighted summation of the value matrix V generated by the self-attention structure of the large model with the attention weight vector of the current element to obtain the global attention representation of the current element.
5. A large model self-attention optimization device, characterized in that, It includes a processor and a memory. The memory stores a computer program, and the computer program can be executed by the processor to implement a large model self-attention optimization method as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, a large model self-attention optimization method as described in any one of claims 1-3 is implemented.
Citation Information
Patent Citations
Weight calculation method and device of attention network, electronic equipment and medium
CN118537577A