Method and device for attention calculation in artificial intelligence model
By determining the effective data location of attention calculation in the AI big model and using blocking and masking processing, the problem of low attention calculation efficiency is solved, and more efficient calculation is achieved.
Patent Information
- Application Number
- CN202410124904.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-07-29
AI Technical Summary
The attention calculation efficiency in existing AI large models is low, mainly because the mask matrix participates in matrix operations, resulting in invalid data calculation, which reduces efficiency.
By determining the effective data location of the Q matrix and K matrix, avoiding invalid data participating in the calculation, and using blocking and masking processing technology to improve the calculation efficiency.
It improves the efficiency of attention calculation, reduces the amount of invalid data calculation, and improves the calculation speed and efficiency.
Smart Images

Figure CN120386969A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method and device for attention calculation in an artificial intelligence model. Background Art
[0002] Transformer belongs to a neural network architecture and is the core structure of current popular large artificial intelligence (AI) models. The attention calculation involved in Transformer can generate text features for the input text of the AI large model, enabling the AI large model to better "understand" the input text so that the AI large model can process the input text more accurately.
[0003] In attention calculation, three matrices Q, K, and V can be generated for the input text first, and then matrix operations are performed on the three matrices Q, K, and V to obtain the text features of the input text.
[0004] However, in current AI large models, a mask matrix is also required to participate in the matrix operations of the three matrices Q, K, and V to filter out invalid data during the operation of the three matrices Q, K, and V, but this also reduces the efficiency of attention calculation. Summary of the Invention
[0005] Embodiments of this application provide a method and device for attention calculation in an artificial intelligence model, which can improve the efficiency of attention calculation. The corresponding technical solutions are as follows:
[0006] In a first aspect, a method for attention calculation in an artificial intelligence model is provided. The method includes: obtaining a plurality of elements that need to perform attention calculation, where there is an order relationship between the plurality of elements, and generating a Q matrix, a K matrix, and a V matrix required for attention calculation corresponding to the plurality of elements. Determining the positions of valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix, where the valid data is used to identify the relationship between each element and the element at a specified position among the plurality of elements. Determining the calculation data for matrix multiplication in the Q matrix and the K matrix according to the determined positions of the valid data in the relationship matrix. Performing matrix multiplication on the determined calculation data of the Q matrix and the K matrix to obtain a relationship matrix. Performing an operation on the relationship matrix and the V matrix to implement attention calculation.
[0007] Among them, the multiple elements that need to perform attention calculation can be multiple characters included in the input text, and the sequential relationship between the multiple elements is the sequential order of each character in the input text. The relationship matrix is the similarity matrix corresponding to the calculation of the Q matrix and the K matrix. The valid data are the elements in the relationship matrix that need to participate in the attention calculation, and the positions of the valid data in the relationship matrix can be set by those skilled in the art.
[0008] In the solution shown in this application, during the process of performing attention calculation, first, when calculating the relationship matrix of the Q matrix and the K matrix, the calculation data for the Q matrix and the K matrix can be determined according to the positions of the valid data in the relationship matrix, and then the valid data in the relationship matrix are calculated based on the calculation data. In this way, during the process of calculating the relationship matrix, invalid data are avoided from participating in the calculation, the calculation efficiency of the relationship matrix is improved, and thus the efficiency of the attention calculation is improved. In addition, since the relationship matrix calculated in this application mainly includes valid data, invalid data can be further avoided from participating in subsequent attention calculations, and the efficiency of the attention calculation can be further improved.
[0009] In an implementable manner, determining the positions of the valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix includes: obtaining the setting information of the invalid data in the artificial intelligence model, where the invalid data are the relationships between multiple elements that do not need to be considered when performing attention calculation. According to the setting information, the positions of the valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix are determined.
[0010] Among them, the setting information of the invalid data in the artificial intelligence model may include the positions of the invalid data in the relationship matrix. In this way, the positions other than the invalid data in the relationship matrix are the positions corresponding to the valid data. In an example, the setting information may be a pre-set mask matrix having the same size as the relationship matrix, and each element value in the mask matrix is used to indicate whether the element at the same position in the relationship matrix is invalid data or valid data. In this way, by determining the positions of the valid data in the relationship matrix, the positions of the calculation data of the Q matrix and the K matrix for calculating the valid data can be deduced through the matrix multiplication operation method. In this way, only by performing the matrix operation of the Q matrix and the K matrix according to the deduced calculation data in the Q matrix and the K matrix, the valid data in the relationship matrix can be calculated, and thus the operation of the invalid data in the relationship matrix can be avoided, and the efficiency of the attention calculation can be improved.
[0011] In an implementable manner, matrix multiplication is performed on the calculated data of the determined Q matrix and K matrix to obtain a relationship matrix, including: partitioning the Q matrix to obtain a first matrix partition, and partitioning the K matrix to obtain a second matrix partition; performing matrix multiplication on the calculated data of the determined Q matrix and K matrix according to the first matrix partition corresponding to the Q matrix and the second matrix partition corresponding to the K matrix to obtain the relationship matrix.
[0012] In the solution shown in the present application, when calculating the valid data in the relationship matrix corresponding to the Q matrix and K matrix, the Q matrix and K matrix can be partitioned to obtain the first matrix partition corresponding to the Q matrix and the second matrix partition corresponding to the K matrix. In this way, by calculating the valid data in the relationship matrix through the first matrix partition and the second matrix partition, the calculation efficiency of the valid data can be improved, and further the efficiency of the attention calculation can be improved.
[0013] In an implementable manner, partitioning the Q matrix to obtain a first matrix partition and partitioning the K matrix to obtain a second matrix partition includes: partitioning the Q matrix and the K matrix respectively according to a specified size to obtain the first matrix partition corresponding to the Q matrix and the second matrix partition corresponding to the K matrix, where the specified size is determined by the size of the Q matrix or the K matrix.
[0014] In the solution shown in the present application, among the respective fifth matrix partitions in the relationship matrix calculated from the first matrix partition and the second matrix partition, there may be some fifth matrix partitions that include invalid data. By determining the partitioning size of the Q matrix and the K matrix according to the size of the Q matrix or the K matrix, the efficiency of calculating the third matrix partition can be ensured, and the number of invalid data included in some of the third matrix partitions can be reduced.
[0015] In an implementable manner, before partitioning the Q matrix and the K matrix respectively according to the specified size, it further includes: obtaining the Q matrix, K matrix, and V matrix generated from multiple sample elements that need to perform attention calculation. Determining a plurality of candidate sizes, and respectively based on each candidate size, partitioning the Q matrix and K matrix generated from the multiple sample elements to obtain the third matrix partition corresponding to the Q matrix generated from the multiple sample elements and the fourth matrix partition corresponding to the K matrix generated from the multiple sample elements, where the candidate size is determined by the size of the Q matrix and the K matrix. Based on the third matrix partition, fourth matrix partition corresponding to each candidate size, and the V matrix generated from the multiple sample elements, sequentially determining the attention calculation results corresponding to the multiple sample elements, and respectively determining the time taken for each obtained attention calculation result. Based on the time taken corresponding to each candidate size, determining the specified size among the multiple candidate sizes.
[0016] In this way, different candidate sizes can be tested through sample elements, and then the specified size with the highest calculation efficiency can be obtained. In practical applications, the Q matrix and the K matrix can be blocked according to the determined specified size, thereby improving the efficiency of calculating the relationship based on the first matrix block and the second matrix block, and improving the efficiency of attention calculation.
[0017] In an implementable manner, the method further includes: determining the fifth matrix block in the same row of the relationship matrix as the sixth matrix block of the relationship matrix, where the fifth matrix block in the relationship matrix is calculated from the first matrix block in the same row of the Q matrix and the second matrix block in the same column of the K matrix. Obtain a Mask matrix corresponding to each sixth matrix block, where each Mask matrix has the same size as the corresponding sixth matrix block. Based on the Mask matrix corresponding to each sixth matrix block, perform masking processing on each sixth matrix block to obtain each masked sixth matrix block.
[0018] In the solution shown in this application, there may be some invalid data in some of the fifth matrix blocks in the relationship matrix calculated from the first matrix block and the second matrix block. Therefore, in this application, the invalid data in the fifth matrix block can be masked row by row to avoid the influence of invalid data on the attention calculation result.
[0019] In an implementable manner, performing an operation on the relationship matrix and the V matrix to implement attention calculation includes: based on each masked sixth matrix block, performing an operation on the relationship matrix and the V matrix to obtain an attention calculation result corresponding to the attention calculation. Since the sixth matrix block mainly includes the valid data in the relationship matrix, performing a matrix multiplication operation on the similarity matrix and the V matrix based on the sixth matrix block can avoid a large amount of invalid data participating in the matrix multiplication operation, improve the efficiency of performing the matrix multiplication operation, and thus improve the efficiency of attention calculation.
[0020] In a second aspect, there is provided an apparatus for attention calculation in an artificial intelligence model, the apparatus includes:
[0021] An acquisition module, configured to acquire a plurality of elements that need to perform attention calculation, and there is an order relationship between the plurality of elements;
[0022] A generation module, configured to generate a Q matrix, a K matrix, and a V matrix required for attention calculation corresponding to the plurality of elements;
[0023] A determination module, configured to determine the positions of valid data in a relationship matrix obtained by performing matrix multiplication on a Q matrix and a K matrix, where the valid data is used to identify the relationship between each element and an element at a specified position among multiple elements;
[0024] A determination module, configured to determine the calculation data for matrix multiplication in the Q matrix and the K matrix according to the positions of the valid data determined in the relationship matrix;
[0025] A calculation module, configured to perform matrix multiplication on the determined calculation data of the Q matrix and the K matrix to obtain a relationship matrix, and perform an operation on the relationship matrix and a V matrix to implement attention calculation.
[0026] In an implementable manner, the determination module is configured to: obtain the setting information of invalid data in the artificial intelligence model, where the invalid data is the relationship between multiple elements that does not need to be considered during attention calculation; determine the positions of the valid data in the relationship matrix obtained by performing matrix multiplication on the Q matrix and the K matrix according to the setting information.
[0027] In an implementable manner, the calculation module is configured to: partition the Q matrix to obtain a first matrix partition, and partition the K matrix to obtain a second matrix partition; perform matrix multiplication on the determined calculation data of the Q matrix and the K matrix according to the first matrix partition corresponding to the Q matrix and the second matrix partition corresponding to the K matrix to obtain a relationship matrix.
[0028] In an implementable manner, the calculation module is configured to: partition the Q matrix and the K matrix respectively according to a specified size to obtain a first matrix partition corresponding to the Q matrix and a second matrix partition corresponding to the K matrix, where the specified size is determined by the size of the Q matrix or the K matrix.
[0029] In an implementable manner, the apparatus further includes a test module, configured to: obtain a Q matrix, a K matrix, and a V matrix generated by multiple sample elements that need to perform attention calculation; determine multiple candidate sizes, and partition the Q matrix and the K matrix generated by the multiple sample elements respectively based on each candidate size to obtain a third matrix partition corresponding to the Q matrix generated by the multiple sample elements and a fourth matrix partition corresponding to the K matrix generated by the multiple sample elements, where the candidate sizes are determined by the sizes of the Q matrix and the K matrix; determine the attention calculation results corresponding to the multiple sample elements in sequence based on the third matrix partition, the fourth matrix partition, and the V matrix generated by the multiple sample elements corresponding to each candidate size, and determine the time taken for each obtained attention calculation result respectively; determine a specified size among the multiple candidate sizes based on the time taken corresponding to each candidate size.
[0030] In an implementable manner, the device further includes a masking processing module, configured to: determine the fifth matrix blocks in the same row of the relationship matrix as the sixth matrix blocks of the relationship matrix, where the fifth matrix blocks in the relationship matrix are calculated from the first matrix blocks in the same row of the Q matrix and the second matrix blocks in the same column of the K matrix; obtain a Mask matrix corresponding to each sixth matrix block, where each Mask matrix has the same size as the corresponding sixth matrix block; and perform masking processing on each sixth matrix block based on the Mask matrix corresponding to each sixth matrix block to obtain each masked sixth matrix block.
[0031] In an implementable manner, the calculation module is configured to: perform an operation between the relationship matrix and the V matrix based on each masked sixth matrix block to obtain an attention calculation result corresponding to the attention calculation.
[0032] In a third aspect, a computing device is provided, which includes a processor and a memory. The processor is configured to execute instructions stored in the memory, so that the computing device executes the method described in the first aspect and / or any implementable manner in the first aspect as above.
[0033] In a fourth aspect, a computer program product including instructions is provided. When the instructions are run on a computing device, the computing device is caused to execute the method described in the first aspect and / or any implementable manner in the first aspect as above.
[0034] In a fifth aspect, a computer-readable storage medium is provided, including computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method described in the first aspect and / or any implementable manner in the first aspect as above. Description of the Drawings
[0035] Figure 1 is a schematic structural diagram of a Transformer provided by an embodiment of the present application;
[0036] Figure 2 is a flowchart of the attention calculation involved in the Transformer;
[0037] Figure 3 is a schematic structural diagram of a computing device provided by an embodiment of the present application;
[0038] Figure 4 is a flowchart of a method for attention calculation in an artificial intelligence model provided by an embodiment of the present application;
[0039] Figure 5It is a schematic diagram of valid data in a relationship matrix provided by an embodiment of the present application;
[0040] Figure 6 It is another schematic diagram of valid data in a relationship matrix provided by an embodiment of the present application;
[0041] Figure 7 It is still another schematic diagram of valid data in a relationship matrix provided by an embodiment of the present application;
[0042] Figure 8 It is a schematic diagram of a matrix multiplication operation provided by an embodiment of the present application;
[0043] Figure 9 It is a schematic diagram of invalid data provided by an embodiment of the present application;
[0044] Figure 10 It is a schematic diagram of a sixth matrix block and a mask matrix provided by an embodiment of the present application;
[0045] Figure 11 It is a flowchart of a method for determining a specified size provided by an embodiment of the present application;
[0046] Figure 12 It is a structural diagram of a device for attention calculation in an artificial intelligence model provided by an embodiment of the present application. Detailed implementation manners
[0047] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0048] Attention calculation is a widely used calculation method in the field of artificial intelligence technology, which can be used to calculate the correlation between multiple elements and can help an AI model "understand" multiple input elements. Generally, the multiple elements that require attention calculation can be each word included in a sentence or each pixel point included in an image.
[0049] Transformer is a neural network architecture, which is the core structure of current popular large AI models and involves attention calculation. Here, taking Transformer as an example, the calculation process of attention calculation will be introduced. Figure 1 It is a schematic diagram of the structure of Transformer. In Figure 1Among them, Input Embedding is the input vector converted from the input data of the large AI model. LayerNorm is Layer normalization, which is used to normalize the input vector. Q, K, and V are the Q matrix, K matrix, and V matrix generated based on the normalized input vector respectively. Among them, Q represents Query, K represents Key, and V represents Value. mask is the mask matrix, "+" represents performing an addition operation on the matrix, and "×" represents performing a multiplication operation on the matrix. The Softmax layer is used to normalize the matrix.
[0050] Figure 2 is the flowchart of the attention calculation involved in the Transformer. See Figure 2 The process of this attention calculation includes:
[0051] S1. After obtaining the Q matrix, K matrix, and V matrix, the multiplication operation can be performed on the Q matrix and the transposed matrix of the K matrix (K.T matrix) to obtain the similarity matrix (Sim matrix) of the Q matrix and the K matrix, which can also be called the relationship matrix.
[0052] S2. After obtaining the similarity matrix Sim, the Sim matrix can be masked by the mask matrix, that is, the matrix addition operation is performed on the mask matrix and the Sim matrix, and then the masked matrix (Sim-mask matrix) is obtained.
[0053] In an example, the values in the lower left triangular region of the mask matrix are all 0, and the values in the upper right triangular region are all negative infinity. Therefore, after performing the matrix addition of the mask matrix and the Sim matrix, the values in the lower left triangular region of the sim_masked matrix are the same as those in the lower left triangular region of the Sim, and the values in the upper right triangular region will become negative infinity. Among them, during the calculation process, negative infinity can be set to a relatively large negative number, such as -10000.
[0054] S3. After obtaining the Sim-mask matrix, the probability matrix is calculated using the softmax calculation formula.
[0055] In step S3, the maximum value of each row of the Sim-mask matrix can be determined row by row, and then the value of each element in each row is subtracted by the maximum value of that row. After the subtraction, the element values in the lower left triangular region of the Sim_masked matrix will be less than or equal to 0, and the upper right triangular region will still be negative infinity. Then, an exponential calculation is performed on each element. According to the characteristics of the exponential calculation, the element values in the lower left triangular region of the Sim_masked matrix will be between 0 and 1, and the values in the upper right triangular region will be all 0. Then, the sum of the elements in each row of the Sim_masked matrix can be calculated, and each element in each row is divided by the sum of the corresponding row to obtain the probability matrix Probs.
[0056] Since the values in the upper right triangular region of the Sim_masked matrix are all 0 after the exponential operation, when calculating the sum of each row, this part is meaningless for the sum of each row, and the value after dividing this part by the sum of the corresponding row is still 0, that is, the values in the upper right triangular region of the probability matrix Probs are all 0.
[0057] S4. After obtaining the probability matrix Probs, the matrix multiplication operation can be performed on the probability matrix Probs and the V matrix to obtain the output matrix Out.
[0058] Since the values in the upper right triangular region of the probability matrix Probs are all 0, the elements in this part have no influence on the result of the matrix multiplication operation. Therefore, it is also meaningless to perform the matrix multiplication operation on the elements in this part.
[0059] As can be seen from the above steps S1 to S4, there are a large number of meaningless calculations in the current attention calculation, so the efficiency of the current attention calculation needs to be further improved.
[0060] The embodiment of the present application provides a method for attention calculation in an artificial intelligence model, which is applicable to attention calculations involved in various scenarios, such as language processing, image processing and other scenarios, can reduce the calculation of invalid data in the attention calculation, and can improve the efficiency of the attention calculation. Figure 3 It is a schematic structural diagram of a computing device for a method for attention calculation in an artificial intelligence model provided by the embodiment of the present application. As Figure 3As shown, the computing device 300 may include: a bus 302, a processor 304, a memory 306. Optionally, the computing device 300 may further include a communication interface 308. The processor 304, the memory 306, and the communication interface 308 communicate with each other via the bus 302. The computing device 300 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 300. The computing device 300 may be a device for running a model, which may be a terminal or a server. When the computing device 300 is a terminal, the computing device 300 includes, but is not limited to, a desktop computer, a mobile phone, a notebook, a tablet computer, etc. When the computing device 300 is a server, the computing device 300 may be a single server, that is, a server that can independently perform AI model training, or any device in a computing cluster that performs AI model training. If the computing device 300 is any device in the cluster, it is used to perform the attention calculation on a part of the data in the AI model. Multiple virtual machines or containers may run on the computing device, and each virtual machine or container may also independently perform the attention calculation.
[0061] The bus 302 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 302 may include a path for transmitting information between various components of the computing device 300 (for example, the memory 306, the processor 304, the communication interface 308).
[0062] The processor 304 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc. Or the processor may be a system-on-chip (SOC) including one or more of the above-mentioned CPU, GPU, MP, etc. Among them, the processor 304 may further include a matrix operation unit, and the matrix operation unit may be involved in matrix operations such as matrix multiplication or matrix addition in the method for implementing the attention calculation in the artificial intelligence model provided in the embodiments of the present application.
[0063] The memory 306 may include a volatile memory, such as a random access memory (RAM). The memory 306 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0064] The executable program code is stored in the memory 306, and the processor 304 executes the executable program code to implement the method for attention calculation in the artificial intelligence model provided by the embodiments of the present application. For example, obtaining a plurality of elements that need to perform attention calculation, there is an order relationship between the plurality of elements, and generating a Q matrix, a K matrix, and a V matrix required for attention calculation corresponding to the plurality of elements. Determining the position of the valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix, where the valid data is used to identify the relationship between each element and the element at a specified position among the plurality of elements. According to the determined position of the valid data in the relationship matrix, determining the calculation data for performing matrix multiplication in the Q matrix and the K matrix. Performing matrix multiplication on the determined Q matrix and K matrix calculation data to obtain a relationship matrix. Performing an operation on the relationship matrix and the V matrix to implement attention calculation, etc.
[0065] The communication interface 308 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 300 and other devices or a communication network.
[0066] Figure 4 It is a flowchart of a method for attention calculation in an artificial intelligence model provided by the embodiments of the present application. This method can be executed by the above-mentioned computing device 300, and further can be executed by the processor 304 included in the computing device 300. Refer to Figure 4 and the method includes:
[0067] Step 401, the processor obtains a plurality of elements that need to perform attention calculation.
[0068] Among them, there is an order relationship between the plurality of elements, and the plurality of elements can be any elements suitable for performing attention calculation.
[0069] In one example, the method for calculating attention provided by the embodiments of the present application is applied to an AI language large model. The multiple elements for which attention calculation is required are the input text of the AI language large model. For example, it is the dialogue statement input by the user to the AI large model, or the input text can also refer to the dialogue text generated by the large model. Each word in the input text is an element, and the order of each word in the input text is the order relationship existing between the multiple elements.
[0070] In another example, the method for calculating attention provided by the embodiments of the present application is applied to an AI image model. The multiple elements for which attention calculation is required are the input images of the AI image model. For example, it is the image to be subjected to semantic segmentation. Each pixel in the input image is an element, and the positional relationship of the multiple pixels in the input image is the order relationship existing between the multiple elements.
[0071] Step 402: Generate Q matrix, K matrix, and V matrix required for attention calculation corresponding to multiple elements.
[0072] As Figure 2 shown, in attention calculation, for the input text, a Q matrix, a K matrix, and a V matrix can be generated. In one example, after the input text undergoes LayerNorm processing to obtain a normalized input vector, the normalized input vector can be respectively subjected to matrix multiplication operations with weight matrices W Q 、W K 、W V to obtain the Q matrix, the K matrix, and the V matrix.
[0073] Step 403: Determine the positions of the valid data in the relationship matrix obtained after the matrix multiplication of the Q matrix and the K matrix. The valid data is used to identify the relationship between each element and the element at a specified position among the multiple elements.
[0074] Among them, the valid data is the data in the relationship matrix that needs to participate in the attention calculation, that is, the relationship between different elements that need to participate in the attention calculation. The specified position is the position of other elements that calculate the relationship with each element. This specified position can be set by those skilled in the art according to the application scenario of the AI model. After the specified position is determined, the position of the corresponding valid data in the relationship matrix is also determined accordingly. Figures 5 to 7 Three specified positions and the positions of the corresponding valid data in the relationship matrix are respectively shown.
[0075] Figure 5 is a schematic diagram of the valid data in a relationship matrix provided by the embodiments of the present application. In Figure 5The elements in the lower left triangular region of the shown relationship matrix are the valid data in the relationship matrix. In one example, the attention calculation does not need to "pay attention" to the complete context, and in the generative model, for the input text, it has the characteristic of "only looking forward, not looking backward". Therefore, in the attention calculation, only the relevance between each word in the input text and itself and the subsequent words needs to be concerned. That is to say, the specified position corresponding to each of the above elements is the positions before the position of this element and the position of this element itself among multiple elements. In the attention calculation, only the data in the lower left triangular region of the relationship matrix needs to participate in the operation to obtain the text features.
[0076] Figure 6 is a schematic diagram of the valid data in another relationship matrix provided by an embodiment of the present application. In Figure 6 the elements located in the diagonal region of the shown relationship matrix are the valid data in the relationship matrix. In one example, when the length of the input text is too long, for example, the input text exceeds the set length threshold, or the input text is an article. In this case, in the attention calculation, only the relevance between each word and its nearby words needs to be concerned. For example, the specified position corresponding to each of the above elements is the n positions before the position of this element and the position of this element itself among multiple elements. Wherein, n can be a preset value. When the number of positions before the position of this element is less than n, the specified position is the positions before this element and the position of this element itself. In the attention calculation, only the relevance between the first n words before each word needs to be concerned.
[0077] Figure 7 is a schematic diagram of the valid data in yet another relationship matrix provided by an embodiment of the present application. In Figure 7 the elements located in the lower left triangular region and the upper left matrix region of the shown relationship matrix are the valid data in the relationship matrix. In one example, the input text includes a dialogue statement input by the user (such as a question) and the dialogue text generated by the AI large model. In this case, in the attention calculation, for the dialogue statement input by the user, the relevance between each word in the dialogue statement can be concerned, and for the dialogue text generated by the AI large model, it can "only look forward, not look backward", that is, only the relevance between each word in the dialogue text and itself and the subsequent words needs to be concerned. That is to say, the specified position corresponding to each of the above elements can include both the positions before the position of this element among multiple elements and the n positions before the position of this element and the position of this element itself among multiple elements. In the attention calculation, only the data located in the lower left triangular region and the upper left matrix region of the relationship matrix needs to participate in the operation to obtain the text features.
[0078] In one example, invalid data configuration information can be obtained from the AI model. Based on this configuration information, the position of valid data in the relationship matrix obtained by performing matrix multiplication on the Q and K matrices is determined. Invalid data refers to data outside the valid data in the relationship matrix, specifically, relationships between multiple elements that are not considered during attention calculations.
[0079] The setting information of invalid data in the artificial intelligence model may include the position of the invalid data in the relationship matrix, so that the position in the relationship matrix other than the invalid data is the position corresponding to the valid data. The setting information can be set according to the above-mentioned specified position. The setting information can be a pre-set mask matrix with the same size as the relationship matrix, and each element value in the mask matrix is used to indicate whether the element at the same position in the relationship matrix is invalid data or valid data. For example, each element in the mask matrix consists of "0" and "-10000", where the element "0" is used to indicate that the element at the same position in the relationship matrix is valid data, and the element "-10000" is used to indicate that the element at the same position in the relationship matrix is invalid data. Therefore, the position of valid data and invalid data in the relationship matrix can be determined through the mask matrix.
[0080] Step 404: Determine the calculation data for matrix multiplication calculation in the Q matrix and the K matrix according to the position of the determined valid data in the relationship matrix.
[0081] After determining the position of the valid data in the relational matrix, the Q matrix and K matrix can be inferred by the matrix multiplication method. T The position of the calculation data used to calculate the effective data in the matrix (the device matrix of the K matrix). For example, if the effective data is in the i-th row and j-th column of the relational matrix, the calculation data for calculating the effective data is the element in the i-th row of the Q matrix and the K T The element in the j-th column of the matrix.
[0082] In this way, it is only necessary to perform matrix operations on the Q matrix and the K matrix based on the calculated data in the inferred Q matrix and the K matrix to calculate the valid data in the relationship matrix, thereby avoiding the operation of invalid data in the relationship matrix and improving the efficiency of attention calculation.
[0083] Step 405: Perform matrix multiplication on the determined Q matrix and K matrix calculation data to obtain a relationship matrix.
[0084] In implementation, a vector operation unit included in the processor may be used to perform a matrix multiplication operation on the Q matrix and the K matrix calculation data to obtain each valid data included in the relational matrix.
[0085] In an implementable manner, to improve the efficiency of calculating the relationship matrix. After obtaining the Q matrix and the K T matrix, the Q matrix and the K T matrix can be first subjected to Tiling (blocking) processing to obtain multiple Tiling blocks included in the Q matrix (which can be referred to as the first matrix blocks) and the K T matrix included in multiple Tiling blocks (which can be referred to as the second matrix blocks). Then, based on the first matrix blocks and the second matrix blocks, the valid data in the relationship matrix corresponding to the Q matrix and the K matrix can be calculated.
[0086] Correspondingly, the processing of step 404 above can be replaced by determining the first matrix block and the second matrix block for matrix multiplication calculation in the Q matrix and the K T matrix according to the position of the determined valid data in the relationship matrix.
[0087] Figure 8 is a schematic diagram of a matrix multiplication operation provided by an embodiment of the present application. In Figure 8 the elements in the lower left triangular region of the relationship matrix are valid data. Therefore, the third matrix block including valid data in the relationship matrix can be calculated according to each row first matrix block in the Q matrix and each column second matrix block in the K T matrix. For example, according to the first matrix block in the i-th row of the Q matrix and the second matrix block in the j-th column of the K T matrix, the third matrix block in the i-th row and j-th column of the relationship matrix is calculated. Wherein, i is greater than or equal to j.
[0088] In implementation, the first matrix block in each row of the Q matrix and the second matrix block in each column of the transposed matrix of the K matrix (the K T matrix) can be calculated in sequence according to the set calculation order, and the fifth matrix block included in the relationship matrix is obtained in sequence.
[0089] In an embodiment of the present application, in the process of performing matrix multiplication operation on the Q matrix and the K T matrix to obtain the corresponding relationship matrix, it is not necessary to calculate all elements in the relationship matrix. Therefore, the calculation amount of matrix multiplication operation in attention calculation is reduced, and the calculation efficiency of attention calculation can be improved.
[0090] Step 406: Perform an operation on the relationship matrix and the V matrix to implement attention calculation.
[0091] Since the relationship matrix lacks invalid data, the matrix multiplication operation of the relationship matrix and the V matrix can be performed only through the invalid data of the relationship matrix. In this way, the calculation of invalid data can be omitted during the calculation process, and thus the efficiency of attention calculation can be improved.
[0092] When the relationship matrix is calculated from the first matrix block of the Q matrix and the second matrix block of the K T matrix, there are partial invalid data in the fifth matrix block of the relationship matrix. For example, Figure 8 each third matrix block located on the diagonal of the relationship matrix includes invalid data. As Figure 9 shown, when the size of the third matrix block is 3×3, the three elements in the upper right of the third matrix block are invalid data. Therefore, in the embodiments of the present application, in order to avoid the influence of this part of invalid data on the attention calculation, a masking processing method is also provided, including:
[0093] Determine the fifth matrix blocks in the same row of the relationship matrix as the sixth matrix blocks. Based on the Mask matrix corresponding to each sixth matrix block, perform masking processing on each sixth matrix block to obtain each sixth matrix block after masking processing.
[0094] In implementation, after calculating the fifth matrix block in the relationship matrix through the first matrix block in the Q matrix and the second matrix block in the K T matrix. Each of the fifth matrix blocks corresponding to the same Tiling row in the relationship matrix can be determined as the sixth matrix block in the relationship matrix. After determining each sixth matrix block, according to the position of each sixth matrix block in the relationship matrix, obtain the mask matrix corresponding to each sixth matrix block.
[0095] Figure 10 FIG. is a schematic diagram of a sixth matrix block and a mask matrix provided by an embodiment of the present application. In an example, the size of each mask matrix is the same as the size of the corresponding sixth matrix block, and each mask matrix includes elements with a value of 0 and elements with a value of negative infinity. Among them, the position of the element with a value of 0 in the mask matrix is the same as the position of the valid data in the corresponding sixth matrix block, and the position of the element with a value of negative infinity in the mask matrix is the same as the position of the invalid data in the corresponding sixth matrix block.
[0096] After the mask matrix corresponding to each sixth matrix block, a matrix addition operation can be performed on each sixth matrix block and the corresponding mask matrix, thereby implementing the masking processing of each fourth matrix block to obtain the sixth matrix block after masking processing. After obtaining the sixth matrix block after masking processing, based on each sixth matrix block after masking processing and the V matrix, perform the operation of the relationship matrix and the V matrix to obtain the attention calculation result corresponding to the attention calculation.
[0097] In implementation, after obtaining the sixth matrix block after mask processing, the elements in each row of each sixth matrix block can be normalized first. That is, determine the maximum value in each row of the sixth matrix block, and then subtract the maximum value of each row from the values of the elements in that row. After that, perform an exponential calculation on each element, then sum the elements in each row, and divide each element in a row by the sum of the corresponding row to obtain the valid data in the probability matrix.
[0098] Among them, after determining the number of multiple elements that need to perform attention calculation, the sizes of Q, K, and V can also be determined, and the specified sizes for Tiling the Q matrix and the K T matrix can also be determined, that is, the sizes of the first matrix block, the second matrix block, and the fifth matrix block can be determined. In this way, the size of each sixth matrix block, the positions of the invalid data in each sixth matrix block can also be determined, and further the mask matrix corresponding to each sixth matrix block can be determined. Therefore, for each sixth matrix block according to its position in the relationship matrix, the corresponding mask matrix is stored in advance, and then after calculating the sixth matrix block, the sixth matrix block can be masked according to the pre-stored mask matrix.
[0099] After obtaining the valid data in the probability matrix, the matrix multiplication operation between the probability matrix and the V matrix can be implemented according to the valid data included in the probability matrix, and then the result of the matrix multiplication operation can be obtained. Then, the result of the matrix multiplication operation can be normalized to obtain the attention calculation result. For example, the attention calculation result can be the text feature corresponding to the input text.
[0100] In the embodiment of the present application, since the sixth matrix block in the relationship matrix mainly includes valid data, when calculating the corresponding probability matrix according to the sixth matrix block, a large amount of invalid data can be avoided from participating in the calculation, and thus the efficiency of generating the probability matrix can be improved. Also, since the probability matrix is mainly calculated through the valid data in the relationship matrix, the calculated probability matrix only includes valid data. Therefore, when performing the matrix multiplication operation between the probability matrix and the V matrix with the valid data in the probability matrix, the participation of invalid data can be avoided, and thus the efficiency of performing the matrix multiplication operation can be improved.
[0101] It can be seen that in the embodiment of the present application, the efficiency of performing the matrix multiplication operation of the Q matrix and the K matrix, the efficiency of the masking process corresponding to the relationship matrix, the efficiency of calculating the probability matrix, and the efficiency of performing the matrix multiplication operation between the probability matrix and the V matrix in the attention calculation can be improved. Therefore, the embodiment of the present application can improve the efficiency of generating the text feature corresponding to the input text.
[0102] In an implementable manner, the specified dimensions for partitioning the Q matrix, K matrix, and V matrix can be determined by the dimensions of the Q matrix or K matrix, that is, by the length of the input text. In one example, the correspondence between the Q matrix or K matrix and the specified dimensions can be pre-stored. During implementation, the specified dimensions for tiling the Q matrix and K matrix can be determined according to this correspondence.
[0103] Figure 11 It is a flowchart of a method for determining the specified dimensions for tiling the Q matrix and K matrix provided by an embodiment of the present application. In one example, this method can be applied in the training phase of an AI large model. Refer to Figure 11 and the method includes:
[0104] Step 1101, obtain the Q matrix, K matrix, and V matrix generated from multiple sample elements that require attention calculation.
[0105] Among them, the multiple sample elements can be the sample texts for training the AI large model. For generating the Q matrix, K matrix, and V matrix corresponding to the sample texts, the processing in step 402 above can be referred to and will not be elaborated here.
[0106] Step 1102, determine multiple candidate dimensions, and respectively partition the Q matrix and K matrix generated from the multiple sample elements based on each candidate dimension to obtain the third matrix partitions corresponding to the Q matrix generated from the multiple sample elements and the fourth matrix partitions corresponding to the K matrix generated from the multiple sample elements. The candidate dimensions are determined by the dimensions of the Q matrix and K matrix.
[0107] Among them, the candidate dimensions can be preset by those skilled in the art according to the dimensions of the Q matrix and K matrix, such as including 2×2, 4×4, 8×8, etc. After determining the candidate dimensions, the transposed matrices (K T matrix) corresponding to the Q matrix and K matrix can be respectively partitioned to obtain each third matrix partition included in the Q matrix and each fourth matrix partition included in the K T matrix. In one example, when the specified dimension is m×n, where m is not equal to n, the Q matrix can be partitioned according to the dimension of m×n, and the K T matrix can be partitioned according to the dimension of n×m.
[0108] Step 1103, based on the third matrix partitions, fourth matrix partitions corresponding to each candidate dimension and the V matrix generated from the multiple sample elements, sequentially determine the attention calculation results corresponding to the multiple sample elements, and respectively determine the time consumed for obtaining the attention calculation results each time.
[0109] Step 1104: Determine a specified size from multiple candidate sizes based on the time consumption corresponding to each candidate size.
[0110] In implementation, during the training of the AI model, the Q matrix and the K matrix can be blocked in sequence according to each candidate size to obtain a third matrix block and a fourth matrix block, and text features corresponding to the test text are generated. Among them, the process of performing attention calculation based on the third matrix block and the fourth matrix block can refer to the process of performing attention calculation based on the first matrix block and the second matrix block in the above steps 404-405, which will not be elaborated here.
[0111] For the third matrix block and the fourth matrix block of each candidate size, the time consumption of generating text features by the third matrix block and the fourth matrix block of each candidate size can be recorded, and then the candidate size with the lowest corresponding time consumption is determined as the specified size in the above embodiments. In this way, it can be avoided that the specified size is too small, resulting in too many matrix blocks and thus a decrease in the efficiency of attention calculation, and it can also be avoided that the specified size is too large, resulting in a large amount of invalid data in the matrix blocks and thus a decrease in the efficiency of attention calculation.
[0112] Based on the same inventive concept, an embodiment of the present application further provides a device for attention calculation in an artificial intelligence model. The device can be the computing device that executes the method for attention calculation in the artificial intelligence model, or a program running in the computing device for executing the method for attention calculation in the artificial intelligence model. Figure 12 It is a schematic structural diagram of a device for attention calculation in an artificial intelligence model provided by an embodiment of the present application. Refer to Figure 12 The device includes:
[0113] An acquisition module 1210, configured to acquire multiple elements that need to perform attention calculation. There is an order relationship among the multiple elements, and it is specifically used to implement the acquisition function of the above step 401 and its hidden steps.
[0114] A generation module 1220, configured to generate a Q matrix, a K matrix, and a V matrix required for attention calculation corresponding to the multiple elements, and is specifically used to implement the generation function of the above step 402 and its hidden steps.
[0115] A determination module 1230 is configured to determine the positions of valid data in the relationship matrix obtained by performing matrix multiplication on the Q matrix and the K matrix. The valid data is used to identify the relationship between each element and the element at a specified position among multiple elements. According to the determined positions of the valid data in the relationship matrix, the calculation data for matrix multiplication in the Q matrix and the K matrix is determined, which is specifically used to implement the determination functions of steps 403-404 and their hidden steps as described above.
[0116] A calculation module 1240 is configured to perform matrix multiplication on the determined calculation data of the Q matrix and the K matrix to obtain a relationship matrix, and perform an operation on the relationship matrix and the V matrix to implement attention calculation, which is specifically used to implement the calculation functions of step 405 and its hidden steps as described above.
[0117] In an implementable manner, the determination module 1230 is configured to: obtain the setting information of invalid data in the artificial intelligence model, where the invalid data is the relationship between multiple elements that does not need to be considered during attention calculation; according to the setting information, determine the positions of the valid data in the relationship matrix obtained by performing matrix multiplication on the Q matrix and the K matrix.
[0118] In an implementable manner, the calculation module 1240 is configured to: partition the Q matrix to obtain a first matrix partition, and partition the K matrix to obtain a second matrix partition; according to the first matrix partition corresponding to the Q matrix and the second matrix partition corresponding to the K matrix, perform matrix multiplication on the determined calculation data of the Q matrix and the K matrix to obtain a relationship matrix.
[0119] In an implementable manner, the calculation module 1240 is configured to: partition the Q matrix and the K matrix respectively according to a specified size to obtain a first matrix partition corresponding to the Q matrix and a second matrix partition corresponding to the K matrix, where the specified size is determined by the size of the Q matrix or the K matrix.
[0120] In an implementable manner, the device further includes a test module, configured to: obtain a Q matrix, a K matrix, and a V matrix generated by a plurality of sample elements that require attention calculation; determine a plurality of candidate sizes, and respectively based on each candidate size, partition the Q matrix and the K matrix generated by the plurality of sample elements to obtain a third matrix partition corresponding to the Q matrix generated by the plurality of sample elements and a fourth matrix partition corresponding to the K matrix generated by the plurality of sample elements, where the candidate sizes are determined by the sizes of the Q matrix and the K matrix; based on the third matrix partition, the fourth matrix partition corresponding to each candidate size, and the V matrix generated by the plurality of sample elements, sequentially determine the attention calculation results corresponding to the plurality of sample elements, and respectively determine the time taken for obtaining the attention calculation results each time; based on the time taken corresponding to each candidate size, determine a specified size among the plurality of candidate sizes.
[0121] In an implementable manner, the device further includes a mask processing module, configured to: determine a sixth matrix partition of the relationship matrix as the fifth matrix partitions in the same row of the relationship matrix, where the fifth matrix partitions in the relationship matrix are calculated by a first matrix partition in the same row of the Q matrix and a second matrix partition in the same column of the K matrix; obtain a Mask matrix corresponding to each sixth matrix partition, where each Mask matrix has the same size as the corresponding sixth matrix partition; based on the Mask matrix corresponding to each sixth matrix partition, perform mask processing on each sixth matrix partition to obtain each sixth matrix partition after mask processing.
[0122] In an implementable manner, the calculation module 1240 is configured to: based on each sixth matrix partition after mask processing, perform an operation between the relationship matrix and the V matrix to obtain an attention calculation result corresponding to the attention calculation.
[0123] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional module may be integrated in a processor, may exist separately physically, or two or more modules may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. In addition, the device for attention calculation in the artificial intelligence model provided in the above embodiments and the method embodiments for attention calculation in the artificial intelligence model belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0124] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a terminal device (which may be a personal computer, a mobile phone, or a network device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0125] An embodiment of this application also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the method for attention calculation in the artificial intelligence model provided by the embodiments of this application.
[0126] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method for attention calculation in the artificial intelligence model provided by the embodiments of this application.
[0127] In this application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same function and role. It should be understood that there is no logical or temporal dependency between "first" and "second", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first matrix block can be referred to as the second matrix block, and similarly, the second matrix block can be referred to as the first matrix block. The first matrix block and the second matrix block can both be collectively referred to as matrix blocks, and in some cases, they can be separate and different matrix blocks.
[0128] In this application, the term "at least one" means one or more, and the term "a plurality of" means two or more.
[0129] The above description is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A method for calculating attention in an artificial intelligence model, characterized in that The method includes: Obtaining a plurality of elements for which attention calculation is to be performed, and there is an order relationship among the plurality of elements; Generating a Q matrix, a K matrix, and a V matrix required for attention calculation corresponding to the plurality of elements; Determining the positions of valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix, where the valid data is used to identify the relationship between each element and the element at a specified position among the plurality of elements; Determining the calculation data for matrix multiplication in the Q matrix and the K matrix according to the determined positions of the valid data in the relationship matrix; Performing matrix multiplication on the determined calculation data of the Q matrix and the K matrix to obtain the relationship matrix; Performing an operation on the relationship matrix and the V matrix to implement the attention calculation.
2. The method according to claim 1, wherein The determining the positions of valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix includes: Obtaining the setting information of invalid data in the artificial intelligence model, where the invalid data is the relationship between the plurality of elements that does not need to be considered when performing the attention calculation; Determining the positions of valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix according to the setting information, where the element at the specified position is determined by the setting information.
3. The method according to claim 1 or 2, characterized in that, The performing matrix multiplication on the determined calculation data of the Q matrix and the K matrix to obtain the relationship matrix includes: Partitioning the Q matrix to obtain a first matrix partition, and partitioning the K matrix to obtain a second matrix partition; Performing matrix multiplication on the determined calculation data of the Q matrix and the K matrix according to the first matrix partition corresponding to the Q matrix and the second matrix partition corresponding to the K matrix to obtain the relationship matrix.
4. The method according to claim 3, characterized in that, The partitioning the Q matrix to obtain a first matrix partition and partitioning the K matrix to obtain a second matrix partition includes: Partitioning the Q matrix and the K matrix respectively according to a specified size to obtain the first matrix partition corresponding to the Q matrix and the second matrix partition corresponding to the K matrix, where the specified size is determined by the size of the Q matrix or the K matrix.
5. The method according to claim 4, wherein Before the partitioning the Q matrix and the K matrix respectively according to the specified size, it further includes: Obtaining a Q matrix, a K matrix, and a V matrix generated from a plurality of sample elements for which attention calculation is to be performed; Determining a plurality of candidate sizes, and respectively partitioning the Q matrix and the K matrix generated from the plurality of sample elements based on each candidate size to obtain a third matrix partition corresponding to the Q matrix generated from the plurality of sample elements and a fourth matrix partition corresponding to the K matrix generated from the plurality of sample elements, where the candidate sizes are determined by the sizes of the Q matrix and the K matrix; Based on the third matrix block corresponding to each candidate size, the fourth matrix block, and the multiple sample elements, generate a V matrix, and sequentially determine the attention calculation results corresponding to the multiple sample elements, and respectively determine the time consumed for obtaining each attention calculation result; Based on the time consumed corresponding to each candidate size, determine a specified size among the multiple candidate sizes.
6. The method according to any one of claims 3 to 5, characterized in that, The method further includes: Determine the fifth matrix block in the same row of the relationship matrix as the sixth matrix block of the relationship matrix. The fifth matrix block in the relationship matrix is obtained by calculating the first matrix block in the same row of the Q matrix and the second matrix block in the same column of the K matrix; Obtain a Mask matrix corresponding to each sixth matrix block, where each Mask matrix has the same size as the corresponding sixth matrix block; Based on the Mask matrix corresponding to each sixth matrix block, perform masking processing on each sixth matrix block to obtain each masked sixth matrix block.
7. The method according to claim 6, wherein The performing an operation on the relationship matrix and the V matrix to implement the attention calculation includes: Based on each masked sixth matrix block, perform the operation on the relationship matrix and the V matrix to obtain the attention calculation result corresponding to the attention calculation.
8. An apparatus for calculating attention in an artificial intelligence model, characterized in that, The apparatus includes: An acquisition module, configured to acquire multiple elements that need to perform attention calculation, and there is an order relationship among the multiple elements; A generation module, configured to generate a Q matrix, a K matrix, and a V matrix required for the attention calculation corresponding to the multiple elements; A determination module, configured to determine the position of valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix in the relationship matrix. The valid data is used to identify the relationship between each element and the element at a specified position among the multiple elements; according to the determined position of the valid data in the relationship matrix, determine the calculation data for performing matrix multiplication in the Q matrix and the K matrix; A calculation module, configured to perform matrix multiplication on the determined calculation data of the Q matrix and the K matrix to obtain the relationship matrix, and perform an operation on the relationship matrix and the V matrix to implement the attention calculation.
9. The device according to claim 8, wherein, The determination module is configured to: Obtain the setting information of invalid data in the artificial intelligence model. The invalid data is the relationship between the multiple elements that does not need to be considered when performing the attention calculation; According to the setting information, determine the position of valid data in the relationship matrix obtained after performing matrix multiplication on the Q matrix and the K matrix, where the element at the specified position is determined by the setting information.
10. The device according to claim 8 or 9, characterized in that, The calculation module is configured to: Perform block division on the Q matrix to obtain a first matrix block, and perform block division on the K matrix to obtain a second matrix block; Perform matrix multiplication on the calculated data of the determined Q matrix and K matrix according to the first matrix block corresponding to the Q matrix and the second matrix block corresponding to the K matrix to obtain the relationship matrix.
11. The device according to claim 10, characterized in that, The calculation module is configured to: Block the Q matrix and the K matrix respectively according to a specified size to obtain the first matrix block corresponding to the Q matrix and the second matrix block corresponding to the K matrix, where the specified size is determined by the size of the Q matrix or the K matrix.
12. The device according to claim 11, characterized in that, The device further includes a test module configured to: Obtain a Q matrix, a K matrix, and a V matrix generated by a plurality of sample elements that need to perform attention calculation; Determine a plurality of candidate sizes, and respectively based on each candidate size, block the Q matrix and the K matrix generated by the plurality of sample elements to obtain a third matrix block corresponding to the Q matrix generated by the plurality of sample elements and a fourth matrix block corresponding to the K matrix generated by the plurality of sample elements, where the candidate sizes are determined by the sizes of the Q matrix and the K matrix; Based on the third matrix block, the fourth matrix block, and the V matrix generated by the plurality of sample elements corresponding to each candidate size, sequentially determine the attention calculation results corresponding to the plurality of sample elements, and respectively determine the time consumed for obtaining the attention calculation results each time; Based on the time consumed corresponding to each candidate size, determine a specified size among the plurality of candidate sizes.
13. The device according to any one of claims 10 to 12, characterized in that, The device further includes a mask processing module configured to: Determine the fifth matrix block in the same row of the relationship matrix as the sixth matrix block of the relationship matrix, where the fifth matrix block in the relationship matrix is calculated from the first matrix block in the same row of the Q matrix and the second matrix block in the same column of the K matrix; Obtain a Mask matrix corresponding to each sixth matrix block, where each Mask matrix has the same size as the corresponding sixth matrix block; Based on the Mask matrix corresponding to each sixth matrix block, perform mask processing on each sixth matrix block to obtain each sixth matrix block after mask processing.
14. The device according to claim 13, characterized in that, The calculation module is configured to: Based on each sixth matrix block after mask processing, perform the operation of the relationship matrix and the V matrix to obtain the attention calculation result corresponding to the attention calculation.
15. A computer program product comprising instructions, characterized in that, When the instruction is run on a computing device, the computing device is caused to execute the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, Includes computer program instructions, and when the computer program instructions are executed by a computing device, the computing device executes the method according to any one of claims 1 to 7.
Citation Information
Cited By
Method and apparatus for attention calculation in artificial intelligence model
WO2025156950A1