Data processing method, computing device and storage medium
By grouping and multi-kernel computing the three-dimensional vectors of the input sequence, the problem of limited attention calculation of ultra-long input sequences in the prior art is solved, and attention calculation for processing longer input sequences without increasing memory is realized.
Patent Information
- Application Number
- CN202410954472.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-16
AI Technical Summary
The prior art is difficult to implement attention calculations for ultra-long input sequences without increasing memory, due to the capacity of processor memory.
By grouping at least one-dimensional vectors in the three-dimensional vectors determined by the input sequence, it is divided into multiple three-dimensional vectors, and using one of the multiple cores to perform attention-related calculations, obtain multiple sets of process data and perform fusion calculations to generate the final attention output.
It realizes attention calculations for ultra-long input sequences without increasing the number of processors, reducing communication costs between processors and improving the ability to process longer input sequences.
Smart Images

Figure CN119003960B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the field of data processing, and more specifically to a data processing method, a computing device, and a storage medium. Background Art
[0002] Attention Mechanism is widely used in machine learning fields such as natural language processing (NLP) and computer vision (CV) to help models pay more attention to important information when processing complex tasks, thereby improving performance. Specifically, the attention mechanism enables the model to dynamically assign different weights to each element in the input sequence when processing the input sequence, so that the model can pay more attention to the information in the input sequence that is relevant to the current task.
[0003] However, in the attention mechanism, the length of the input sequence that the model can process is limited by the processor memory. In other words, without increasing the memory, existing data processing methods find it difficult to implement attention calculations for ultra-long input sequences. Summary of the invention
[0004] In response to the above problems, the present invention provides a data processing method, which enables attention calculation for ultra-long input sequences without increasing memory.
[0005] According to a first aspect of the present invention, a data processing method is provided, characterized in that it includes: determining a three-dimensional vector based on an input sequence, the three-dimensional vector including: a query vector, a key vector and a value vector; grouping at least one-dimensional vector in the three-dimensional vector so as to divide the three-dimensional vector into a plurality of three-dimensional vector groups; for each three-dimensional vector group, performing attention-related calculations using one of a plurality of kernels to obtain a set of process data corresponding to each three-dimensional vector group; and performing fusion calculations on the calculated multiple sets of process data to determine final data for generating attention outputs corresponding to the input sequence.
[0006] In some embodiments, grouping at least one-dimensional vectors in a three-dimensional vector includes: grouping at least one-dimensional vectors in a three-dimensional vector to obtain multiple sub-vectors of the at least one-dimensional vector; and grouping the three-dimensional vectors based at least on the multiple sub-vectors to obtain multiple three-dimensional vector groups.
[0007] In some embodiments, each three-dimensional vector grouping includes at least one sub-vector of each dimension of the grouped three-dimensional vector.
[0008] In some embodiments, each sub-vector is included in only one three-dimensional vector grouping among the plurality of three-dimensional vector groups.
[0009] In some embodiments, each group of process data includes: first process data, second process data, and third process data. In these embodiments, obtaining a group of process data corresponding to each three-dimensional vector grouping includes: performing maximum value calculation on the query vector and the key vector in each three-dimensional vector grouping to obtain the first process data; performing sum calculation on the query vector and the key vector in each three-dimensional vector grouping to obtain the second process data; and performing product calculation on the query vector, the key vector, and the value vector in each three-dimensional vector grouping to obtain the third process data.
[0010] In some embodiments, the final data includes: first data, second data, and third data. In these embodiments, determining the final data for generating the attention output corresponding to the input sequence includes: performing maximum value calculation for all first process data in the plurality of sets of process data to determine the first data; determining the second data based on at least all second process data and the first data in the plurality of sets of process data; and determining the third data based on at least all third process data, the first data, and the second data in the plurality of sets of process data.
[0011] In some embodiments, the data processing method provided by the first aspect of the present invention also includes: at least based on the final data, calculating the gradient of each dimensional vector in the three-dimensional vector determined based on the input sequence to generate an attention output.
[0012] In some embodiments, calculating the gradient of each dimension of a three-dimensional vector determined based on an input sequence includes: dividing a plurality of variables obtained based on operations on the input sequence, the plurality of variables including: a three-dimensional vector, final data, and a gradient variable obtained by performing gradient operations based on the input sequence; and performing reverse calculation based on the divided plurality of variables to determine the gradient of each dimension of the three-dimensional vector.
[0013] In some embodiments, determining the gradient of each dimensional vector in a three-dimensional vector includes: arranging the divided multiple variables in a predetermined order to obtain an intermediate matrix for reverse calculation; for each row of the intermediate matrix, performing reverse calculation using one of a plurality of kernels to obtain multiple intermediate results, wherein each intermediate result is related to the gradient of a one-dimensional vector in the three-dimensional vector; and performing at least one of addition and concatenation operations on the intermediate results related to the gradient of the same-dimensional vector to determine the gradient of the dimensional vector.
[0014] According to a second aspect of the present invention, a computing device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can execute the method of the first aspect of the present invention.
[0015] According to a third aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method of the first aspect of the present invention.
[0016] According to a fourth aspect of the present invention, there is provided a computer program product tangibly stored on a non-transitory computer readable medium and comprising machine executable instructions which, when executed, cause a machine to perform the steps in the method of the first aspect of the present invention.
[0017] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements.
[0019] Figure 1 A flow chart of a data processing method according to an embodiment of the present invention is shown.
[0020] FIG. 2A to FIG. 2C A schematic diagram showing an exemplary embodiment of a data processing method of the present invention.
[0021] Figure 3 A flow chart of a data processing method for reverse calculation according to an embodiment of the present invention is shown.
[0022] FIG. 4A to FIG. 4C A schematic diagram showing an exemplary embodiment of a data processing method of the present invention.
[0023] Figure 5 A block diagram of a computing device suitable for implementing embodiments of the present invention is schematically shown. DETAILED DESCRIPTION
[0024] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0025] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0026] In the field of natural language processing, commonly used deep learning model architectures include: Transformer model based on the attention mechanism. Regarding the attention mechanism, it allows the model to focus on information at different positions in the input sequence when processing text, so as to help the model filter out information that is helpful to the model from a large amount of information. In other words, the core purpose of the attention mechanism is to filter out information that is more critical to the current task, so that attention can be focused on the filtered information.
[0027] According to the basic principle of the attention mechanism, a three-dimensional vector is usually generated based on the input sequence input into the model. The three-dimensional vector includes: query vector Q, key vector K, and value vector V. On this basis, based on the generated query vector Q, key vector K, and value vector V, attention calculation can be performed according to the following formula (1) to calculate the corresponding attention weight and attention output.
[0028]
[0029] Among them, d k Represents the dimension size of the key vector K. Specifically, the query vector Q and the key vector K can be subjected to a scaled dot product calculation such as calculation to obtain an attention weight matrix, and then the value vector V is weighted and summed based on the attention weight matrix to obtain the attention output.
[0030] In practice, the aforementioned attention calculation can be performed on processors such as a graphics processing unit (GPU) and a general-purpose graphics processing unit (GPGPU). Taking the GPU as an example, the existing data processing method for attention calculation includes: reading the input data from the GPU's high bandwidth memory (HBM) to the GPU's static random access memory (SRAM), so as to perform attention calculation on the input data in the SRAM, and return the calculation result to the HBM. Usually, the above-mentioned data processing method for attention calculation is completed in a core of the GPU. Since a lot of time is consumed in this process to continuously exchange data between the GPU's HBM and SRAM, the calculation speed is slow and the efficiency is low. In addition, since the GPU's SRAM capacity is small, and the size of the memory space required for attention calculation increases quadratically with the growth of the length N of the input sequence, for input sequences exceeding the threshold length, the attention calculation cannot be realized due to the GPU memory limitation.
[0031] In summary, the shortcoming of existing data processing methods is that it is impossible to implement attention calculation for ultra-long input sequences without increasing memory.
[0032] In some embodiments, in order to be able to process input sequences exceeding a threshold length, the processor memory used for attention calculations can be increased by increasing the number of processors. That is, the attention calculation to be performed can be split into multiple parts, such as multiple parts equal to the number of processors (such as GPUs), so that each GPU is responsible for a part of the attention calculation, and then the final attention output is obtained based on the calculation results of each GPU. However, when the number of GPUs reaches the threshold number, due to the memory limitations of each GPU, the attention calculations that can be performed will also be limited, making it difficult to process longer input sequences, such as input sequences with a length of more than one million tokens.
[0033] In order to at least partially solve one or more of the above problems and other potential problems, an exemplary embodiment of the present invention proposes a data processing scheme. In the scheme, at least one-dimensional vectors in the three-dimensional vectors determined based on the input sequence are grouped so as to divide the three-dimensional vectors into multiple three-dimensional vector groups; for each three-dimensional vector group, one of multiple cores is used to perform attention-related calculations to obtain a set of process data corresponding to each three-dimensional vector group; and the multiple sets of calculated process data are fused to determine the final data for generating the attention output corresponding to the input sequence, so that multiple cores of the same processor can be used to perform different parts of the attention-related calculations, so that each processor can process a longer input sequence, thereby realizing the execution of attention calculations for ultra-long input sequences without increasing the number of processors.
[0034] The following will be combined Figures 1 to 4C The data processing scheme of the present invention is described. Figure 1 The flowchart of the data processing method 100 according to the embodiment of the present invention is shown. It should be understood that the method 100 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this respect.
[0035] In step 102, a three-dimensional vector is determined based on an input sequence.
[0036] Regarding the three-dimensional vector, it refers to: query vector Q, key vector K and value vector V.
[0037] According to an embodiment of the present invention, a linear transformation may be performed on an input sequence to obtain a query vector Q, a key vector K, and a value vector V corresponding to the input sequence.
[0038] In step 104, at least one-dimensional vector in the three-dimensional vector is grouped so as to divide the three-dimensional vector into a plurality of three-dimensional vector groups.
[0039] Regarding grouping at least one dimensional vector in the three-dimensional vector, it means that at least one dimensional vector in the query vector Q, the key vector K, and the value vector V can be grouped to obtain multiple sub-vectors of the at least one dimensional vector. For example, in some embodiments, only the query vector Q in the three-dimensional vector can be grouped to obtain multiple query sub-vectors Q of the query vector Q. n In some other embodiments, the key vector K and the value vector V in the three-dimensional vector may be grouped to obtain a plurality of key sub-vectors K of the key vector K: n and multiple value subvectors V of the value vector V n In some other embodiments, the query vector Q, key vector K and value vector V in the three-dimensional vector may be grouped to obtain multiple query sub-vectors Q n , multiple key subvectors K n and multiple value subvectors V n .
[0040] According to an embodiment of the present invention, the grouped vectors in the three-dimensional vector may include two or more sub-vectors. For example, if the query vector Q is grouped, in some examples, the grouped query vector Q may include two query sub-vectors (such as the first query sub-vector Q1 and the second query sub-vector Q2); in another example, the grouped query vector Q may include three query sub-vectors (such as the first query sub-vector Q1, the second query sub-vector Q2, and the third query sub-vector Q3). The present invention is not limited to this.
[0041] Further, according to an embodiment of the present invention, when a multidimensional vector in a three-dimensional vector is grouped, the number of sub-vectors obtained for each vector grouping may be different or the same. For example, in the example of grouping a key vector K and a value vector V in a three-dimensional vector as described above, the key sub-vector K obtained is n The number of subvectors V n The same number of.
[0042] Regarding obtaining a plurality of three-dimensional vector groups, it may include: grouping the three-dimensional vectors based on at least a plurality of sub-vectors obtained by grouping at least one-dimensional vector in the three-dimensional vectors to obtain a plurality of three-dimensional vector groups.
[0043] For example, in the example of grouping only the query vector Q as described above, it is assumed that the query vector Q is grouped to obtain the first query sub-vector Q1 and the second query sub-vector Q2. In this case, the three-dimensional vectors can then be grouped based on the first query sub-vector Q1 and the second query sub-vector Q2, and a first three-dimensional vector grouping and a second three-dimensional vector grouping can be obtained, wherein the first three-dimensional vector grouping includes: the first query sub-vector Q1, the key vector K and the value vector V; the second three-dimensional vector grouping includes: the second query sub-vector Q2, the key vector K and the value vector V.
[0044] For another example, in another example of grouping the key vector K and the value vector V in the three-dimensional vector as described above, it is assumed that after grouping the key vector K, the first key sub-vector K1 and the second key sub-vector K2 are obtained, and after grouping the value vector V, the first value sub-vector V1 and the second value sub-vector V2 are obtained. In this case, the three-dimensional vector can then be grouped based on the first key sub-vector K1, the second key sub-vector K2, the first value sub-vector V1 and the second value sub-vector V2, and two three-dimensional vector groups are obtained, wherein the first three-dimensional vector group includes: query vector Q, the first key sub-vector K1 and the first value sub-vector V1; the second three-dimensional vector group includes: query vector Q, the second key sub-vector K2 and the second value sub-vector V2.
[0045] From the above, it can be seen that in an embodiment of the present invention, each three-dimensional vector grouping may include at least one sub-vector of each dimensional vector grouped in the three-dimensional vector, and further, according to some embodiments of the present invention, each sub-vector is included in only one three-dimensional vector grouping among multiple three-dimensional vector groups.
[0046] In step 106, for each three-dimensional vector group, a kernel among a plurality of kernels is used to perform attention-related calculations to obtain a set of process data corresponding to each three-dimensional vector group.
[0047] Regarding a set of process data, according to an embodiment of the present invention, it may include first process data (Max) determined based on maximum value calculation, second process data (Sum) determined based on summation calculation, and third process data (Q') determined based on vector product calculation.
[0048] Specifically, according to an embodiment of the present invention, obtaining a set of process data corresponding to each three-dimensional vector grouping may include: performing a maximum value calculation on the query vector and the key vector in each three-dimensional vector grouping to obtain first process data; performing a sum calculation on the query vector and the key vector in each three-dimensional vector grouping to obtain second process data; and performing a product calculation on the query vector, the key vector and the value vector in each three-dimensional vector grouping to obtain third process data.
[0049] Taking a three-dimensional vector grouping including a query vector Q, a first key sub-vector K1, and a first value sub-vector V1 as an example, according to an embodiment of the present invention, obtaining a set of process data corresponding to the three-dimensional vector grouping may include: performing a maximum value calculation on the query vector Q and the first key sub-vector K1 to obtain the first process data Max1; performing a sum calculation on the query vector Q and the first key sub-vector K1 to obtain the second process data Sum1; and performing a product calculation on the query vector Q, the first key sub-vector K1, and the first value sub-vector V1 to obtain the third process data Q1*. That is, in this example, a set of process data obtained for the three-dimensional vector grouping including the query vector Q, the first key sub-vector K1, and the first value sub-vector V1 includes: the first process data Max1 obtained based on the maximum value calculation for the query vector Q and the first key sub-vector K1; the second process data Sum1 obtained based on the sum calculation for the query vector Q and the first key sub-vector K1; and the third process data Q1' obtained based on the vector product calculation for the query vector Q, the first key sub-vector K1, and the first value sub-vector V1.
[0050] According to the inventive concept of the present invention, it is necessary to determine the process data corresponding to each three-dimensional vector grouping obtained in step 106. That is, when N three-dimensional vector groups are obtained in step 106, the process data determined in step 108 is N groups in total, where N is an integer greater than 1. For example, in some embodiments, two three-dimensional vector groups are obtained in step 106, and two groups of process data are determined in step 108.
[0051] In summary, according to the inventive concept of the present invention, multiple sub-vectors can be obtained by grouping at least one-dimensional vectors in the three-dimensional vector, and the three-dimensional vector can be divided based on these sub-vectors, so that the three-dimensional vector can be divided into multiple three-dimensional vector groups. On this basis, these three-dimensional vector groups can then be assigned to different kernels for attention-related calculations to obtain multiple groups of process data corresponding to these three-dimensional vector groups. As a result, each kernel only needs to load at least one of the vectors and sub-vectors included in the corresponding three-dimensional vector group, without having to load all three-dimensional vectors determined based on the input sequence, significantly reducing communication costs.
[0052] In step 108, a fusion calculation is performed on the calculated multiple sets of process data to determine final data for generating an attention output corresponding to the input sequence.
[0053] Regarding the final data, it may include first data (new_Max), second data (new_Sum) and third data (new_Q'), wherein the first data may be determined based on all first process data in multiple sets of process data; the second data may be determined based on at least all second process data in multiple sets of process data; and the third data may be determined based on at least all third process data in multiple sets of process data.
[0054] Specifically, according to an embodiment of the present invention, determining final data for generating an attention output corresponding to an input sequence includes: performing a maximum value calculation on all first process data in multiple sets of process data to determine first data; determining second data based on at least all second process data and first data in the multiple sets of process data; and determining third data based on at least all third process data, first data, and second data in the multiple sets of process data.
[0055] For example, in some embodiments, assuming that N groups of process data are obtained based on the aforementioned steps and the N groups of process data include N first process data (i.e., first process data Max1 to first process data MaxN), N second process data (i.e., second process data Sum1 to second process data SumN) and N third process data (i.e., third process data Q1' to third process data QN'), the first data new_Max, the second data new_Sum and the third data new_Q' in the final data can be calculated according to the following formulas (2)-(4).
[0056] new_Max = max (Max1, Max2, …, MaxN) (2)
[0057]
[0058] According to an embodiment of the present invention, the final result for generating the attention output can then be determined based on the final data. Specifically, the final result (Final_R) can then be calculated based on the second data new_Sum and the third data new_Q' in the final data according to the following formula (5).
[0059] Final_R = new_Q' / new_Sum (5)
[0060] According to some embodiments of the present invention, the above fusion calculation may be performed on other cores among the multiple cores except the core used in step 106 .
[0061] In summary, the present invention divides the three-dimensional vectors determined based on the input sequence into multiple three-dimensional vector groups, and distributes these three-dimensional vector groups to multiple cores for attention-related calculations; and introduces process data, and for each three-dimensional vector group, uses one of the multiple cores to perform attention-related calculations to determine the process data corresponding to each three-dimensional vector group, and performs fusion calculations on all the calculated process data to obtain the final result of the attention calculation, so that the ultra-long input sequence can be reasonably divided, and multiple cores of the same processor can be used to perform attention-related calculations, so that each processor can better process longer input sequences, thereby realizing the execution of attention calculations for ultra-long input sequences without increasing the number of processors.
[0062] As mentioned above, since the data needs to be read into SRAM for attention calculation, and the capacity of SRAM is usually small, according to some embodiments of the present invention, each dimensional vector in the three-dimensional vector grouping written into each core for attention-related calculation can be further tiled (Tiling), and then the three-dimensional vector grouping can be divided according to the tiled vectors to further divide the three-dimensional vector grouping into multiple small groups. On this basis, the process data of all small groups are traversed and calculated in an iterative manner, so as to calculate the process data corresponding to the three-dimensional vector grouping.
[0063] For example, the first process data Max, the second process data Sum and the third process data Q' can be configured for each core first, and the configured data are initialized, that is, Max = 0, Sum = 0, Q' = 0. In one example, based on the block operation of the vector, the three-dimensional vector group is divided into two small groups, which are recorded as the first group and the second group, wherein each group includes its own three-dimensional vector q, k, v. For the sake of clarity, the three-dimensional vector of the first group is recorded as q1, k1, v1, and the three-dimensional vector of the second group is recorded as q2, k2, v2. In this example, according to the inventive concept of the present invention, then the attention-related calculation can be performed based on the three-dimensional vector q1, k1, v1 of the first group, and the first intermediate process data corresponding to the first group is determined. In response to determining the intermediate process data corresponding to the first group, the attention-related calculation is then performed based on the three-dimensional vector of the second group, which is recorded as q2, k2, v2, and the second intermediate process data is calculated on the basis of the first intermediate process data. The second intermediate process data here can be regarded as the process data corresponding to the three-dimensional vector group. The specific data calculation method can be found in the previous article and will not be repeated here.
[0064] The following will be combined FIG. 2A to FIG. 2C Exemplary embodiments of the data processing scheme of the present invention are described.
[0065] FIG. 2A to FIG. 2C A schematic diagram showing an exemplary embodiment of a data processing method of the present invention.
[0066] like Figure 2A As shown, the three-dimensional vector determined based on the input sequence includes: a query vector Q, a key vector K and a value vector V. Further, in the Figure 2A In the illustrated embodiment, the key vector K and the value vector V are respectively divided into two groups of sub-vectors, that is, the key vector K is divided into a first key sub-vector K1 and a second key sub-vector K2, and the value vector V is divided into a first value sub-vector V1 and a second value sub-vector V2.
[0067] like Figure 2B As shown, according to an embodiment of the present invention, a first kernel kernel_1 may be used to calculate a first set of process data, namely Max1, Sum1 and Q1', based on a query vector Q, a first key sub-vector K1 and a first value sub-vector V1.
[0068] like Figure 2C As shown, according to an embodiment of the present invention, a second kernel kernel_2 may be used to calculate a second set of process data, namely Max2, Sum2 and Q2', based on the query vector Q, the second key sub-vector K2 and the second value sub-vector V2.
[0069] According to an embodiment of the present invention, the third kernel can then be used to calculate Figure 2B The first set of process data Max1, Sum1 and Q1' calculated in Figure 2C The second group of process parameters Max2, Sum2 and Q2' calculated in the above step are fused and calculated to determine the final data, namely the first data new_Max, the second data new_Sum and the third data new_Q'.
[0070] new_Max = max (Max1, Max2) (6)
[0071] new_Sum = Sum1*exp(Max1-new_Max) + Sum2*exp(Max2-new_Max) (7)
[0072] new_Q' = (Q1'*exp(Max1-new_Max)*Sum1 + Q'*exp(Max2-new_Max)*Sum2 (8)
[0073] Based on the calculated final data, the final result for generating the attention output can then be determined according to formula (5) as shown above, which will not be repeated here.
[0074] Generally, the above detailed description of the forward calculation part related to attention calculation based on the data processing scheme of the present invention is implemented. According to the inventive concept of the present invention, in such as Figure 1 On the basis of the data processing scheme of the data processing method 100 shown in the figure, the gradient of the three-dimensional vector determined based on the input sequence can also be calculated based on the final data determined for generating the attention output corresponding to the input sequence, so as to realize the reverse calculation part related to the attention calculation. Specifically, according to an embodiment of the present invention, the gradient of each dimension of the three-dimensional vector determined based on the input sequence can also be calculated based on at least the final data, so as to generate the attention output. Figure 3 Detailed description.
[0075] Figure 3 The flowchart of the data processing method 300 for reverse calculation according to the embodiment of the present invention is shown. It should be understood that the method 300 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this respect.
[0076] In step 302, multiple variables obtained based on the input sequence operation are divided, and the multiple variables include: a three-dimensional vector determined based on the input sequence, final data determined based on the three-dimensional vector determined based on the input sequence, and a gradient variable obtained by performing a gradient operation based on the input sequence.
[0077] Regarding the final data, the method of determining it can be found as above. Figure 1 The data processing method 100 shown is not described in detail here.
[0078] Regarding the gradient variable obtained by performing a gradient operation based on the input sequence, it may refer to a variable obtained by performing a gradient operation on the result obtained by performing operations related to attention calculation and loss function operation on the input sequence.
[0079] According to an embodiment of the present invention, multiple variables obtained by operation based on the input sequence can be divided according to the number of sub-vectors of the three-dimensional vector determined based on the input sequence during the forward calculation process. In some embodiments, the first data, the second data, and the third data in the final data obtained by the forward calculation can be divided according to the number of sub-vectors of the query vector Q. For example, in the forward calculation process, the query vector Q is divided into two groups, namely the first query sub-vector Q1 and the second query sub-vector Q2, then the first data, the second data, and the third data in the final data can be divided into two parts respectively; and when the query vector Q is not grouped, the first data, the second data, and the third data in the final data are not divided. In these embodiments, the query vector Q determined based on the input sequence and the gradient variable dO obtained by gradient operation based on the input sequence can also be divided according to the number of sub-vectors of the query vector Q. Similarly, if the query vector Q is divided into two groups during the forward calculation process, the query vector Q determined based on the input sequence and the gradient variable dO obtained by gradient operation based on the input sequence can be divided into two parts respectively. In these embodiments, the key vector K and the value vector V determined based on the input sequence may also be divided according to the number of subvectors of the key vector K. For example, in the forward calculation process, the key vector K is divided into two groups, namely the first key subvector K1 and the second key subvector K2, and accordingly the key vector K and the value vector V determined based on the input sequence may be divided into two parts respectively.
[0080] In step 304, a reverse calculation is performed based on the divided plurality of variables to determine the gradient of each dimensional vector in the three-dimensional vector determined based on the input sequence.
[0081] According to an embodiment of the present invention, determining the gradient of each dimensional vector in a three-dimensional vector determined based on an input sequence may include: arranging the divided multiple variables in a predetermined order to obtain an intermediate matrix for reverse calculation; for each row of the intermediate matrix, performing reverse calculation using one of the multiple kernels to obtain multiple intermediate results, wherein each intermediate result is related to the gradient of a one-dimensional vector in the three-dimensional vector; and performing at least one of addition and concatenation operations on the intermediate results related to the gradient of the same-dimensional vector to determine the gradient of the dimensional vector. FIG. 4A to FIG. 4C Detailed description.
[0082] FIG. 4A to FIG. 4C A schematic diagram showing an exemplary embodiment of the data processing method of the present invention is shown. It should be noted that, in this embodiment, in the preceding forward calculation process, the query vector Q and the key vector K determined based on the input sequence are divided into two groups.
[0083] like Figure 4A As shown, multiple variables obtained based on the input sequence operation can be divided respectively. Specifically, the first data new_Max and the second data new_Sum in the final data determined based on the three-dimensional vector determined based on the input sequence are divided into two parts respectively, wherein the first data new_Max is divided into new_Max_1 and new_Max_2, and the second data new_Sum is divided into new_Sum_1 and new_Sum_2; the query vector Q determined based on the input sequence, the third data new_Q' in the final data determined based on the three-dimensional vector determined based on the input sequence, and the gradient variable dO obtained by performing gradient operation based on the input sequence are divided into two parts respectively, wherein the query vector Q is divided into Q_1 and Q_2, the third data new_Q' is divided into new_Q'_1 and new_Q'_2, and the gradient variable dO is divided into dO_1 and dO_2; the key vector K and the value vector V determined based on the input sequence are divided into two parts respectively, wherein the key vector K is divided into K_1 and K_2, and the value vector V is divided into V_1 and V_2.
[0084] like Figure 4B As shown, the divided multiple variables can be arranged in a predetermined order to obtain an intermediate matrix for reverse calculation. The predetermined order may be related to the operation rules of the reverse calculation. Figure 4B An example of a predetermined order is shown. In some embodiments, the predetermined order may be based on Figure 4B The order shown is Figure 4A The obtained multiple variables are arranged to obtain an intermediate matrix for reverse calculation. It should be understood by those skilled in the art that the predetermined order here can be adjusted according to the operation rules of reverse calculation, and no limitation is made here.
[0085] Furthermore, reverse calculations can be performed for each row of the intermediate matrix to obtain multiple intermediate results. For example, Figure 4B The intermediate matrix shown in FIG. 1 includes four rows. Therefore, in this embodiment, performing reverse calculation for each row of the intermediate matrix is equivalent to performing four reverse calculations based on the intermediate matrix. Specifically, assuming that the first reverse calculation is performed for the first row of the intermediate matrix, the second reverse calculation is performed for the second row of the intermediate matrix, and so on, then in Figure 2B In the illustrated embodiment, taking the first row of the intermediate matrix as an example, the first row includes variables new_Max_1, new_Max_2, Q_1, new_Q'_1, dO_1, K_1, and V_1, and then three intermediate results can be calculated based on the aforementioned seven variables, wherein the intermediate result related to the gradient of the query vector Q is recorded as dQ_1, the intermediate result related to the gradient of the key vector K is recorded as dK_1, and the intermediate result related to the gradient of the value vector V is recorded as dV_1. Thus, according to four reverse operations based on the intermediate matrix, four intermediate results related to the gradient of the query vector Q (i.e., dQ_1, dQ_2, dQ_3, and dQ_4), four intermediate results related to the gradient of the key vector K (i.e., dK_1, dK_2, dK_3, and dK_4), and four intermediate results related to the gradient of the value vector V (i.e., dV_1, dV_2, dV_3, and dV_4) can be obtained. In addition, according to the inventive concept of the present invention, the reverse calculation for each row of the intermediate matrix can be distributed to different cores to improve the operation efficiency.
[0086] Then, at least one of the addition and concatenation operations may be performed on the intermediate results related to the gradient of the same-dimensional vector to determine the gradient of the dimensional vector. For example, at least one of the addition and concatenation operations may be performed on the intermediate results related to the gradient of the query vector Q to determine the gradient dQ of the query vector Q. That is, by performing at least one of the addition and concatenation operations on the intermediate results dQ_1, dQ_2, dQ_3, and dQ_4, at least one of the addition and concatenation operations may be obtained. Figure 4C The calculation process for the intermediate results dQ_1, dQ_2, dQ_3, and dQ_4 is shown by way of example. It should be understood that the calculation process can be modified and adjusted according to different intermediate results, and no limitation is made here.
[0087] like Figure 4CAs shown, the calculation process of determining the gradient dQ of the query vector Q includes: adding dQ_1 and dQ_2 to obtain a first addition result dQ_12; adding dQ_3 and dQ_4 to obtain a second addition result dQ_34; and concatenating the first addition result dQ_12 and the second addition result dQ_34 to obtain the gradient dQ of the query vector Q. Similarly, the key vector K and the value vector V can be calculated as follows: Figure 4C The same operations as shown will not be repeated here.
[0088] In summary, the data processing scheme of the present invention enables a single processor to handle attention calculations for ultra-long input sequences, reducing the communication cost between processors, and the scheme can realize forward and backward calculations of attention calculations without modifying the processor's core and interface.
[0089] Figure 5 The block diagram of a computing device 500 suitable for implementing an embodiment of the present invention is schematically shown. The device 500 may be a computer program product for implementing an embodiment of the present invention. Figure 1 The method 100 and Figure 3 The method 300 is shown in the apparatus. Figure 5 As shown, the device 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 502 or computer program instructions loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0090] Multiple components in the device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, and a storage unit 508. The processing unit 501 performs the various methods and processes described above, such as performing method 100 or method 300. For example, in some embodiments, method 100 or method 300 may be implemented as a computer software program, which is stored in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by CPU 501, one or more operations of method 100 or method 300 described above may be performed. Alternatively, in other embodiments, CPU 501 may be configured to perform one or more actions of method 100 and / or method 300 in any other appropriate manner (e.g., by means of firmware).
[0091] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0092] The computer program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present invention.
[0093] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of other programmable data processing devices, thereby producing a machine, so that when these instructions are executed by the processing unit of a computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other device to work in a specific manner.
[0094] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0095] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data processing method, characterized in that: include: Determine a three-dimensional vector based on the input sequence, the three-dimensional vector comprising: a query vector, a key vector, and a value vector; Grouping at least one-dimensional vector of the three-dimensional vectors so as to divide the three-dimensional vectors into a plurality of three-dimensional vector groups; For each three-dimensional vector group, use one of the multiple cores to perform attention-related calculations, so as to assign the multiple three-dimensional vector groups to different cores to perform attention-related calculations, so as to obtain a set of process data corresponding to each three-dimensional vector group; and The calculated multiple sets of process data are fused to determine final data for generating an attention output corresponding to the input sequence.
2. The method according to claim 1, characterized in that Grouping at least one-dimensional vector in the three-dimensional vector comprises: Grouping at least one-dimensional vector in the three-dimensional vector to obtain a plurality of sub-vectors of the at least one-dimensional vector; and The three-dimensional vectors are grouped at least based on the multiple sub-vectors to obtain the multiple three-dimensional vector groups.
3. The method according to claim 2, characterized in that Each of the three-dimensional vector groups includes at least one sub-vector of each dimension of the grouped three-dimensional vector.
4. The method according to claim 3, characterized in that Each sub-vector is included in only one three-dimensional vector grouping among the plurality of three-dimensional vector groups.
5. The method according to claim 1, characterized in that Each set of process data includes: first process data, second process data and third process data, The process data corresponding to each three-dimensional vector group is obtained as follows: Performing maximum value calculation on the query vector and the key vector in each three-dimensional vector group to obtain the first process data; performing a sum calculation on the query vector and the key vector in each three-dimensional vector group to obtain the second process data; and A product calculation is performed on the query vector, the key vector and the value vector in each three-dimensional vector group to obtain the third process data.
6. The method according to claim 5, characterized in that The final data includes: first data, second data and third data, Wherein, determining final data for generating an attention output corresponding to the input sequence includes: Performing maximum value calculation on all first process data in the plurality of sets of process data to determine the first data; determining the second data based on at least all second process data in the plurality of sets of process data and the first data; and The third data is determined based on at least all third process data in the plurality of sets of process data, the first data, and the second data.
7. The method according to claim 6, characterized in that Also includes: Based at least on the final data, a gradient of each dimension of the three-dimensional vector determined based on the input sequence is calculated for generating the attention output.
8. The method according to claim 7, characterized in that Calculating the gradient of each dimensional vector in the three-dimensional vector determined based on the input sequence includes: Dividing a plurality of variables obtained based on an input sequence operation, the plurality of variables comprising: the three-dimensional vector, the final data, and a gradient variable obtained based on a gradient operation of the input sequence; A reverse calculation is performed based on the divided plurality of variables to determine a gradient of each dimensional vector in the three-dimensional vector.
9. The method according to claim 8, characterized in that Determining the gradient of each dimension of the three-dimensional vector includes: Arranging the divided multiple variables in a predetermined order to obtain an intermediate matrix for the reverse calculation; For each row of the intermediate matrix, use one of the plurality of kernels to perform a reverse calculation to obtain a plurality of intermediate results, wherein each intermediate result is related to a gradient of a one-dimensional vector in the three-dimensional vector; and At least one of an addition operation and a concatenation operation is performed on intermediate results related to the gradient of the same-dimensional vector to determine the gradient of the same-dimensional vector.
10. A computing device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method of any one of claims 1 to 9.
12. A computer program product, characterized in that The computer program product is tangibly stored on a non-transitory computer readable medium and comprises machine executable instructions which, when executed, cause a machine to perform the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Method and device for performing attention operation, and storage medium
CN117707791A