Long sequence data processing method and related equipment
By decomposing the long sequence data into multiple sets of target matrices and calculating the attention results, the problem of high computational complexity of the traditional transformer model is solved, and the calculation efficiency and model accuracy are improved.
Patent Information
- Application Number
- CN202311852495.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
The traditional transformer model has high computational complexity when processing long sequence data, resulting in inefficient computing.
The long sequence data is processed into M-group target matrix, and the attention results of each target matrix are calculated separately, and then the final attention results are obtained, reducing the matrix size and calculation complexity.
By reducing the matrix scale and computational complexity, the computing efficiency is improved, and the training accuracy and inference accuracy of the transformer model are improved.
Smart Images

Figure CN120235255A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI), and in particular, to a method for processing long sequence data and related devices. Background Art
[0002] AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. Simply put, artificial intelligence studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. With the development of AI technology, the transformer model plays an important role.
[0003] The transformer model applies the self-attention mechanism and needs to calculate the attention result for better training of the model and reasoning. However, in the traditional solution, the input sequence is regarded as a whole for calculating the attention result. The longer the length of the input sequence, the higher the computational complexity, resulting in serious consumption of computing power resources and low computing efficiency. Summary of the Invention
[0004] The embodiments of this application provide a method for processing long sequence data and related devices. The first sequence is processed to obtain M groups of target matrices, and M first attention results corresponding to the M groups of target matrices are calculated respectively, and then the second attention result corresponding to the first sequence is obtained by splicing. Compared with directly calculating the attention result by taking the first sequence as a whole, when calculating the first attention result of each group of target matrices, the scale of the matrix is reduced, thereby reducing the computational complexity and improving the computational efficiency.
[0005] In a first aspect, this application provides a method for processing long sequence data. The method is applied to a transformer model and includes:
[0006] Obtain a first sequence with a length of N, where N is an integer greater than or equal to 2. Based on the first sequence, obtain M groups of target matrices. Each group of target matrices includes a query matrix, a key matrix, and a value matrix, and M is an integer greater than or equal to 2. That is, by processing the first sequence, M groups of target matrices are obtained, and the scale of each group of target matrices is smaller than the scale of the matrix directly obtained by feature mapping of the first sequence. Calculate M first attention results corresponding to the M groups of target matrices respectively, and then splice the M first attention results to obtain the second attention result of the first sequence.
[0007] In this application, M groups of target matrices are obtained by processing the first sequence. M first attention results corresponding to the M groups of target matrices are calculated respectively, and then the second attention result corresponding to the first sequence is obtained by splicing. Compared with directly calculating the attention result by taking the first sequence as a whole, when calculating the first attention result of each group of target matrices, the scale of the matrix is reduced, so the calculation complexity is reduced and the calculation efficiency is improved.
[0008] In some optional implementation manners of the first aspect, obtaining M groups of target matrices based on the first sequence includes: first, obtaining the matrix corresponding to the first sequence by means of feature mapping. Then, dividing the matrix corresponding to the first sequence into M groups of target matrices.
[0009] In some optional implementation manners of the first aspect, obtaining M groups of target matrices based on the first sequence includes: first, splitting the first sequence to obtain M groups of second sequences. Then, mapping the features of the M groups of second sequences to the feature space to obtain M groups of target matrices, and the M groups of target matrices correspond to the M groups of second sequences one by one.
[0010] In this application, there are multiple ways to obtain M groups of target matrices. It can be either the way of first performing feature transformation on the first sequence and then splitting, or the way of first splitting and then performing feature transformation, which enriches the implementation manners of the technical solution of this application.
[0011] In some optional implementation manners of the first aspect, calculating the M first attention results corresponding to the M groups of target matrices includes:
[0012] Calculating the first attention result according to the first query matrix, the first key matrix, and the first value matrix. Among them, at least two of the first query matrix, the first key matrix, and the first value matrix correspond to different target matrices. In addition, the first query matrix is included in the M query matrices included in the M groups of target matrices; the first key matrix is the M key matrices included in the M groups of target matrices, and the first value matrix is included in the M value matrices included in the M groups of target matrices.
[0013] It can be understood that at least two of the first query matrix, the first key matrix, and the first value matrix correspond to different target matrices, which means that when calculating a first attention result, the inter-group interaction of different groups of target matrices is considered, and the inter-group dependence relationship of different groups of target matrices in the first sequence can be reflected. No matter in which field the transformer model is applied, it is beneficial to improve the training accuracy and inference accuracy of the model, so as to better achieve the training effect of the model and the inference accuracy in downstream tasks.
[0014] In some alternative embodiments of the first aspect, the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are different, including: the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are adjacent matrices.
[0015] In this application, in the solution where the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are adjacent in the first sequence, since the inter-group dependency relationship of different groups of target matrices at adjacent positions is stronger, this solution can more accurately reflect the inter-group dependency relationship of different groups of target matrices, further improving the accuracy of the transformer model.
[0016] In some alternative embodiments of the first aspect, the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are different, including: the target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are not adjacent to each other.
[0017] In this application, there are various possible cases where the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are different, enriching the implementation methods and application scenarios of the technical solutions of this application and improving the flexibility of the technical solutions.
[0018] In some alternative embodiments of the first aspect, calculating M first attention results corresponding to M groups of target matrices includes: calculating the first attention result according to the second query matrix, the second key matrix, and the second value matrix. Wherein, the target matrices corresponding to the second query matrix, the second key matrix, and the second value matrix are the same, and the second query matrix is included in the M query matrices included in the M groups of target matrices; the second key matrix is included in the M key matrices included in the M groups of target matrices, and the second value matrix is included in the M value matrices included in the M groups of target matrices.
[0019] In this application, it is also possible to determine a second query matrix corresponding to the same target matrix and the second key matrix to calculate a first attention result, which can be simpler during calculation and further simplifies the operation.
[0020] In some alternative embodiments of the first aspect, the matrix corresponding to the first sequence is divided into M groups of target matrices, including: evenly dividing the matrix corresponding to the first sequence to obtain M groups of target matrices. Wherein, the matrix corresponding to the first sequence is obtained by performing feature mapping on the first sequence, or is obtained by performing feature mapping on the padded first sequence. The length of the padded first sequence can be evenly divided into M parts.
[0021] In this application, the M groups of target matrices are obtained by evenly dividing the matrix corresponding to the first sequence, and the matrix scales of each group of target matrices are the same. Then, when calculating the first attention result corresponding to each group of target matrices, the same algorithm logic can be adopted, making the calculation process more convenient.
[0022] In some alternative embodiments of the first aspect, processing the first sequence to obtain M groups of second sequences includes: when the length N of the first sequence is evenly divisible by M, then evenly dividing the first sequence to obtain M groups of second sequences. That is to say, the lengths of each group of second sequences are the same. Or, when the length N is not evenly divisible by M, then padding the first sequence, and the length of the padded first sequence is an integer multiple of M. Then evenly divide the padded first sequence to obtain M groups of second sequences.
[0023] In this application, the matrix scales of the M groups of target matrices are the same. Then, when calculating the first attention result corresponding to each group of target matrices, the same algorithm logic can be adopted, making the calculation process more convenient. In the scenario where the first sequence cannot be evenly divided, padding the first sequence not only enables the padded first sequence to be evenly divided, but also does not affect the accuracy of the attention result.
[0024] In a second aspect, this application provides a processing device for long sequence data. The device applies a transformer model. The device includes:
[0025] An acquisition unit for acquiring a first sequence with a length of N, where N is an integer greater than or equal to 2;
[0026] A processing unit for obtaining M groups of target matrices based on the first sequence, where each group of target matrices includes a query matrix, a key matrix, and a value matrix, and M is an integer greater than or equal to 2.
[0027] The processing unit is further configured to calculate M first attention results corresponding to the M groups of target matrices, and splice the M first attention results to obtain a second attention result of the first sequence.
[0028] The processing device for long sequence data is used to implement the foregoing first aspect, or any possible implementation manner of the first aspect. For details, see the foregoing, and will not be elaborated here.
[0029] In a third aspect, the present application provides a processing device for long sequence data, including a processor and a memory. The processor stores instructions, and when the instructions stored in the memory run on the processor, the method shown in the foregoing first aspect or any possible implementation manner of the first aspect is implemented.
[0030] In a fourth aspect, the present application provides a computer-readable storage medium in which instructions are stored, and when the instructions run on a processor, the method shown in the foregoing first aspect or any possible implementation manner of the first aspect is implemented.
[0031] In a fifth aspect, the present application provides a computer program product, and when the computer program product is executed on a processor, the method shown in the foregoing first aspect or any possible implementation manner of the first aspect is implemented.
[0032] The beneficial effects shown in any one of the second aspect to the fifth aspect are similar to those in the first aspect or any possible implementation manner of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of an artificial intelligence entity framework provided by an embodiment of the present application;
[0034] Figure 2 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0035] Figure 3 It is a schematic flowchart of a method for processing long sequence data provided by an embodiment of the present application;
[0036] Figure 4 It is a schematic diagram of an algorithm summary provided by an embodiment of the present application;
[0037] Figure 5 It is a schematic diagram of an experimental result provided by an embodiment of the present application;
[0038] Figure 6a It is a schematic diagram of pseudocode provided by an embodiment of the present application;
[0039] Figure 6b It is a schematic diagram of pseudocode provided by an embodiment of the present application;
[0040] Figure 7 It is a schematic structural diagram of a processing device for long sequence data provided by an embodiment of the present application;
[0041] Figure 8 It is another schematic structural diagram of a processing device for long sequence data provided by an embodiment of the present application;
[0042] Figure 9A schematic structural diagram of the chip provided by the embodiment of the present application. Detailed implementation manners
[0043] The embodiment of the present application provides a method for processing long sequence data and related devices. The first sequence is processed to obtain M groups of target matrices, and M first attention results corresponding to the M groups of target matrices are calculated respectively, and then spliced to obtain the second attention result corresponding to the first sequence. Compared with directly calculating the attention result by taking the first sequence as a whole, when calculating the first attention result of each group of target matrices, the scale of the matrix is reduced, so the calculation complexity is reduced and the calculation efficiency is improved.
[0044] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0045] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices. In addition, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (item)" or a similar expression thereof refers to any combination of these items, including any combination of single (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0046] Please refer to Figure 1 , Figure 1 which shows a schematic diagram of an artificial intelligence main framework. This main framework describes the overall working process of the artificial intelligence system and is applicable to the general requirements of the artificial intelligence field.
[0047] The above artificial intelligence theme framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis).
[0048] The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom".
[0049] The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0050] (1) Infrastructure.
[0051] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external through sensors; the computing power is provided by intelligent chips, which include hardware acceleration chips such as central processing unit (CPU), neural-network processing unit (NPU), graphics processing unit (GPU), application specific integrated circuit (ASIC), or field programmable gate array (FPGA); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.
[0052] (2) Data.
[0053] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves Internet of Things data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, humidity, etc.
[0054] (3) Data processing.
[0055] Data processing usually includes methods such as data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0056] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0057] Inference refers to the process of simulating the intelligent reasoning mode of humans in a computer or intelligent system, and using formal information for machine thinking and problem-solving according to the inference control strategy. The typical function is search and matching.
[0058] Decision-making refers to the process of making decisions after the intelligent information is inferred, and usually provides functions such as classification, sorting, prediction, etc.
[0059] (4) General capabilities.
[0060] After the data is processed as mentioned above, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0061] (5) Intelligent products and industry applications.
[0062] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize the intelligent information decision-making, and realize the landing application. Its application fields mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, smart city, intelligent terminal, etc.
[0063] Next, please refer to Figure 2 , Figure 2 which is a schematic diagram of the application scenario of the long sequence data processing method provided by the embodiment of the present application.
[0064] The long sequence data processing method provided by the embodiment of the present application is applied to the transformer model. As Figure 2 shown, similar to the traditional model, the transformer model also needs to go through model training. After training until the model converges or reaches the preset accuracy, according to different downstream tasks, model inference is performed to obtain the inference result.
[0065] In practical applications, the transformer model is widely used in fields such as text, speech, and video. In different fields, the downstream tasks are also different. For example, in the text field, it can be used in scenarios such as text recognition, intelligent question answering, and text analysis. In the video field, it can be used in scenarios such as object recognition and path analysis. The present application does not limit this. When applying the transformer model to process long sequences, the long sequence data processing method provided by the embodiment of the present application can be applied.
[0066] Next, please refer to Figure 3 , Figure 3 , which is a schematic flowchart of a method for processing long sequence data provided by an embodiment of the present application, including:
[0067] 301. Obtain a first sequence of length N, where N is an integer greater than or equal to 2.
[0068] In model training or model inference using a transformer, whether in the field of NLP, CV, or speech processing, etc., the first device may obtain a first sequence of length N. The first sequence is processed through the neural network architecture of the transformer to obtain a training result or an inference result. The first device mentioned here is a device that applies the transformer.
[0069] Among them, the length of the first sequence being N means that the first sequence includes N symbols (tokens). In different fields, tokens may have different meanings. For example: in the field of NLP, a token can be understood as the smallest unit in text, and a token can be a word, number, punctuation mark, single letter, a Chinese character, or any other single symbol that can be called text analysis. Another example: in the field of CV, a token can be understood as a sequence of non-overlapping small patches obtained by cutting an image.
[0070] 302. Obtain M groups of target matrices based on the first sequence, where each group of target matrices includes a query matrix, a key matrix, and a value matrix, and M is an integer greater than or equal to 2.
[0071] After the first device obtains the first sequence, it processes the first sequence to obtain M groups of target matrices. In a specific implementation process, either the first sequence can be feature-transformed first and then segmented, or the first sequence can be segmented first and then feature-transformed. The present application does not limit the specific implementation manner, and the following will be described separately.
[0072] 1. Feature-transform the first sequence first and then segment it.
[0073] In some alternative implementation manners, the first device first obtains the matrix corresponding to the first sequence, and then divides the matrix corresponding to the first sequence into M groups of target matrices.
[0074] The first device obtains the matrix corresponding to the first sequence by feature-mapping the first sequence, mapping the features of the first sequence into the feature space to obtain the matrix corresponding to the first sequence. Among them, the matrix corresponding to the first sequence includes matrices in three dimensions: a query matrix, a key matrix, and a value matrix.
[0075] It can be understood that dividing the matrix corresponding to the first sequence into M groups of target matrices is to split the matrix corresponding to the long sequence, thereby reducing the floating point operations per second (Flops) and improving the efficiency of the transformer model in processing long sequences.
[0076] In the embodiments of the present application, there are multiple methods to divide the first sequence, which will be described separately below.
[0077] Optionally, if the length N of the first sequence can be evenly divided by M, then the matrix corresponding to the first sequence can be evenly divided into M groups of target matrices.
[0078] Optionally, if the length N of the first sequence cannot be evenly divided by M, then the first sequence can be padded. After padding, the length of the first sequence is an integer multiple of M. Perform feature mapping on the padded first sequence to obtain the matrix corresponding to the first sequence. Then evenly divide the matrix corresponding to the first sequence to obtain M groups of second sequences. That is to say, when N cannot be divided evenly by M, the first sequence will be padded. The padding operation is to expand the first sequence and fill in 0 in the first sequence, or in other words, the padding value is 0.
[0079] Second, first split the first sequence and then perform feature transformation.
[0080] In some optional embodiments, the first device first processes the first sequence to obtain M groups of second sequences. Then map the features of the M groups of second sequences to the feature space to obtain M groups of target matrices, and the M groups of target matrices correspond one-to-one with the M groups of second sequences.
[0081] Among them, dividing the first sequence with length N into M groups of second sequences is to divide the long sequence into multiple groups of short sequences, and from the perspective of reducing Flops, improve the efficiency of the transformer model in processing long sequences. In the embodiments of the present application, there are multiple methods to divide the first sequence, which will be described separately below.
[0082] In some optional embodiments, the first sequence is evenly divided into M groups of second sequences. That is to say, the length of each group of second sequences is the same. The prerequisite for achieving even division is that the length N of the first sequence can be evenly divided by M, that is, N is an integer multiple of M. Exemplarily, assume that the length of the first sequence is 12, then it can be divided into 4 groups of second sequences with 3 lengths in each group, and the length of each second sequence is 12 / 4 = 3.
[0083] In some alternative embodiments, if the length N of the first sequence cannot be evenly divided by M, then the first sequence is padded so that the length of the padded first sequence is an integer multiple of M. Then, the padded first sequence is evenly divided to obtain M groups of second sequences. That is to say, when N is not divisible by M, the first sequence will be padded. The padding operation is to expand the first sequence by filling it with 0s, or in other words, the padding value is 0.
[0084] Optionally, the first sequence can be divided in order. The length of each of the previous groups of second sequences is m. Since the length N cannot be evenly divided by M, the length of the last group of second sequences in the first sequence is less than m. A padding operation is performed on this last group so that the length of the last group is also m.
[0085] For example, assume that the length of the first sequence is 20 and the first sequence is divided into 7 groups of second sequences. Since 20 is not divisible by 7, when dividing the first sequence of length 20, the tail of the first sequence can be padded to make the length of the first sequence 21. The padded first sequence is evenly divided into 7 groups of second sequences, and each group of second sequences includes 3 subsequences. The above process can also be understood as dividing the first sequence in order, dividing it into a group of second sequences every time it reaches a length of 3. The length of the last group of second sequences is 2, and this last group is padded so that the length of the last group of second sequences is also 3.
[0086] Optionally, in the scheme where N is not divisible by M and the first sequence needs to be padded, the padded part can also be at other positions in the first sequence instead of at the tail, such as the head, the middle, etc. of the first sequence. The specific position is not limited here.
[0087] Exemplarily, assume that the length of the first sequence is 14 and the length of each group of second sequences is 5. Then, the length of the first sequence needs to be padded to 15. When padding, in addition to filling a subsequence of length 1 at the end of the first sequence, a subsequence of length 1 can also be filled at any position in the first sequence, such as at the beginning of the first sequence, in the middle of the first sequence, etc. The specific position is not limited here.
[0088] In some alternative embodiments, in the field of NLP, the filling position can be determined according to the meaning of the text corresponding to the first sequence. For example, assume that the text corresponding to the first sequence is "The weather today is really nice". Taking each character as a token, the length of the first sequence is 9. If the first sequence is divided into 5 groups of second sequences, then the length of the first sequence needs to be filled to 10 so that the length of each group of second sequences is the same, which is 2. Considering the meaning of the text corresponding to the first sequence, the first sequence can be divided into 5 groups of second sequences corresponding to the texts "today", "of", "weather", "really", and "nice". The group that needs to be filled is the second group, that is, the group corresponding to the text "of". Or rather, the determined filling position is between the subsequences corresponding to the texts "today" and "of", or between the subsequences corresponding to the texts "of" and "weather". A subsequence with a length of 1 is filled at the filling position so that the length of each group of second sequences is the same.
[0089] In summary, in this application, there are multiple ways to obtain M groups of target matrices. It can be the way of first performing feature transformation on the first sequence and then splitting, or the way of first splitting and then performing feature transformation, which enriches the implementation ways of the technical solutions of this application. In addition, the matrix scales of the M groups of target matrices are the same. Then, when calculating the first attention results corresponding to each group of target matrices, the same algorithm logic can be adopted, making the calculation process more convenient. In the scenario where the first sequence cannot be evenly divided, the first sequence is filled. The filling operation not only makes the filled first sequence divisible, but also does not affect the accuracy of the attention results.
[0090] It should be noted that in the descriptions of "1. First perform feature transformation on the first sequence and then split" and "2. First split the first sequence and then perform feature transformation" above, the matrix scales of the M groups of target matrices are the same. However, in practical applications, no matter which of the above schemes is shown, the matrix scales of the M groups of target matrices can also be different.
[0091] In some alternative embodiments, when the length N of the first sequence cannot be evenly divided by M, the first sequence can also be directly processed without filling operation to obtain M groups of target matrices.
[0092] Exemplarily, in the scenario of "1. First perform feature transformation on the first sequence and then split" introduced above, when the length N of the first sequence cannot be evenly divided by M, the first sequence can be directly subjected to feature mapping to obtain the matrix corresponding to the first sequence, and then the first matrix is divided into M groups of target matrices. The matrix scales of each group of target matrices can be partially the same. For example, the matrix scales of M - 1 groups of target matrices are the same, and the scale of 1 group of target matrices is different from that of the other M - 1 groups of target matrices.
[0093] Exemplarily, in the scenario of "2. First, segment the first sequence and then perform feature transformation" introduced above, when the length N of the first sequence cannot be evenly divided by M, the lengths of the M groups of second sequences obtained by segmenting the first sequence are not all the same. For example, assume the length of the first sequence is 20, and the first sequence is divided into 7 groups of second sequences. Among these 7 groups of second sequences, 6 groups have a length of 3 each, and 1 group has a length of 2.
[0094] In some alternative embodiments, in the field of NLP, when dividing the first sequence into M groups of second sequences, the division can be combined with the meaning of the text corresponding to the first sequence, and it is not necessary for the lengths of each group of second sequences to be the same. For example, assume the text corresponding to the first sequence is "Look, the scenery here is really beautiful!", according to the meaning of this text, the first sequence can be divided into 8 groups of second sequences, and the texts corresponding to these 8 groups of second sequences are "Look", ",", "here", "the", "scenery", "really", "beautiful", "!" respectively.
[0095] In this application, when dividing the first sequence with length N into M groups of second sequences, there are multiple possible ways, which enriches the implementation methods of the technical solution of this application and improves the flexibility of the technical solution of this application.
[0096] 303. Calculate M first attention results corresponding to the M groups of target matrices.
[0097] Each group of target matrices includes a query matrix, a key matrix, and a value matrix. Then the M groups of target matrices include M query matrices, M key matrices, and M value matrices. When calculating a first attention result, it is based on a query matrix, a key matrix, and a value matrix. In the embodiments of this application, when calculating the first attention result, it can either consider the dependency relationship between different groups of target matrices or not, which will be described separately below:
[0098] 1. The multiple matrices for calculating each first attention result correspond to different target matrices.
[0099] Generally speaking, according to the first query matrix, the first key matrix, and the first value matrix, calculate the first attention result. Among them, the first query matrix is included in the M query matrices included in the M groups of target matrices; the first key matrix is the M key matrices included in the M groups of target matrices, and the first value matrix is included in the M value matrices included in the M groups of target matrices.
[0100] Specifically, it includes the following processes: First, according to the first query matrix and the first key matrix, determine the attention score of the target matrix corresponding to the first query matrix. Then, perform normalization processing on the attention score to obtain the probability (P) of the target matrix corresponding to the first query matrix. Finally, according to this probability and the first value matrix, determine the first attention result of the target matrix corresponding to the first query matrix. Here, the target matrix corresponding to the first query matrix refers to the target matrix including the first query matrix.
[0101] The calculation process of each target matrix is similar to the process shown above. Calculate the first attention result corresponding to each target matrix respectively, so as to obtain M first attention results.
[0102] Among them, the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are different. That is to say, there are various relationships among the target matrices corresponding to at least two of the matrices in the position of the first sequence. Specifically, it includes: the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are adjacent matrices. Or, the target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are not adjacent to each other.
[0103] It should be noted that there are various possibilities for the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix to be adjacent in the first sequence: It can be that the 3 matrices correspond to 3 groups of adjacent target matrices. It can also be that the 3 matrices correspond to 2 groups of adjacent second sequences. It can also be that there are two groups of second sequences that are adjacent among the 3 groups of second sequences corresponding to the 3 matrices. Specifically, it is not limited here.
[0104] It should be noted that the position of the target matrix in the first sequence can also be understood as the position of the second sequence corresponding to the target matrix in the first sequence. The second sequence corresponding to the target matrix refers to the matrix obtained by performing feature mapping on this second sequence. That is to say, there is a corresponding relationship between a second sequence and the target matrix obtained by performing feature mapping on this second sequence.
[0105] It can be understood that the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are different, which means that when calculating a first attention result, the inter-group interaction of different groups of target matrices is considered, and the inter-group dependence relationship of different groups of target matrices in the first sequence can be reflected. Regardless of the field where the transformer model is applied, it is beneficial to improve the training accuracy and inference accuracy of the model, so as to better achieve the training effect of the model and the inference accuracy in downstream tasks.
[0106] There are multiple possible cases where the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are different, which enriches the implementation methods and application scenarios of the technical solution of this application and improves the flexibility of the technical solution. Further, in the solution where the target matrices corresponding to at least two of the first query matrix, the first key matrix, and the first value matrix are adjacent in the first sequence, since the inter-group dependence relationship of different groups of target matrices at adjacent positions is stronger, this solution can more accurately reflect the inter-group dependence relationship of different groups of target matrices and further improve the accuracy of the transformer model.
[0107] Second, the multiple matrices for calculating each first attention result correspond to the same target matrix.
[0108] Generally speaking, according to the second query matrix, the second key matrix, and the second value matrix, the first attention result is calculated. Among them, the second query matrix is included in the M query matrices included in the M groups of target matrices; the second key matrix is included in the M key matrices included in the M groups of target matrices, and the second value matrix is included in the M value matrices included in the M groups of target matrices.
[0109] Specifically, it includes the following process: First, according to the second query matrix and the second key matrix, determine the attention scores of the target matrix corresponding to the second query matrix. Then perform normalization processing on the attention scores to obtain the probability of the target matrix corresponding to the second query matrix. Finally, according to this probability and the second value matrix, determine the first attention result of the target matrix corresponding to the second query matrix. Among them, the target matrix corresponding to the second query matrix mentioned here refers to the target matrix including the second query matrix.
[0110] Among them, the target matrices corresponding to the second query matrix, the second key matrix, and the second value matrix are the same. That is to say, the second query matrix, the second key matrix, and the second value matrix are included in the same target matrix.
[0111] The calculation process of each target matrix is similar to the process shown above. Calculate the first attention results corresponding to each target matrix respectively, so as to obtain M first attention results.
[0112] In this application, based on the second query matrix and the second key matrix corresponding to the same target matrix, determine the second query matrix to calculate a first attention result, which is simpler in calculation and further simplifies the operation.
[0113] III. For multiple matrices calculating a part of the first attention results, they correspond to different target matrices, and for multiple matrices calculating the other part of the first attention results, they correspond to the same target matrix.
[0114] In this solution, the calculation of M first attention results combines the above two methods of calculating a first attention result. Which first attention results are calculated based on the query matrix, key matrix, and value matrix included in the same target matrix, and which first attention results are calculated based on the query matrix, key matrix, and value matrix included in different target matrices are not limited in this application.
[0115] It should be noted that no matter which of the above methods is used to calculate M first attention results, the M query matrices, M key matrices, and M value matrices included in the M groups of target matrices all participate in the operation.
[0116] The following combines examples to elaborate on the process of calculating a first attention result. Please refer to Figure 4 , Figure 4 which is a schematic diagram of the algorithm outline provided by the embodiment of this application.
[0117] In Figure 4 the shown embodiment, taking the length of the first sequence as 12 and processing the first sequence to obtain 4 groups of target matrices as an example. As Figure 4 shown, the first sequence includes `X1 to X 12 these 12 subsequences, and the matrix size corresponding to the first sequence is 12×128. Divide the first sequence into 4 groups of second sequences, namely G1, G2, G3, and G4. Among them, each group of second sequences includes 3 subsequences, and the matrix size corresponding to each group of target matrices is 3×128.
[0118] The query matrices, key matrices, and value matrices corresponding to these 4 groups of target matrices are as follows:
[0119]
[0120] Among them, Q represents the query matrix, and Q1 to Q4 respectively correspond to the second sequences G1 to G4; K represents the key matrix, and K1 to K4 respectively correspond to the second sequences G1 to G4; V represents the value matrix, and V1 to V4 respectively correspond to the second sequences G1 to G4. The value of i is any integer among 1, 2, 3, and 4. That is to say, the scales of the query matrix, key matrix, and value matrix included in any group of target matrices are all 3×128.
[0121] Taking the calculation of the first attention result of the target matrix A1 as an example, the calculation process of calculating a first attention result is described as follows:
[0122] First, calculate the attention score of a target matrix A1 That is to say, the attention score is the product of a query matrix and the transpose of a key matrix. In Figure 4 the illustrated embodiment, if the obtained result is represented as a matrix, the matrix scale is 3×3.
[0123] Then, perform normalization processing on Si based on the softmax function to obtain the probability P of this target matrix i = softmax(S i , dim = -1). In Figure 4 the illustrated embodiment, if the probability of the target matrix A1 is represented as a matrix, the matrix scale is 3×3.
[0124] Finally, the first attention result O i = P i ·V y . That is to say, the first attention result of the target matrix A1 is the product of the probability of the target matrix A1 and the first value matrix. In Figure 4 the illustrated embodiment, if the obtained result is represented as a matrix, the matrix scale is 3×128.
[0125] In the embodiment shown above, when calculating the first attention result corresponding to the target second sequence Gi, the query matrix, key matrix, and value matrix used are Qi, Kj, and Vy respectively. In Figure 4 the illustrated embodiment, the value ranges of i, j, and y are all any integer from 1 to 4.
[0126] It should be noted that in Figure 4 the schematic diagram, taking j = (i + 1) % M = y, that is, i and j are adjacent, and y is the same as j as an example for illustration. In practical applications, there are various possibilities for the values of i, j, and y. They can be partially or completely the same, or all different. Details are not elaborated here.
[0127] Optionally, Qi, Kj, and Vy may be the first query matrix, the first key matrix, and the first value matrix introduced above. That is to say, at least two of the target matrices corresponding to Qi, Kj, and Vy are different, and at least two of the values of i, j, and y are different. Specifically, at least two of the values of i, j, and y are adjacent, or the values of i, j, and y are all different. The so-called adjacent values mean that the value ranges are adjacent end to end. For example, in Figure 4 the illustrated embodiment, 1 and 4 are adjacent values.
[0128] Exemplarily, if at least two of the target matrices corresponding to Qi, Kj, and Vy are adjacent, then there are multiple possibilities for Qi, Kj, and Vy.
[0129] For example, it may be that 3 matrices correspond to 3 groups of adjacent second sequences. For instance, in Figure 4 the illustrated embodiment, the first query matrix, the first key matrix, and the first value matrix correspond to the second sequences G1, G2, and G3 in sequence.
[0130] For example, it may be that 3 matrices correspond to 2 groups of adjacent second sequences. For instance, in Figure 4 the illustrated embodiment, the first query matrix and the first key matrix correspond to the same group of second sequence G1; the first value matrix corresponds to the second sequence G3.
[0131] For example, it may be that two of the three groups of second sequences corresponding to 3 matrices are adjacent. For instance, in Figure 4 the illustrated embodiment, the first query matrix corresponds to the second sequence G1, the first key matrix corresponds to the second sequence G2, and the first value matrix corresponds to the second sequence G4.
[0132] Optionally, Qi, Kj, and Vy may be the second query matrix, the second key matrix, and the second value matrix introduced above. That is to say, the target matrices corresponding to Qi, Kj, and Vy are the same, and the values of i, j, and y are the same.
[0133] In addition, it should be noted that the adjacent target matrices mentioned in the embodiments of the present application refer to adjacent end to end, that is, the subsequences corresponding to the target matrices are adjacent end to end in the first sequence. For example, in Figure 4 the illustrated embodiment, the second sequence G4 and the second sequence G1 are adjacent, that is, the target matrix corresponding to the second sequence G4 and the target matrix corresponding to the second sequence G1 are adjacent.
[0134] In some alternative embodiments, the query matrix, key matrix, and value used to calculate the first group of attention may be defined by the following formula:
[0135] j = (i ± m) % M, y = (i ± m) % M, where 1 ≤ m < M and m is an integer. Here, i + m means taking values clockwise starting from i, and i - m means taking values counterclockwise starting from i.
[0136] In addition, it should be noted that when calculating the M first attention results corresponding to the M groups of target matrices in the embodiments of this application, parallel execution or serial execution can be performed. The following will be described separately:
[0137] Optionally, when the first device includes multiple AI cores, the process of calculating the M first attention results can be performed in parallel between different AI cores. Each AI core is used to calculate one or more first attention results, so that these multiple AI cores obtain the M first attention results. Among them, the number of first attention results calculated by each AI core can be the same or different, and specific details are not limited here.
[0138] In addition, when an AI core calculates multiple first attention results, the operations of multiple first attention results can be performed in parallel through a multi-threaded method; or the operations of these multiple first attention results can be performed serially. In the serial execution scheme, when calculating each first attention result, this AI core can first obtain the query matrix, key matrix, and value matrix required for calculating the current first attention result, reducing the resources used for each calculation. It can also first obtain all the query matrices, key matrices, and value matrices required for calculating these multiple first attention results, and then start calculating a single attention result to prepare a data basis for subsequent calculation of attention results, improving the practicality of the scheme.
[0139] Optionally, when the first device includes multiple AI cores, one of these multiple AI cores can calculate the M first attention results. The calculation of the M first attention results by this one AI core can be performed serially or in parallel, and specific details are not limited here.
[0140] Optionally, when the first device includes 1 AI core, this 1 AI core can calculate the M first attention results in parallel or serially, and specific details are not limited here.
[0141] It can be understood that in the serial calculation scheme, the data volume required for each calculation is small, the calculation complexity is low, the performance requirements for the device are not high, and the occupation of computing resources is reduced. In the parallel calculation scheme, multiple first attention results are calculated simultaneously, and the calculation efficiency is fast.
[0142] In summary, in the embodiments of the present application, there are various possibilities for the target matrices corresponding to the query matrix, key matrix, and value matrix used to calculate a first attention result, enriching the implementation manners of the technical solutions of the present application and enhancing the flexibility of the technical solutions of the present application. In the solutions where the target matrices corresponding to these matrices are different, the inter-group interaction of different groups of target matrices is considered, which can reflect the inter-group dependence relationship of different groups of target matrices in the first sequence, thereby improving the accuracy of the transformer model. In the solutions where the second sequences corresponding to these matrices are the same, the calculation method is simple and convenient, enhancing the practicability of the solutions.
[0143] Taking the case where the scale of each group of target matrices in M groups of target matrices is the same as an example, the beneficial effects of the technical solutions of the present application are further described below. Assume that the length of the first sequence is N and the dimension of the token is D.
[0144] The calculation process of traditional self-attention (i.e., directly calculating the attention result by taking the first sequence as a whole) is as follows:
[0145] Step 1, calculate S: the number of addition operations (Add) is N 2 ·(D - 1), and the number of multiplication operations (Mul) is N 2 ·D.
[0146] Step 2, calculate P: the number of exponential operations (Exp) is N 2 , the number of Add is N·(N - 1), and the number of division operations (Div) is N 2 .
[0147] Step 3, calculate O: the number of Add is N·D·(D - 1), and the number of Mul is N·D 2 .
[0148] The calculation process of group-attention provided in the embodiments of the present application (including processing the first sequence to obtain M groups of target matrices and calculating M first attention results) is as follows:
[0149] Step 1, calculate S: the number of Add is N 2 ·(D - 1) / M, and the number of Mul is N 2 ·D / M.
[0150] Step 2, calculate P: the number of Exp is the number of Add is the number of Div is
[0151] Step 3, calculate O: the number of Add is N·D·(D - 1) / M, and the number of Mul is N·D 2 / M.
[0152] The comparison results are shown in Table 1 as follows:
[0153] Table 1
[0154] Operation type Number of operations of traditional method Number of operations of the method of this application Add ND(N + D - 1) - N (ND(N + D - 1) - N) / G Mul ND(N + D) ND(N + D) / G Div <![CDATA[N 2 > <![CDATA[N 2 / G]]> Exp <![CDATA[N 2 > <![CDATA[N 2 / G]]>
[0155] Based on the comparison in Table 1, in the method provided by the embodiments of the present application, the flops in the calculation process are reduced by M times compared with the traditional method.
[0156] In addition, it should be noted that Figure 4 in the schematic diagram and the example in Table 1, it is taken as an example that the length of each group of target matrices in M groups of target matrices is the same. In actual applications, the matrix scales of different groups of target matrices can also be different. At this time, the calculation logic and process of the first attention result corresponding to each group of target matrices are similar to those in Figure 4 the embodiments shown, except that the matrix scales of each first attention result are not necessarily the same, but compared with directly calculating the attention result of the first sequence as a whole, the matrix scale during the operation is also reduced, thereby reducing the calculation complexity.
[0157] 304. Concatenate the M first attention results to obtain the second attention result of the first sequence.
[0158] After obtaining the M first attention results, the first device concatenates these M first attention results in order to obtain the attention result of the first sequence. The so-called concatenating in order means concatenating according to the positions of the second sequences corresponding to the M first attention results in the first sequence.
[0159] Exemplarily, taking the Figure 4 embodiments shown as an example. The 4 groups of second sequences obtained by dividing the first sequence are G1, G2, G3, and G4 respectively, and the positions of the 4 groups of second sequences in the first sequence are also sorted according to G1, G2, G3, and G4. Then, during concatenation, it is concatenated in the order of "the first attention result O1 corresponding to G1 → the first attention result O2 corresponding to G2 → the first attention result O3 corresponding to G3 → the first attention result O4 corresponding to G4" to ensure the accurate calculation of the attention result of the first sequence.
[0160] From the above description, it can be seen that in the embodiments of the present application, the first sequence is processed to obtain M groups of target matrices, the M first attention results corresponding to the M groups of target matrices are calculated respectively, and then concatenated to obtain the second attention result corresponding to the first sequence. Compared with directly calculating the attention result by taking the first sequence as a whole, when calculating the first attention result of each group of target matrices, the matrix scale is reduced, the calculation complexity is reduced, and the calculation efficiency is improved.
[0161] Furthermore, in combination with the experimental results, the beneficial effects of the technical solution of this application are described. Please refer to Figure 5 , Figure 5 which is a schematic diagram of the experimental results provided by an embodiment of this application.
[0162] As Figure 5 shown, when the sequence length is 8192 and the dimension of the token is 128, divided into 4 groups, the time consumption of the method (group attention) provided by the embodiment of this application is about 22.8% of the original method, and the performance improvement is about 80%
[0163] Through end-to-end (E2E) verification of the impact of this algorithm on model convergence, text classification is selected as the downstream task to verify the effect of the improved transformer model. Taking THUCNews as the dataset, the results obtained by training the classification network respectively are shown in Table 2 below:
[0164] Table 2
[0165] Method Accuracy Traditional method (self - attention) 0.898 Method of this application (group - attention) 0.903
[0166] It can be seen from Table 2 that the method for processing long sequence data provided by the embodiment of this application has improved accuracy in the classification task.
[0167] In addition, it should be noted that the main idea of the method for calculating the attention result provided by the embodiment of this application is to process a long sequence to obtain target matrices corresponding to multiple short sequences. In what circumstances to trigger this method can be determined according to the actual application needs, and specific details are not limited here.
[0168] Optionally, a length threshold can be set. When the length of the obtained sequence is greater than or equal to this length threshold, the sequence is determined as a long sequence, and the method for calculating the attention result provided by the embodiment of this application is adopted. For example, after step 301 and before step 302, add step 305. Determine whether N is greater than or equal to the length threshold. If so, execute steps 303 to 304; if not, calculate the attention result of the first sequence. At this time, the attention result of the first sequence is determined according to the matrix obtained by feature mapping of the first sequence, and its principle is similar to that of determining the first attention result described above, and will not be elaborated here.
[0169] In addition, it should be noted that the method for processing long sequence data provided by the embodiment of this application can be applied to different devices through an algorithm software development kit (SDK).
[0170] Exemplarily, please refer to Figure 6a and Figure 6b , Figure 6aand Figure 6b are both schematic diagrams of the pseudocode provided by the embodiments of the present application.
[0171] For the Transformer model under the PyTorch framework, such as Figure 6a , the processing method of the long sequence data provided by the embodiments of the present application is called through instructions. The specific operation instructions are implemented by the Figure 6b -shown pseudocode. Among them, Figure 6b the shown pseudocode takes the scheme where the calculation of the first attention result includes the interaction between groups as an example. In practical applications, it can also be a scheme that does not consider the interaction between groups, which is not limited here.
[0172] It can be understood that Figure 6a and Figure 6b are only schematic diagrams of the pseudocode and do not constitute a limitation to the SDK provided by the embodiments of the present application. In practical applications, other codes can also be used to implement the processing method of the long sequence data provided by the embodiments of the present application.
[0173] Next, please refer to Figure 7 , Figure 7 , which is a schematic structural diagram of the long sequence data processing device provided by the embodiments of the present application. As Figure 7 shown, the long sequence processing device 700 includes an acquisition unit 701 and a processing unit 702. The long sequence processing device 700 applies a transformer model.
[0174] In some optional implementation manners, the acquisition unit 701 is configured to acquire a first sequence with a length of N, where N is an integer greater than or equal to 2.
[0175] The processing unit 702 is configured to obtain M groups of target matrices based on the first sequence, where each group of target matrices includes a query matrix, a key matrix, and a value matrix, and M is an integer greater than or equal to 2; calculate M first attention results corresponding to the M groups of target matrices; and splice the M first attention results to obtain a second attention result of the first sequence.
[0176] In some optional implementation manners, the processing unit 702 is specifically configured to: acquire the matrix corresponding to the first sequence; and divide the matrix corresponding to the first sequence into M groups of target matrices.
[0177] In some optional implementation manners, the processing unit 702 is specifically configured to: process the first sequence to obtain M groups of second sequences; and map the features of the M groups of second sequences to the feature space to obtain M groups of target matrices, and the M groups of target matrices correspond to the M groups of second sequences one by one.
[0178] In some alternative embodiments, the processing unit 702 is specifically configured to: calculate a first attention result according to a first query matrix, a first key matrix, and a first value matrix, where at least two of the target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are different; wherein, the first query matrix is included in the M query matrices included in the M groups of target matrices; the first key matrix is included in the M key matrices included in the M groups of target matrices, and the first value matrix is included in the M value matrices included in the M groups of target matrices.
[0179] In some alternative embodiments, that at least two of the target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are different includes: at least two of the target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are adjacent matrices.
[0180] In some alternative embodiments, that at least two of the target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are different includes: the target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are not adjacent to each other.
[0181] In some alternative embodiments, the processing unit 702 is specifically configured to: calculate a first attention result according to a second query matrix, a second key matrix, and a second value matrix, where the target matrices corresponding to the second query matrix, the second key matrix, and the second value matrix are the same; wherein, the second query matrix is included in the M query matrices included in the M groups of target matrices; the second key matrix is one of the M key matrices included in the M groups of target matrices, and the second value matrix is included in the M value matrices included in the M groups of target matrices.
[0182] In some alternative embodiments, the processing unit 702 is specifically configured to: evenly divide the matrix corresponding to the first sequence to obtain M groups of target matrices; wherein, the matrix corresponding to the first sequence is obtained by performing feature mapping on the first sequence, or is obtained by performing feature mapping on the padded first sequence.
[0183] In some alternative embodiments, the processing unit 702 is specifically configured to: if the length N of the first sequence is evenly divisible by M, then evenly divide the first sequence to obtain M groups of second sequences; or, if the length N is not evenly divisible by M, then pad the first sequence so that the length of the padded first sequence is an integer multiple of M; evenly divide the padded first sequence to obtain M groups of second sequences.
[0184] The processing device 700 for long sequence data is used to implement the foregoing Figures 1 to 6bThe processing method of the long sequence data shown is not described herein again.
[0185] Please refer to Figure 8 , Figure 8 , which is a schematic structural diagram of a long sequence data processing device provided by an embodiment of the present application. The long sequence data processing device 800 includes a processor 801, a memory 802, a communication interface 803, and a bus 804. Among them, the processor 801, the memory 802, and the communication interface 803 communicate through the bus 804, and can also achieve communication through other means such as wireless transmission. The memory 802 stores program codes, and the processor 801 can call the program codes stored in the memory 802 to implement the Figures 1 to 6b processing method of the long sequence shown, which is not described herein again.
[0186] It should be understood that in the embodiment of the present application, the processor 801 may be a CPU, and the processor 801 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0187] The memory 802 may include a read-only memory and a random access memory, and provide instructions and data to the processor 801. The memory 802 may also include a non-volatile random access memory. For example, the memory 802 may also store information about the device type.
[0188] The memory 802 can be a volatile memory, a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0189] In addition to including a data bus, the bus 804 can also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clear illustration, all kinds of buses are labeled as bus 804 in the figure. The bus 840 can be a Peripheral Component Interconnect Express (PCIe) bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus 840 can be divided into an address bus, a data bus, a control bus, etc.
[0190] The processing device 800 for long sequence data can also include one or more communication interfaces, one or more operating systems, such as Windows Server TM , Mac OS X TM, Unix TM , Linux TM , FreeBSD TM etc.
[0191] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the chip provided by this application. The chip can be embodied as a neural network processor 900. The neural network processor 900 is mounted on the main CPU as a coprocessor and tasks are assigned by the host CPU. The core part of the neural network processor 900 is the arithmetic circuit 903. The controller 904 can control the arithmetic circuit 903 to extract data from the weight memory 902 or the input memory 901 and perform operations.
[0192] In some implementations, the arithmetic circuit 903 includes multiple processing units (process engine, PE) inside. In some implementations, the arithmetic circuit 903 can be a two-dimensional systolic array. The arithmetic circuit 903 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 903 is a general matrix processor.
[0193] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 902 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 901 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 908.
[0194] The unified memory 906 is used to store input data and output data. The weight data is directly transported through the direct memory access controller (DMAC) 905, and the DMAC transports it to the weight memory 902. The input data is also transported to the unified memory 906 through the DMAC.
[0195] The bus interface unit 910 (bus interface unit, BIU) can be used to realize the interaction between the main CPU, the DMAC, and the instruction fetch buffer 909 (instruction fetch buffer, IFB) through the bus. The instruction fetch buffer 909 is used to store the instructions used by the controller 904.
[0196] The vector calculation unit 907 includes multiple operation processing units, which further process the output of the operation circuit as needed, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as pixel-level summation, upsampling of feature planes, etc.
[0197] Among them, Figures 1 to 6b The operations of each layer in the neural network shown in the corresponding embodiments can be executed by the operation circuit 903 or the vector calculation unit 907.
[0198] Among them, the processor mentioned anywhere above can be a central processing unit, a microprocessor, or one or more integrated circuits for controlling the execution of the program of the method in the first aspect above.
[0199] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0200] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.
[0201] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0202] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0203] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
Claims
1. A method for processing long sequence data, characterized in that, The method is applied to a Transformer model, and the method includes: Obtain a first sequence of length N, where N is an integer greater than or equal to 2; Based on the first sequence, obtain M groups of target matrices, where each group of target matrices includes a query matrix, a key matrix, and a value matrix, and M is an integer greater than or equal to 2; Calculate M first attention results corresponding to the M groups of target matrices; Concatenate the M first attention results to obtain a second attention result of the first sequence.
2. The method according to claim 1, wherein The obtaining of the M groups of target matrices includes: Obtain the matrix corresponding to the first sequence; Divide the matrix corresponding to the first sequence into the M groups of target matrices.
3. The method according to claim 1, characterized in that, The obtaining of the M groups of target matrices includes: Process the first sequence to obtain M groups of second sequences; Map the features of the M groups of second sequences to a feature space to obtain the M groups of target matrices, and the M groups of target matrices correspond one-to-one to the M groups of second sequences.
4. The method according to any one of claims 1 to 3, characterized in that The calculating of the M first attention results corresponding to the M groups of target matrices includes: Calculate the first attention result according to a first query matrix, a first key matrix, and a first value matrix, where at least two of the first query matrix, the first key matrix, and the first value matrix correspond to different target matrices; Wherein, the first query matrix is included in the M query matrices included in the M groups of target matrices; the first key matrix is the M key matrices included in the M groups of target matrices, and the first value matrix is included in the M value matrices included in the M groups of target matrices.
5. The method according to claim 4, wherein At least two of the first query matrix, the first key matrix, and the first value matrix corresponding to different target matrices includes: at least two of the first query matrix, the first key matrix, and the first value matrix correspond to adjacent target matrices.
6. The method according to claim 4, wherein At least two of the first query matrix, the first key matrix, and the first value matrix corresponding to different target matrices includes: The target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are not adjacent to each other.
7. The method according to any one of claims 1 to 3, characterized in that, The calculating of the M first attention results corresponding to the M groups of target matrices includes: Calculate the first attention result according to a second query matrix, a second key matrix, and a second value matrix, where the target matrices corresponding to the second query matrix, the second key matrix, and the second value matrix are the same; Wherein, the second query matrix is included in the M query matrices included in the M groups of target matrices; the second key matrix is included in the M key matrices included in the M groups of target matrices, and the second value matrix is included in the M value matrices included in the M groups of target matrices.
8. The method according to any one of claims 2, 4 to 7, characterized in that The dividing of the matrix corresponding to the first sequence into the M groups of target matrices includes: Divide the matrix corresponding to the first sequence evenly to obtain the M groups of target matrices; Among them, the matrix corresponding to the first sequence is obtained by performing feature mapping on the first sequence, or is obtained by performing feature mapping on the padded first sequence.
9. The method according to any one of claims 3 to 7, characterized in that, The processing of the first sequence to obtain M groups of second sequences includes: If the length N of the first sequence is divisible by M, divide the first sequence evenly to obtain the M groups of second sequences; or, If the length N is not divisible by M, pad the first sequence so that the length of the padded first sequence is an integer multiple of M; Divide the padded first sequence evenly to obtain the M groups of second sequences.
10. A processing device for long sequence data, characterized in that, The device applies a transformer model, and the device includes: An acquisition unit for acquiring a first sequence of length N, where N is an integer greater than or equal to 2; A processing unit for obtaining M groups of target matrices based on the first sequence, where each group of target matrices includes a query matrix, a key matrix, and a value matrix, and M is an integer greater than or equal to 2; The processing unit is further configured to calculate M first attention results corresponding to the M groups of target matrices; The processing unit is further configured to splice the M first attention results to obtain a second attention result of the first sequence.
11. The device according to claim 10, characterized in that, Specifically, the processing unit is configured to: Obtain the matrix corresponding to the first sequence; Divide the matrix corresponding to the first sequence into the M groups of target matrices.
12. The device according to claim 10, wherein Specifically, the processing unit is configured to: Process the first sequence to obtain M groups of second sequences; Map the features of the M groups of second sequences to the feature space to obtain the M groups of target matrices, and the M groups of target matrices correspond to the M groups of second sequences one by one.
13. The device according to any one of claims 10 to 12, characterized in that, Specifically, the processing unit is configured to: Calculate the first attention result according to a first query matrix, a first key matrix, and a first value matrix, where at least two of the first query matrix, the first key matrix, and the first value matrix correspond to different target matrices; Among them, the first query matrix is included in the M query matrices included in the M groups of target matrices; the first key matrix is included in the M key matrices included in the M groups of target matrices, and the first value matrix is included in the M value matrices included in the M groups of target matrices.
14. The device according to claim 13, wherein, At least two of the first query matrix, the first key matrix, and the first value matrix corresponding to different target matrices includes: At least two of the first query matrix, the first key matrix, and the first value matrix correspond to adjacent target matrices.
15. The device according to claim 13, characterized in that, At least two of the first query matrix, the first key matrix, and the first value matrix corresponding to different target matrices includes: The target matrices corresponding to the first query matrix, the first key matrix, and the first value matrix are not adjacent to each other.
16. The device according to any one of claims 10 to 12, characterized in that, Specifically, the processing unit is configured to: Calculate the first attention result according to the second query matrix, the second key matrix, and the second value matrix, where the target matrices corresponding to the second query matrix, the second key matrix, and the second value matrix are the same; Among them, the second query matrix is included in the M query matrices included in the M groups of target matrices; the second key matrix is the M key matrices included in the M groups of target matrices, and the second value matrix is included in the M value matrices included in the M groups of target matrices.
17. The device according to any one of claims 11, 13 to 16, characterized in that, The processing unit is specifically configured to: Divide the matrix corresponding to the first sequence evenly to obtain the M groups of target matrices; Among them, the matrix corresponding to the first sequence is obtained by performing feature mapping on the first sequence, or by performing feature mapping on the padded first sequence.
18. The device according to any one of claims 12 to 16, characterized in that, The processing unit is specifically configured to: If the length N of the first sequence is divisible by M, divide the first sequence evenly to obtain the M groups of second sequences; or, If the length N is not divisible by M, pad the first sequence so that the length of the padded first sequence is an integer multiple of M; Divide the padded first sequence evenly to obtain the M groups of second sequences.
19. A processing device for long sequence data, characterized in that, It includes a processor, and the processor is coupled to a memory; Instructions are stored in the memory, and when the instructions run on the processor, the long sequence data processing device implements the method according to any one of claims 1 to 9.
20. A computer program product, characterized in that, When the computer program product is executed on a computer, the method according to any one of claims 1 to 9 is implemented.