Data parallel processing method and device

By converting the input data into an input sequence and splicing, the processing sequence is obtained based on parallel parameter segmentation, the problem of waste of computing resources caused by filling in the prior art is solved, and more efficient computing processing is achieved.

CN120144186AActive Publication Date: 2025-06-13SHANGHAI XIYU TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510629657.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

When processing input sequences of different lengths, the prior art fills other sequences by the length of the longest sequence, resulting in waste of computing resources and reducing processing efficiency.

Method used

The total sequence is obtained by converting the input data into an input sequence and splicing it according to the order of the input sequences. Then, the total sequence is segmented based on parallel parameters to obtain the processing sequence, and the processing sequence is processed in parallel through parallel devices to obtain ring attention.

Benefits of technology

The fill amount of input sequence is reduced, the waste of computing resources is reduced, the computing efficiency is improved, and the amount of data required to calculate the ring attention is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144186A_ABST
    Figure CN120144186A_ABST
Patent Text Reader

Abstract

The invention discloses a data parallel processing method and device, and the method comprises the steps: determining at least two input sequences according to at least one piece of input data, and splicing the input sequences according to the sequence of the input sequences, so as to obtain a total sequence; wherein each input sequence comprises at least one lexical element; parallel parameters are obtained, the total sequence is segmented based on the parallel parameters, at least two processing sequences are obtained, and the parallel parameters are determined at least according to the number of parallel devices; and performing parallel processing on the processing sequence through at least two parallel devices to obtain the ring attention. The method can reduce the filling amount of the input sequence, reduces the waste of computing resources, and improves the computing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a data parallel processing method and apparatus. Background Art

[0002] With the development of natural language model technology, natural language models have been increasingly widely used. When processing input data, a natural language model usually divides the input data into sentences, then converts the sentences into tokens, and forms sequences of different lengths according to the indexes of the tokens. When performing parallel processing on the mixed sequences formed by the input data, usually based on the length of the longest sequence, the ends of other sequences are padded. For example, if the mixed sequences are: [0,1,2,3,4,5,6,7,8,9,10], [0,1,2,3], [0,1,2], [0,1,2,3,4], and taking the longest first sequence as the benchmark to pad the other sequences, we will get: [0,1,2,3,4,5,6,7,8,9,10], [0,1,2,3,p,p,p,p,p,p,p], [0,1,2,p,p,p,p,p,p,p,p], [0,1,2,3,4,p,p,p,p,p,p], where p represents the padding bit. Finally, the padded sequences are processed separately by a parallel device.

[0003] After padding other sequences based on the length of the longest sequence and then performing parallel processing on the mixed sequences, the computing resources for processing the padded parts are wasted. Especially when the length differences of the mixed input sequences are large and the number of input sequences is large, a huge amount of computing resources will be wasted. Summary of the Invention

[0004] The present invention provides a data parallel processing method and apparatus to reduce the padding amount of input sequences, reduce waste of computing resources, and improve computing efficiency.

[0005] In a first aspect, an embodiment of the present invention provides a data parallel processing method, which includes:

[0006] Determine at least two input sequences according to at least one input data, and splice the input sequences according to the order of each input sequence to obtain a total sequence;

[0007] Wherein, each input sequence includes at least one token;

[0008] Obtain parallel parameters, and divide the total sequence based on the parallel parameters to obtain at least two processing sequences, where the parallel parameters are determined at least according to the number of parallel devices;

[0009] Perform parallel processing on a processing sequence through at least two parallel devices to obtain circular attention.

[0010] In a second aspect, an embodiment of the present invention further provides a data parallel processing device, which includes:

[0011] A sequence splicing module, configured to determine at least two input sequences, and splice the input sequences according to the order of each input sequence to obtain a total sequence;

[0012] Wherein, each input sequence includes at least one token;

[0013] A sequence splitting module, configured to obtain parallel parameters, and split the total sequence based on the parallel parameters and the input sequences to obtain at least two processing sequences, where the parallel parameters are determined at least according to the number of parallel devices;

[0014] A sequence parallel processing module, configured to perform parallel processing on the processing sequences through at least two parallel devices to obtain circular attention.

[0015] The technical solution of the embodiment of the present invention converts the input data into input sequences, splices the input sequences according to the order of the input sequences to obtain a total sequence, splits the total sequence through parallel parameters to obtain at least two processing sequences, and performs parallel processing on the processing sequences through at least two parallel devices to obtain circular attention. It solves the problem in the prior art that when processing at least two sequences of different lengths, it is necessary to pad other sequences based on the length of the longest sequence and then perform parallel processing on the mixed sequences, wasting computing resources and reducing the processing speed. It reduces the padding amount of the input sequences, reduces the waste of computing resources, and improves the computing efficiency; compared with other circular attention calculation methods in the prior art, it also reduces the required data volume for calculating circular attention and further improves the computing efficiency.

[0016] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 is a flowchart of a data parallel processing method provided in Embodiment 1 of the present invention;

[0019] Figure 2 It is a flowchart of a data parallel processing method provided in the second embodiment of the present invention;

[0020] Figure 3 It is a schematic diagram of masks of each processing sequence of a causal task provided in the second embodiment of the present invention;

[0021] Figure 4 It is a schematic diagram of masks of each processing sequence of a non-causal task provided in the second embodiment of the present invention;

[0022] Figure 5 It is a schematic diagram of pre-communication of a causal task provided in the second embodiment of the present invention;

[0023] Figure 6 It is a schematic diagram of pre-communication of a non-causal task provided in the second embodiment of the present invention;

[0024] Figure 7 It is a schematic diagram of the structure of a data parallel processing device provided in the third embodiment of the present invention;

[0025] Figure 8 It is a schematic diagram of the structure of an electronic device provided in the fourth embodiment of the present invention. Detailed implementation manners

[0026] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the embodiments of the present application, certain industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be regarded as exemplary. Their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0028] In the technical solution of the present application, the acquisition, transmission, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.

[0029] Embodiment 1

[0030] Figure 1 The flowchart of a data parallel processing method is provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation of parallel processing after converting input data into input sequences. This method can be executed by a data parallel processing device, which can be implemented in the form of hardware and / or software, and the data parallel processing device can be configured in an electronic device, and a natural language model is deployed in the electronic device.

[0031] As Figure 1 shown, the method includes:

[0032] S110. Determine at least two input sequences according to at least one input data, and splice the input sequences according to the order of each input sequence to obtain a total sequence.

[0033] Wherein, each input sequence includes at least one token.

[0034] Wherein, the input data refers to the input data of the natural language model, and the forms of the input data include but are not limited to various forms such as text, image, video, and audio. At the same time, the input data in this embodiment also supports multimodal mixing, such as text combined with image, etc. Converting the input data into input sequences, and converting multimodal input data into input sequences can be implemented in a conventional manner, and this embodiment does not limit this.

[0035] A token is a unit obtained after processing input data. Specifically, if the input data is text, the text can be tokenized or segmented to obtain tokens. For example, if the text is "I love you", the tokens obtained after tokenization are "I", "love", and "you". If the input data is an image, the image can be divided into patches, feature extraction can be performed on each patch, and vectorization conversion can be carried out. The cosine distance or Euclidean distance, etc., is calculated between the obtained vectors and a pre-set token vocabulary, and the token closest to each patch is found as the token corresponding to that patch.

[0036] The input sequence is an ordered sequence composed of at least one token. The applicable scenario of this embodiment is mainly used to process mixed input sequences, especially mixed input sequences of different lengths. Therefore, there are at least two input sequences in this embodiment.

[0037] According to the order of each input sequence, each input sequence is concatenated to obtain a total sequence. Therefore, the total sequence contains all the tokens obtained after the conversion of the input data, and each token is arranged in order. Exemplarily, taking the input sequence including the following four sequences as an example: [0,1,2,3,4,5,6,7,8,9,10], [0,1,2,3], [0,1,2], [0,1,2,3,4], the total sequence is [0,1,2,3,4,5,6,7,8,9,10,0,1,2,3,0,1,2,0,1,2,3,4]. It should be noted that 0-10 in the above input sequence and total sequence are only used to refer to the initial positions of each token in the original input sequence, and are not used to represent the actual content of the token.

[0038] In this embodiment, the input data of the natural language model is converted into an input sequence. For multiple input sequences of different lengths, they are concatenated in order to form a total sequence, which is convenient for subsequent sequence segmentation based on the total sequence. This way of sequence segmentation based on the total sequence formed by concatenating each input sequence in order to obtain multiple segmented sequences can reduce the padding amount compared to the method of padding each sequence according to the length of the longest sequence, thereby saving computing resources and improving the processing efficiency of the input data.

[0039] S120. Obtain parallel parameters, and based on the parallel parameters, segment the total sequence to obtain at least two processing sequences.

[0040] Among them, the parallel parameter is determined at least according to the number of parallel devices. The number of parallel devices can be determined in advance, for example, it can be 4. The parallel parameter can be the same as the number of parallel devices, or it can be an integer multiple of the number of parallel devices, such as 2 times. Further, when the natural language model performs a causal task on the input data, pre-communication needs to be performed on each parallel device. At this time, the value of the parallel parameter can be set to 2 times the number of parallel devices.

[0041] The processing sequences are the sequences obtained after dividing the total sequence based on the parallel parameter. When the parallel parameter is the same as the number of parallel devices, the number of processing sequences is the same as the number of input sequences, but the tokens corresponding to the processing sequences and the input sequences at the same position may be different. When the parallel parameter is an integer multiple of the number of parallel devices, the number of processing sequences is also an integer multiple of the input sequences.

[0042] In the prior art, since each sequence is padded according to the length of the longest sequence and then distributed to each parallel device for processing, for each parallel device, it is necessary to consume the computing resources required to process the longest sequence. Especially when the lengths of the mixed input sequences are different and the difference is large, it will cause a great waste of computing resources and reduce the processing efficiency of the input data.

[0043] In this embodiment, dividing the total sequence based on the parallel parameter can, on the one hand, ensure less sequence padding and reduce waste of computing resources; on the other hand, through token allocation and pre-communication, it can ensure that the number of tokens processed by each parallel device is balanced, thereby avoiding idleness or congestion of one or more parallel devices, ensuring the load balance of each parallel device, and improving the computing efficiency.

[0044] Further, S120 can further include:

[0045] S121. Determine the processing sequence length according to the total number of tokens in the total sequence and the parallel parameter;

[0046] S122. Divide the total sequence according to the processing sequence length to obtain at least two processing sequences.

[0047] In this embodiment, to ensure that each token in the total sequence can be processed, when dividing the total sequence based on the parallel parameter, if the total number of tokens in the total sequence cannot be evenly divided by the parallel parameter, the end of the total sequence can be padded, and then divided based on the parallel parameter, so that the total number of tokens in the padded total sequence can be evenly divided by the parallel parameter; or the value obtained by dividing the total number of tokens by the parallel parameter can be rounded up, and the obtained value can be used as the number of tokens in the processing sequence, for dividing the processing sequence, and padding the end of the last processing sequence.

[0048] In this embodiment, padding and splitting are performed based on the total sequence. Such a setting can ensure that the number of padded elements does not exceed the value obtained by dividing the total number of tokens by the parallel parameter, reducing the number of sequence padding; at the same time, the number of tokens in each processed sequence obtained after padding and splitting is the same, achieving load balancing among parallel devices; and the number of tokens in each processed sequence and the padding amount for sequence padding based on an integer multiple of the parallel parameter are less than the number of tokens in the input sequence after padding each input sequence based on the length of the longest sequence in the prior art, thereby reducing the consumption of computing resources and improving the processing efficiency of input data.

[0049] Further, S121 can further include: if the total number of tokens in the total sequence is an integer multiple of the parallel parameter, then the value obtained by dividing the total number of tokens in the total sequence by the parallel parameter is used as the length of the processed sequence; otherwise, the length of the processed sequence is calculated by the following formula: len = ┌a┐, a = T / D, where len represents the length of the processed sequence, ┌ ┐ represents ceiling calculation, T represents the total number of tokens in the total sequence, and D represents the parallel parameter.

[0050] In this embodiment, if the total number of tokens in the total sequence is an integer multiple of the parallel parameter, then directly use the value obtained by dividing the total number of tokens in the total sequence by the parallel parameter as the length of the processed sequence without padding. Otherwise, to ensure that all tokens are evenly processed by the parallel devices, the value obtained by dividing the total number of tokens by the parallel parameter is rounded up as the length of the processed sequence.

[0051] Exemplarily, taking the above total sequence [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 0, 1, 2, 3, 0, 1, 2, 0, 1, 2, 3, 4] as an example, the total number of tokens is 23. If the parallel parameter is 4, then the total number of tokens in the total sequence is not an integer multiple of the parallel parameter. The value obtained by dividing the total number of tokens by the parallel parameter (23 / 4 = 5.75) is rounded up to get 6 as the length of the processed sequence.

[0052] Further, S122 can further include: if the total number of tokens in the total sequence is an integer multiple of the parallel parameter, then the total sequence is split according to the length of the processed sequence; otherwise, the total sequence is split according to the length of the processed sequence, and the end of the last processed sequence is padded to make the number of tokens in the last processed sequence consistent with the length of the processed sequence.

[0053] In this embodiment, if the total number of tokens in the total sequence is an integer multiple of the parallel parameter, there is no need to pad the sequence, and the total sequence can be directly split according to the length of the processed sequence.

[0054] Exemplarily, if the total number of tokens in the total sequence is 24 and the parallel parameter is 4, then directly form a processed sequence with 6 tokens in order.

[0055] If the total number of tokens in the total sequence is not an integer multiple of the parallel parameter, the total sequence is segmented according to the processed sequence length obtained by rounding up the value obtained by dividing the total number of tokens by the parallel parameter. For the last processed sequence, the number of tokens is less than the processed sequence length. Therefore, the end of the last processed sequence is padded so that the number of tokens in the last processed sequence is the same as the processed sequence length. Exemplarily, taking the above total sequence [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 0, 1, 2, 3, 0, 1, 2, 0, 1, 2, 3, 4] as an example, the total number of tokens is 23, the parallel parameter is 4, and the processed sequence length is 6. Then the processed sequences after segmentation are: [0, 1, 2, 3, 4, 5], [6, 7, 8, 9, 10, 0], [1, 2, 3, 0, 1, 2], and [0, 1, 2, 3, 4]. Except for the last processed sequence, the number of tokens in the other processed sequences is 6. Therefore, the end of the last processed sequence is padded to get [0, 1, 2, 3, 4, p]. Here, p refers to padding, and the content of the padding is not limited in this embodiment. Preferably, the padding content is 0.

[0056] S130. Parallelly process the processed sequence through at least two parallel devices to obtain ring attention.

[0057] In this embodiment, each parallel device parallelly processes the processed sequence based on the ring attention mechanism to obtain the ring attention corresponding to the input data. Ring Attention is a distributed optimization technology aimed at breaking through the memory and computing bottlenecks of the traditional attention mechanism when processing long sequences. Its core idea is to improve the sequence parallel processing and memory efficiency based on the ring communication topology and block computing strategy.

[0058] This embodiment adopts the ring attention mechanism to calculate the ring attention corresponding to the input data. Since when performing the total sequence segmentation, tokens that originally belonged to the same input sequence may be segmented into at least two processed sequences respectively, and by adopting the ring attention mechanism, the relevance between the context can be noticed, and there is still a causal relationship within the segmented processed sequences. Therefore, splitting an input sequence into multiple segments will not significantly reduce the quality of natural language model training or inference.

[0059] The technical solution of the embodiment of the present invention converts the input data into an input sequence, splices the input sequences according to the order of the input sequences to obtain a total sequence, divides the total sequence by parallel parameters to obtain at least two processing sequences, and performs parallel processing on the processing sequences by at least two parallel devices to obtain circular attention. It solves the problem in the prior art that after padding other sequences based on the length of the longest sequence and then performing parallel processing on the mixed sequence, it wastes computing resources and reduces the processing speed, reduces the padding amount of the input sequence, reduces the waste of computing resources, and improves the computing efficiency.

[0060] Embodiment 2

[0061] Figure 2 It is a flowchart of a data parallel processing method provided by Embodiment 2 of the present invention. On the basis of the above embodiment, the present invention further specifies the process of obtaining circular attention by performing parallel processing on each processing sequence by each parallel device.

[0062] As Figure 2 shown, the method includes:

[0063] S210. Determine at least two input sequences according to at least one input data, and splice the input sequences according to the order of each input sequence to obtain a total sequence.

[0064] S220. Obtain parallel parameters, and divide the total sequence based on the parallel parameters to obtain at least two processing sequences.

[0065] The method of converting the input data into an input sequence, splicing each input sequence, and then dividing based on the parallel parameters to obtain the processing sequence has been described in the above embodiment, and will not be elaborated here in this embodiment.

[0066] S230. Obtain the task type matching the input data.

[0067] Among them, the task type includes causal tasks and non-causal tasks. A causal task refers to a task whose completion depends on the changes of the input data at different times and the observation of the order of the input data before and after, and a causal task is implemented based on a causal model. A non-causal task refers to a task whose completion mainly focuses on the correlation or pattern of the input data, and a non-causal task is implemented based on a non-causal model.

[0068] S240. Construct a mask matching each processing sequence according to the task type and the input sequence.

[0069] The mask refers to the attention mask, which is used to shield the information that does not need to be attended to or highlight the information that needs to be attended to when calculating the attention weights, enabling the artificial intelligence model to focus on specific parts of the processing sequence.

[0070] The attention mask is usually a binary tensor with the same length as the processing sequence. When calculating the attention scores subsequently, the mask is multiplied by the attention score tensor. The attention scores corresponding to the positions where the value in the mask is 0 will be set to an extremely small value (usually negative infinity). After passing through the softmax function, the attention weights at these positions will approach 0 and thus be ignored by the natural language model; while the positions where the value in the mask is 1 allow the normal calculation of attention weights, enabling the model to attend to the information corresponding to these positions.

[0071] In this embodiment, there are certain differences in constructing the mask according to different task types. This is related to the principles of causal tasks and non-causal tasks themselves. Causal models model and predict data based on causal relationships. When constructing the mask, the purpose is to ensure that when predicting the output at the current moment, the model can only utilize the information from the past and the current moment, and cannot utilize future information, in order to conform to the logic of causal relationships. Non-causal models do not rely on strict causal relationships. The purpose of constructing their masks is usually to control the model's attention to different information, or to handle some special structures in the data, such as masking out some irrelevant regions or padding parts in images or texts.

[0072] Furthermore, if the task type is a causal task, S240 can further include: constructing a mask matching each processing sequence based on the unidirectional attention mechanism;

[0073] If the task type is a non-causal task, S240 can further include: constructing a mask matching each processing sequence based on the bidirectional attention mechanism;

[0074] The unidirectional attention mechanism and the bidirectional attention mechanism are two different ways for processing sequence data in fields such as natural language processing. The unidirectional attention mechanism is the forward attention mechanism. In the unidirectional attention mechanism, when processing the input data, the information only flows in one direction, from the starting position of the sequence to the ending position. Thus, when calculating the attention weights, only the sequence information before the current position is considered to calculate the attention degree for each position. In the bidirectional attention mechanism, when processing the input data, the information can flow in two directions, both from the starting position of the sequence to the ending position and from the ending position to the starting position. When calculating the attention weights, all the sequence information before and after the current position will be comprehensively considered.

[0075] Therefore, causal tasks conform to the unidirectional attention mechanism, and the masks of the processing sequences of causal tasks usually take the form of a lower triangular matrix. Figure 3 A schematic diagram of the mask of the processing sequences of a causal task is provided. Figure 3 Taking the processed sequences after segmentation as [0, 1, 2, 3, 4, 5], [6, 7, 8, 9, 10, 0], [1, 2, 3, 0, 1, 2], and [0, 1, 2, 3, 4] as examples, the mask is described. Specifically, in the second processed sequence, each token comes from two different input sequences. Therefore, the second processed sequence can be further decomposed into [6, 7, 8, 9, 10] and [0]. Similarly, for the third processed sequence, it can be further decomposed into [1, 2, 3] and [0, 1, 2]. Therefore, in this example, circular attention needs to be calculated for six processed sequences, and the mask of the causal task is as Figure 3 shown by the red diagonal shaded part in, presenting the form of six lower triangles, and the six lower triangles respectively correspond to the six processed sequences after decomposition.

[0076] As Figure 3 shown, the computational resources required for adopting the technical solution of this embodiment can be compared with the prior art through the total area of the lower triangles. It should be noted that the total area of the lower triangles in this embodiment has no actual meaning and is not used to represent the actual required computational resources, but is only used to compare the computational resource consumption with the prior art. Adopting the technical solution of this embodiment, taking Figure 3 as an example, the sum of the areas of the 6 lower triangles is 6×6 / 2 + 5×5 / 2 + 1×1 / 2 + 3×3 / 2 + 3×3 / 2 + 6×6 / 2 = 58. If the prior art is adopted, the four input sequences after padding based on the longest input sequence are: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10], [0, 1, 2, 3, p, p, p, p, p, p, p], [0, 1, 2, p, p, p, p, p, p, p, p], [0, 1, 2, 3, 4, p, p, p, p, p, p]. At this time, the sum of the areas of the lower triangles is 11×11×4 / 2 = 242. At this time, not only more computational resources need to be occupied, but also a large amount of computational resources will be wasted in processing the invalid padding part. Comparing the area of the lower triangles in this embodiment with the area of the lower triangles of the prior art, it can be seen that this embodiment will greatly reduce the consumption and waste of computational resources and improve the processing efficiency of input data.

[0077] Non-causal tasks conform to the bidirectional attention mechanism, and the masks of the processing sequences of non-causal tasks usually take the form of a two-dimensional matrix. Figure 4 A schematic diagram of the mask of the processing sequences of a non-causal task is provided. Figure 4Taking the processed sequences after splitting as: [0, 1, 2, 3, 4, 5], [6, 7, 8, 9, 10, 0], [1, 2, 3, 0, 1, 2] and [0, 1, 2, 3, 4] as an example, the mask will be described. Similarly, Figure 4 According to the input sequence to which each token belongs, each processed sequence is further decomposed to obtain six processed sequences. In this example, when performing circular attention calculation on the six processed sequences, the mask for non-causal tasks is as Figure 4 shown by the blue square shaded part in, presenting the form of six squares, and the six squares respectively correspond to the six processed sequences after decomposition.

[0078] As Figure 4 shown, the computational resources required for adopting the technical solution of this embodiment can be compared with the prior art through the total area of the squares. It should be noted that the total area of the squares in this embodiment has no actual meaning and is not used to represent the actually required computational resources, but is only used for comparing the computational resource consumption with the prior art. Adopting the technical solution of this embodiment, the sum of the areas of the squares is 6×6 + 5×5 + 1×1 + 3×3 + 3×3 + 6×6 = 116. If the prior art is adopted, the four input sequences after padding based on the longest input sequence are: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10], [0, 1, 2, 3, p, p, p, p, p, p, p], [0, 1, 2, p, p, p, p, p, p, p, p], [0, 1, 2, 3, 4, p, p, p, p, p, p], and the sum of the areas of the squares at this time is 11×11×4 = 484. At this time, not only more computational resources need to be occupied, but also a large amount of computational resources will be wasted in processing the invalid padding part. Comparing the area of the squares in this embodiment with the area of the squares adopting the prior art, it can be seen that this embodiment will greatly reduce the consumption and waste of computational resources and improve the processing efficiency of input data.

[0079] In this embodiment, the part necessary for calculating circular attention is highlighted through the mask, reducing the data volume and computational complexity, and improving the computational efficiency. Combining the mask with the sequence padding and splitting methods in this embodiment can not only greatly reduce the occupation of computational resources, improve the processing efficiency of input data, but also reduce the invalid padding part and the computational resource waste caused by processing the invalid padding part.

[0080] Furthermore, if the task type is a causal task, this embodiment further includes: performing pre-communication on each parallel device, and the pre-communication includes: allocating the query vectors, key vectors, and value vectors of each token required for processing the causal task to each parallel device.

[0081] In this embodiment, pre-communication is used to achieve load balancing among parallel devices in causal tasks and improve computing efficiency.

[0082] Specifically, according to the number of parallel devices, the rearrangement of each token in each processing sequence is performed. Each token is divided into the number of parallel devices × 2 parts and is represented by different Q (Queues). Figure 5 A schematic diagram of pre-communication for causal tasks is provided, as Figure 5 shown. Taking four parallel devices as an example, the tokens in each processing sequence are divided into 8 parts, respectively named Q 0 -Q 7 . According to Figure 5 shown, Q 0 +Q 7 (red diagonal shaded part), Q 1 +Q 6 (yellow diagonal shaded part), Q 2 +Q 5 (green diagonal shaded part), and Q 3 +Q 4 (blue diagonal shaded part). The area represents the amount of data to be processed. Through the above distribution, the area corresponding to each color is the same. Therefore, Q 0 +Q 7 , Q 1 +Q 6 , Q 2 +Q 5 , and Q 3 +Q 4 are respectively allocated to four parallel devices for processing, which can ensure the load balancing of each parallel device.

[0083] In Figure 5 , KV 0 -KV 7 are respectively the key-value pairs (Key-Value) corresponding to Q 0 -Q 7 . Therefore, during pre-communication, when allocating tokens to each parallel device, the query vector (Query Vector), key vector (Key Vector), and value vector (Value Vector) corresponding to the tokens need to be pre-allocated to each parallel device. Taking Figure 5 as an example, the KV 0 and KV 7 corresponding to Q 0 and Q 7 are allocated to the first parallel device, and the KV 1 and KV 6 corresponding to Q 1 and Q 6Allocate to the second parallel device, and so on.

[0084] It should be noted that when calculating the circular attention of non-causal tasks, as Figure 6 shown, the number of tokens processed by each parallel device is the same, and there is no problem of uneven load between parallel devices. Therefore, for non-causal tasks, no pre-communication is required.

[0085] S250. Through at least two parallel devices, perform parallel processing on each processing sequence to determine the attention scores of each token in each processing sequence.

[0086] Among them, the attention score of each token is used to represent the correlation or importance between this token and other tokens in the sequence. Since other tokens are involved, when calculating the attention score, according to the different task types, the types of other tokens involved are different, thus distinguishing different attention score calculation methods.

[0087] Furthermore, if the task type is a causal task, determining the attention scores of each token in each processing sequence includes: calculating the attention scores of each token based on the query vector, key vector, and value vector of each token corresponding to the unidirectional attention of each processing sequence.

[0088] In a causal task, due to the use of the unidirectional attention mechanism, when calculating the attention score at each position, only the elements at and before this position can be focused on. Therefore, when calculating the attention score of a certain token, only this token and the tokens before it can be focused on. At this time, based on the query vector, key vector, and value vector of this token and the tokens before it, perform a correlation calculation to calculate the attention score of the current token. The calculation of the attention score can adopt the conventional attention score calculation method, and this embodiment will not elaborate here.

[0089] Furthermore, if the task type is a non-causal task, determining the attention scores of each token in each processing sequence includes: calculating the attention scores of each token based on the query vector, key vector, and value vector of each token corresponding to the bidirectional attention of each processing sequence.

[0090] In a non-causal task, due to the use of the bidirectional attention mechanism, when calculating the attention score at each position, all positions in the sequence can be focused on, including itself and all other elements in the circular structure. Therefore, for the current token, perform a correlation calculation based on the query vector, key vector, and value vector of all tokens from the start position to the end position in the sequence to obtain the attention score of the current token.

[0091] S260. Calculate circular attention based on the masks matching each processing sequence and the attention scores of each token in each processing sequence.

[0092] In this embodiment, applying the mask to the calculation process of circular attention enables the model to output circular attention without focusing on invalid positions when outputting circular attention, thereby reducing the computational complexity and improving the computational efficiency.

[0093] Furthermore, S260 can further include:

[0094] S261. Determine the attention weights of each token according to the masks matching each processing sequence and the attention scores of each token.

[0095] S262. Perform weighted summation on the attention weights of each token in each processing sequence according to the masks matching each processing sequence and the attention weights of each token to obtain circular attention.

[0096] Specifically, perform a softmax operation on the attention scores after applying the mask to obtain the attention weights of each token. The attention scores corresponding to the positions with a value of 0 in the mask (invalid padding positions) will be set to an extremely small value (usually negative infinity). After passing through the softmax function, the attention weights of these invalid padding positions will approach 0 and thus be ignored by the natural language model; while the positions with a value of 1 in the mask allow normal calculation of the attention weights, enabling the model to focus on the information corresponding to these positions.

[0097] After obtaining the attention weights corresponding to each position, integrate the information corresponding to each token according to its attention weight to obtain circular attention. Circular attention is an attention representation containing circular structure information, which comprehensively considers the degree of attention of each position in the sequence to other positions and the additional information flow brought by the circular structure.

[0098] The technical solution of this embodiment converts the input data into an input sequence, and splices the input sequences according to the order of the input sequences to obtain a total sequence. The total sequence is segmented by parallel parameters to obtain at least two processing sequences. According to different task types, different types of mask constructions are performed on each processing sequence respectively. Combining the mask with the sequence padding and segmentation methods can not only greatly reduce the occupation of computing resources and improve the processing efficiency of the input data, but also reduce the ineffective padding part and the waste of computing resources caused by processing the ineffective padding part. When the task type is a causal task, pre-communication is performed to ensure the load balance of each parallel device. Each parallel device performs parallel processing on each processing sequence, and circular attention is calculated based on the mask matching each processing sequence and the attention scores of each token in each processing sequence. The technical solution of this embodiment is more suitable for processing mixed input sequences of different lengths. Compared with the conventional circular attention calculation method, it saves more computing power, reduces computing power waste, and can ensure the load balance between parallel devices, making the data processing steps of the artificial intelligence model more efficient.

[0099] Embodiment III

[0100] Figure 7 FIG. is a schematic structural diagram of a data parallel processing device provided in Embodiment III of the present invention. As Figure 7 shown, the device includes:

[0101] A sequence splicing module 310, configured to determine at least two input sequences, and splice the input sequences according to the order of each input sequence to obtain a total sequence;

[0102] Wherein, each input sequence includes at least one token;

[0103] A sequence segmentation module 320, configured to obtain parallel parameters, and segment the total sequence based on the parallel parameters and the input sequence to obtain at least two processing sequences, where the parallel parameters are determined at least according to the number of parallel devices;

[0104] A sequence parallel processing module 330, configured to perform parallel processing on the processing sequences through at least two parallel devices to obtain circular attention.

[0105] The technical solution of the embodiment of the present invention converts the input data into an input sequence, splices the input sequence according to the order of the input sequence to obtain a total sequence, divides the total sequence by parallel parameters to obtain at least two processing sequences, and performs parallel processing on the processing sequences through at least two parallel devices to obtain circular attention. It solves the problem in the prior art that after filling other sequences based on the length of the longest sequence and then performing parallel processing on the mixed sequence, it wastes computing resources and reduces the processing speed. It reduces the filling amount of the input sequence, reduces the waste of computing resources, and improves the computing efficiency.

[0106] Based on the above embodiment, optionally, the sequence splitting module 320 includes:

[0107] A processing sequence length determination unit for determining the length of the processing sequence according to the total number of tokens in the total sequence and the parallel parameter;

[0108] A total sequence splitting unit for splitting the total sequence according to the length of the processing sequence to obtain at least two processing sequences.

[0109] Based on the above embodiment, optionally, the processing sequence length determination unit is specifically used for:

[0110] If the total number of tokens in the total sequence is an integer multiple of the parallel parameter, the value obtained by dividing the total number of tokens in the total sequence by the parallel parameter is used as the length of the processing sequence;

[0111] Otherwise, the length of the processing sequence is calculated by the following formula: len = ┌a┐, a = T / D, where len represents the length of the processing sequence, ┌ ┐ represents rounding up calculation, T represents the total number of tokens in the total sequence, and D represents the parallel parameter.

[0112] Based on the above embodiment, optionally, the total sequence splitting unit is specifically used for:

[0113] If the total number of tokens in the total sequence is an integer multiple of the parallel parameter, the total sequence is split according to the length of the processing sequence;

[0114] Otherwise, the total sequence is split according to the length of the processing sequence, and the end of the last processing sequence is padded so that the number of tokens in the last processing sequence is the same as the length of the processing sequence.

[0115] Based on the above embodiment, optionally, the sequence parallel processing module 330 includes:

[0116] A task type determination unit for obtaining a task type matching the input data, where the task type includes a causal task and a non-causal task;

[0117] A mask determination unit, configured to construct a mask that matches each processing sequence according to the task type and the input sequence;

[0118] An attention score determination unit, configured to perform parallel processing on each processing sequence through at least two parallel devices to determine the attention scores of each token in each processing sequence;

[0119] A circular attention determination unit, configured to calculate circular attention based on the mask that matches each processing sequence and the attention scores of each token in each processing sequence.

[0120] Based on the above embodiments, optionally, if the task type is a causal task, the mask determination unit is specifically configured to:

[0121] Construct a mask that matches each processing sequence based on a unidirectional attention mechanism;

[0122] The attention score determination unit is specifically configured to:

[0123] Calculate the attention scores of each token based on the query vector, key vector, and value vector of each token corresponding to the unidirectional attention of each processing sequence.

[0124] Based on the above embodiments, optionally, if the task type is a non-causal task, the mask determination unit is specifically configured to:

[0125] Construct a mask that matches each processing sequence based on a bidirectional attention mechanism;

[0126] The attention score determination unit is specifically configured to:

[0127] Calculate the attention scores of each token based on the query vector, key vector, and value vector of each token corresponding to the bidirectional attention of each processing sequence.

[0128] Based on the above embodiments, optionally, the circular attention determination unit is specifically configured to:

[0129] Determine the attention weights of each token according to the mask that matches each processing sequence and the attention scores of each token;

[0130] Perform weighted summation on the attention weights of each token in each processing sequence according to the mask that matches each processing sequence and the attention weights of each token to obtain circular attention.

[0131] Based on the above embodiments, optionally, the apparatus further includes:

[0132] A pre - communication module is used for pre - communicating with each parallel device. The pre - communication includes: allocating the query vectors, key vectors, and value vectors of each token required for processing causal tasks to each parallel device.

[0133] The data parallel processing device provided by the embodiments of the present invention can execute the data parallel processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0134] Embodiment Four

[0135] Figure 8 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0136] As Figure 8 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read - only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read - only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0137] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0138] Processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the data parallel processing method.

[0139] In some embodiments, the data parallel processing method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the data parallel processing method described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the data parallel processing method by any other suitable means (e.g., by means of firmware).

[0140] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0141] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0142] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0143] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0144] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0145] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0146] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0147] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data parallel processing method, characterized in that: include: Determine at least two input sequences according to at least one input data, and splice the input sequences according to the order of the input sequences to obtain a total sequence; Wherein, each input sequence includes at least one word unit; Acquire a parallel parameter, and divide the total sequence based on the parallel parameter to obtain at least two processing sequences, wherein the parallel parameter is determined at least according to the number of parallel devices; The processing sequence is processed in parallel by at least two parallel devices to obtain circular attention.

2. The method according to claim 1, characterized in that The total sequence is divided based on the parallel parameter to obtain at least two processing sequences, including: Determine the processing sequence length according to the total number of word units in the total sequence and the parallel parameter; The total sequence is divided according to the length of the processing sequence to obtain at least two processing sequences.

3. The method according to claim 2, characterized in that Determining the processing sequence length according to the total number of word units in the total sequence and the parallel parameter includes: If the total number of words in the total sequence is an integer multiple of the parallel parameter, the value obtained by dividing the total number of words in the total sequence by the parallel parameter is used as the processing sequence length; Otherwise, the processing sequence length is calculated by the following formula: len=┌a┐, a=T / D, where len represents the processing sequence length, ┌ ┐ represents rounding up, T represents the total number of tokens in the total sequence, and D represents the parallel parameter.

4. The method according to claim 2, characterized in that: According to the length of the processing sequence, the total sequence is divided to obtain at least two processing sequences, including: If the total number of word units in the total sequence is an integer multiple of the parallel parameter, the total sequence is segmented according to the length of the processing sequence; Otherwise, the total sequence is divided according to the length of the processing sequence, and the end of the last processing sequence is padded so that the number of tokens in the last processing sequence is consistent with the length of the processing sequence.

5. The method according to claim 1, characterized in that The processing sequence is processed in parallel by at least two parallel devices to obtain a circular attention, including: Acquire a task type that matches the input data, where the task type includes a causal task and a non-causal task; According to the task type and input sequence, a mask matching each processing sequence is constructed; Processing each processing sequence in parallel by at least two parallel devices to determine an attention score for each word in each processing sequence; Ring attention is calculated based on the mask matched to each processing sequence and the attention score of each token in each processing sequence.

6. The method according to claim 5, characterized in that If the task type is a causal task, a mask matching each processing sequence is constructed according to the task type, including: Based on the unidirectional attention mechanism, a mask matching each processing sequence is constructed; Determine the attention score for each token in each processing sequence, including: Based on the query vector, key vector, and value vector of each word corresponding to the unidirectional attention of each processing sequence, the attention score of each word is calculated.

7. The method according to claim 5, characterized in that If the task type is a non-causal task, a mask matching each processing sequence is constructed according to the task type, including: Based on the bidirectional attention mechanism, a mask matching each processing sequence is constructed; Determine the attention score for each token in each processing sequence, including: Based on the query vector, key vector, and value vector of each word corresponding to the bidirectional attention of each processing sequence, the attention score of each word is calculated.

8. The method according to any one of claims 5 to 7, characterized in that: Based on the mask matching each processing sequence and the attention score of each token in each processing sequence, the ring attention is calculated, including: Determine the attention weight of each word unit according to the mask matched with each processing sequence and the attention score of each word unit; According to the mask matching each processing sequence and the attention weight of each word unit, the attention weight of each word unit in each processing sequence is weighted and summed to obtain the ring attention.

9. The method according to claim 5, characterized in that If the task type is a causal task, before the processing sequences are processed in parallel by at least two parallel devices, the process further includes: Pre-communication is performed on each parallel device, and the pre-communication includes: distributing the query vector, key vector, and value vector of each word element required for processing the causal task to each parallel device.

10. A data parallel processing device, characterized in that: include: A sequence splicing module, used to determine at least two input sequences, and splice the input sequences according to the order of the input sequences to obtain a total sequence; Wherein, each input sequence includes at least one word unit; A sequence segmentation module, used for obtaining a parallel parameter, and segmenting the total sequence based on the parallel parameter and the input sequence to obtain at least two processing sequences, wherein the parallel parameter is determined at least according to the number of parallel devices; The sequence parallel processing module is used to process the processing sequence in parallel through at least two parallel devices to obtain circular attention.

Citation Information

Patent Citations

  • Sparse processing method and device of sparse attention network, and electronic equipment

    CN118520912A

  • Large-scale language model-oriented super-long text sequence processing method and device

    CN118821730A

  • Sequence processing method, electronic equipment and storage medium

    CN119005177A

  • Speculation decoding method and device based on large model, equipment and storage medium

    CN119806649A

  • Techniques for enabling bit-parallel wide string matching with a SIMD register

    US20140281371A1