Large model reasoning optimization method and device, storage medium and electronic equipment
By dividing the batch data set of the large model into microbatch data sets, and concurrently executing operators with different computing resources and network bandwidth resources, the time delay problem of large model inference tasks is solved and the execution efficiency is improved.
Patent Information
- Application Number
- CN202510963877.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-11
AI Technical Summary
The large model has a time delay problem when performing inference tasks, which limits its widespread use in application scenarios with high real-time requirements.
The batch data set is divided into at least two microbatch data sets, and different types of operators are executed concurrently, using different self-attention operators, feedforward neural network operators and Allreduce operators with different computing resources and network bandwidth resources to reduce the waiting time for calculation and communication.
By executing different types of operators concurrently, the hardware resource utilization rate of large models is improved, the inference delay is reduced, and the overall execution efficiency is improved.
Smart Images

Figure CN120471177A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of artificial intelligence technology, and in particular, to a large-model reasoning optimization method, device, storage medium, and electronic device. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, the field of natural language processing has achieved remarkable success. Among them, the Transformer architecture, based on the self-attention mechanism, has demonstrated remarkable results on a variety of tasks due to its superior ability to model long sequences, becoming a hot topic of research and application. Since its introduction, the Transformer architecture has undergone continuous evolution and innovation, gradually spawning three major variants. Among them, closed-source models such as the GPT series and open-source models such as Llama have gradually shown a convergence in architectural design, with a common tendency towards the decoder-only architecture. With its simple and efficient structure, the decoder-only architecture provides excellent support for language generation tasks and has become a key development direction for current language model architectures.
[0003] However, due to the large scale of large models, each inference requires a complete model inference process, which results in a high time delay when performing inference tasks through large models, thus limiting the widespread use of models in application scenarios with high real-time requirements.
[0004] Therefore, how to improve the efficiency of large models in performing reasoning tasks is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a large model reasoning optimization method is proposed, including: Obtaining a batch data set, wherein the batch data set includes input data corresponding to a plurality of tasks to be executed; Splitting the batch data set into at least two micro-batch data sets; Input data and operator description information corresponding to each to-be-executed task contained in each micro-batch data set are submitted to an inference device, so that when the inference device executes a first type of operator to process the input data corresponding to the to-be-executed task in one micro-batch data set, it concurrently executes a second type of operator to process the input data corresponding to the to-be-executed task in another micro-batch data set, and the first type of operator and the second type of operator require different hardware resources for execution.
[0006] According to a second aspect of one or more embodiments of this specification, a large model reasoning optimization device is proposed, including: An acquisition module is used to acquire a batch data set, wherein the batch data set contains input data corresponding to multiple tasks to be executed; A splitting module, configured to split the batch data set into at least two micro-batch data sets; An inference module is configured to submit input data and operator description information corresponding to each to-be-executed task contained in each micro-batch data set to an inference device, so that the inference device executes a first-type operator to process the input data corresponding to the to-be-executed task in one micro-batch data set while concurrently executing a second-type operator to process the input data corresponding to the to-be-executed task in another micro-batch data set, where the first-type operator and the second-type operator require different hardware resources for execution.
[0007] According to the third aspect of one or more embodiments of this specification, an electronic device is proposed, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the above-mentioned large model inference optimization method by running the executable instructions.
[0008] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is proposed, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the large model reasoning optimization method as described above are implemented.
[0009] According to a fifth aspect of one or more embodiments of this specification, a computer program product is proposed, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the large model reasoning optimization method as described above.
[0010] As can be seen from the above embodiments, this specification obtains a batch processing data set containing input data corresponding to multiple tasks to be executed, and divides the obtained batch processing data set into at least two micro-batch processing data sets, and then submits the input data and operator description information corresponding to each task to be executed contained in each micro-batch processing data set to the inference device, so that when the inference device executes a first type of operator to process the input data corresponding to the task to be executed in one micro-batch processing data set, it concurrently executes a second type of operator to process the input data corresponding to the task to be executed in another micro-batch processing data set, wherein the hardware resources required for the execution of the first type of operator and the second type of operator are different.
[0011] In this method, a batch data set can be divided into at least two micro-batch data sets, so that operators requiring communication resources and operators requiring computing resources contained in different micro-batch data sets can be executed concurrently. By masking each other, hardware resource utilization can be improved and inference latency can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a schematic diagram of a Decoder-Only model provided by an exemplary embodiment.
[0013] Figure 2 is a schematic diagram of a tensor parallel process provided by an exemplary embodiment.
[0014] Figure 3 It is a schematic diagram of a Transformer layer execution process provided by an exemplary embodiment.
[0015] Figure 4 It is a flowchart of a large model reasoning optimization method provided by an exemplary embodiment.
[0016] Figure 5 It is a schematic diagram of concurrent execution of multiple micro-batch data set calculation and communication processes provided by an exemplary embodiment.
[0017] Figure 6 This is a schematic diagram of the decomposition of the calculation process of the self-attention operator provided in an exemplary embodiment.
[0018] Figure 7A is a schematic diagram of an execution queue provided in an exemplary embodiment.
[0019] Figure 7B is a schematic diagram of an execution queue provided in another exemplary embodiment.
[0020] Figure 8 It is a schematic diagram of a method for adding an event object provided in an exemplary embodiment.
[0021] Figure 9 is a schematic diagram of a large model reasoning process provided in an exemplary embodiment.
[0022] Figure 10 It is a schematic structural diagram of a device provided by an exemplary embodiment.
[0023] Figure 11 It is a block diagram of a large model reasoning optimization device provided by an exemplary embodiment. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0026] With the rapid development of artificial intelligence (AI), large models using decoder-only architectures have achieved remarkable success in fields such as natural language processing and are widely used in tasks such as language generation, text summarization, and machine translation. However, these large models typically have billions or even tens of billions of parameters, which places significant pressure on their storage and computing resources, posing severe challenges to their actual inference deployment and application.
[0027] Among them, the reasoning process of the Decoder-Only model can be divided into two typical processes: prefill and decoding. Figure 1 shown.
[0028] Figure 1 is a schematic diagram of a Decoder-Only model provided by an exemplary embodiment.
[0029] Combine Figure 1As can be seen, during the prefill process, the Decoder-Only model first receives the complete input sequence. This input sequence can be natural language text, code snippets, or other data to be processed, depending on the model's application scenario. The model then uses an embedding layer and a self-attention layer (which can be a self-attention layer or a multi-head self-attention layer) to encode each token in the input sequence (i.e., the smallest semantic unit after the model segments the input text, which can be a word, subword (such as "un" + "happy"), character, or even punctuation). The model then calculates the association weights between each token and other tokens in the input sequence. This allows the model to understand the relationships and context between tokens in the input sequence.
[0030] After the self-attention calculation, a feed-forward neural network (FFN) performs further feature extraction and transformation on each token. After these calculations, the model generates the first output token based on the context of the input sequence. This token is generated based on a comprehensive understanding of the entire input sequence and marks the beginning of the output sequence.
[0031] In the above content, FFN usually contains two layers of linear transformation and activation function to increase the nonlinear expression ability of the model.
[0032] In the above content, in order to avoid problems such as gradient disappearance and training instability in the large model training process, model residual connections and normalization layers Add&Normalize layers can be set between the above-mentioned multi-head self-attention mechanism layer and the feedforward neural network layer, and between the feedforward neural network and the output layer to improve the model's convergence speed and generalization performance, thereby improving the overall model effect.
[0033] Furthermore, during the decoding process, the model takes the tokens generated in the previous iteration as input. These tokens are sequentially fed into the model to generate the next new token. The input tokens of the first iteration are the tokens generated in the prefill process.
[0034] To improve computational efficiency, the model leverages the KV (Key-Value) history stored in the KV cache during previous computations during decoding. This KV history contains the key and value information from the previously calculated self-attention mechanism. This eliminates the need to recalculate the attention vector matrix for all historical tokens when generating new tokens, saving computational resources and time.
[0035] For ease of understanding, the following takes the input sequence "Once upon a time, there was a scholar" as an example to explain the above process in detail. After the above input sequence is input into the large model, during the prefill process of the large model, calculations can be performed based on all tokens contained in the input sequence (that is, each unit in "Once upon a time, there was a scholar") to generate the first token of the output sequence, such as "he".
[0036] In the subsequent decoding process, the large model can run the model inference process again each time based on the token generated by the previous round of calculation and the KV historical records stored in the previous calculation process to obtain a new token, such as "bitter", until the inference is completed and the output sequence "Once upon a time, there was a scholar who studied poetry and books, hoping that one day he would be able to pass the imperial examination" is obtained.
[0037] During this process, researchers and engineers have tried various acceleration methods to reduce the inference latency of large models and improve system throughput. Among them, Tensor Parallelism (TP) technology can be used to split the model's weight matrix into multiple inference cards according to different dimensions. Each inference card only needs to store and calculate part of the weights, effectively reducing the memory pressure of a single card and improving computing efficiency. This accelerates the model inference process and, to a certain extent, alleviates the resource and latency issues of large-scale model inference. Figure 2 shown.
[0038] Figure 2 is a schematic diagram of a tensor parallel process provided by an exemplary embodiment.
[0039] Combine Figure 2 As can be seen from process a, when using TP technology to accelerate the inference process of a large model, the weight matrix of the feedforward neural network layer FFN (or multi-layer perceptron MLP) in each Transformer layer contained in the large model can be split into two parts according to different dimensions, and calculated on two inference cards respectively. After the calculation is completed, the communication process g is executed through the AllReduce operator to integrate the calculation results of the two inference cards to achieve data synchronization. Figure 2Process b is the process of dividing the weight matrix of the self-attention layer into two parts according to different dimensions and calculating it on two inference cards. It is similar to process a above and will not be repeated in this manual.
[0040] Combine Figure 1 and Figure 2 It can be seen that after the large model completes the calculation process of the self-attention layer and the feedforward neural network layer, there are still residual connection and normalization layer calculation processes. The normalization calculation process requires the calculation of the complete implicit representation (hidden_state) of a token. This requires that before the normalization calculation process, the shard calculation results of multiple inference cards must be aggregated. The most commonly used aggregation process is through Figure 2 The Allreduce operator in executes the communication process g to synchronize the shard calculation results on each inference card to other inference cards for integration to achieve data synchronization.
[0041] That is to say, in the process of executing reasoning tasks in a large model, each self-attention layer and feedforward neural network layer calculation process must be followed by an AllReduce communication, so that 2*layers AllReduce communications are required in one reasoning process.
[0042] The Allreduce operator is used to transfer and aggregate data between multiple inference cards. The latency calculation method can refer to the following formula:
[0043] In the above formula, start latency is the startup delay, link latency is the link delay, N is the number of communication cards, M is the amount of communication data, and BW is the bandwidth between cards. When the start latency and link latency are small, and the amount of data is small, the Allreduce latency is mainly limited by the link latency, while when the amount of data is large, the Allreduce performance is limited by the bandwidth between cards. This means that on machines without high-speed interconnection between cards, such as L20 and L40s, the Allreduce operation will become a serious performance bottleneck, significantly affecting the inference efficiency of the model. Figure 3 shown.
[0044] Figure 3 It is a schematic diagram of a Transformer layer execution process provided by an exemplary embodiment.
[0045] Combine Figure 3 It can be seen that during the execution of the Transformer layer of the large model, the execution time of the Allreduce communication process accounts for more than 63%, significantly affecting the reasoning efficiency of the model.
[0046] Based on this, this specification provides a large model reasoning optimization method. The following is a detailed description of the technical solutions provided by each embodiment of this specification in conjunction with the accompanying drawings.
[0047] Figure 4 This is a flowchart of a large model reasoning optimization method provided by an exemplary embodiment, including: S400: Acquire a batch data set, where the batch data set includes input data corresponding to a plurality of tasks to be executed.
[0048] S402: Split the batch data set into at least two micro-batch data sets.
[0049] Combine Figure 3 It can be seen that in the large model, the reasoning process of a Transformer layer mainly includes: self-attention operator, Allreduce operator 1, feedforward neural network operator, Allreduce operator 2, and there is a strict sequence dependency between these operators. It can be understood that the next operator can only start executing after the previous operator is completed.
[0050] Among them, the self-attention operator here may include the self-attention layer in the above content and the entire calculation process of the residual connection and normalization layer after the self-attention layer, and the feedforward neural network operator here may include the feedforward neural network layer in the above content and the entire calculation process of the residual connection and normalization layer after the feedforward neural network layer. In this specification, the self-attention layer in the above content and the entire calculation process of the residual connection and normalization layer after the self-attention layer, as well as the feedforward neural network layer in the above content and the entire calculation process of the residual connection and normalization layer after the feedforward neural network layer can be regarded as an "operator", so that different inference engines can be adapted to avoid optimization coupling for specific architectures.
[0051] In the above description, the large model can adopt a decoder-only structure. The self-attention operator and feedforward neural network operator in the large model mainly require computing resources when executing, and do not require network bandwidth resources. Allreduce operators 1 and 2 mainly require network bandwidth resources when executing, and require fewer computing resources.
[0052] Based on this, during large-model inference, the business platform can concurrently execute the self-attention operator and Allreduce operator 1 on the inference device, and concurrently execute the feedforward neural network operator and Allreduce operator 2 on the inference device, fully utilizing different hardware resources and effectively improving hardware resource utilization. Because the self-attention operator and feedforward neural network operator primarily require different hardware resources than Allreduce operators 1 and 2, concurrent execution avoids intense competition for hardware resources, effectively improving the efficiency of large-model inference tasks.
[0053] In the above context, inference devices can refer to hardware devices used for inference in artificial intelligence and deep learning tasks, such as inference cards, training cards, and graphics processing units (GPUs).
[0054] Furthermore, for a single pending task, the inference device executes the operators in the following order: self-attention operator, Allreduce operator 1, feedforward neural network operator, and Allreduce operator 2, showing a strict dependency relationship. However, in actual application scenarios, the input data corresponding to multiple pending tasks can be grouped into a batch data set for batch processing, and the operators used to process the input data corresponding to different pending tasks contained in the batch data set do not have the aforementioned dependency relationship.
[0055] Therefore, when the business platform executes the operator for processing the input data corresponding to the task to be executed through the inference device, it can simultaneously obtain the input data corresponding to multiple tasks to be executed to form a batch data set, and then the batch data set can be divided into at least two micro-batch data sets, so that when the inference device executes the self-attention operator or feedforward neural network operator for processing the input data corresponding to a task to be executed in one of the micro-batch data sets, it concurrently executes the Allreduce operator 1 or Allreduce operator 2 for processing the input data corresponding to a task to be executed in another micro-batch data set. In this way, the calculation process and the communication process can be carried out in parallel, so that the time required for the calculation process and the time required for the communication process cover each other, thereby improving the overall execution efficiency of the large model. Figure 5 shown.
[0056] Figure 5 It is a schematic diagram of concurrent execution of multiple micro-batch data set calculation and communication processes provided by an exemplary embodiment.
[0057] Combine Figure 5It can be seen that while the business platform executes the inference process for the pending task n in one micro-batch processing dataset through the inference device, it concurrently executes the inference process for the pending task m in another micro-batch processing dataset.
[0058] Among them, when executing the self-attention operator An of the task to be executed n, the Allreduce operator 2 of the previous task to be executed m-1 before the task to be executed m is executed concurrently, that is, Figure 4 Ar(m-1)2 in. When executing the Allreduce operator 1 of the task to be executed n, that is, Figure 5 Arn1 in the task m is executed concurrently with the self-attention operator Am. When executing the feedforward neural network operator FFNn of the task n, the Allreduce operator 1 of the task m is executed concurrently, that is, Figure 5 When executing the Allreduce operator 2 of the task n to be executed, that is, Figure 5 Arn2 in , concurrently executes the feedforward neural network operator FFNm of the task to be executed m, and so on.
[0059] In the above content, the overall delay can be calculated by referring to the following formula:
[0060] In the above formula, This is the overall delay. As can be seen here, the overall delay of the above process is the sum of the maximum of the delays of An and Ar(m-1)2, the maximum of the delays of Arn1 and Am, the maximum of the delays of FFNn and Arm1, and the maximum of the delays of Arn2 and FFNm.
[0061] During the original computation, all operators of the aforementioned pending tasks n and m are combined by the scheduler into a batch of data and executed sequentially. The overall latency can be calculated using the following formula:
[0062] It can be seen from the above formula that the above overall delay is the sum of the delay of An, the delay of Ar(m-1)2, the delay of Arn1, the delay of Am, the delay of FFNn, the delay of Arm1, the delay of Arn2 and the delay of FFNm.
[0063] From this, it can be seen that when the first type of operator is executed by the inference device to process the input data corresponding to the task to be executed in a micro-batch processing dataset, the second type of operator is concurrently executed to process the input data corresponding to the task to be executed in another micro-batch processing dataset. This can make the execution time of the operator that takes longer than the other type of operator mask the execution time of the other operator, thereby improving the efficiency of executing inference tasks on large models.
[0064] In the above description, the business platform can split a batch dataset into at least two micro-batch datasets by splitting the batch dataset into at least two micro-batch datasets based on a preset split granularity. The split granularity here includes the granularity of the pending task and the granularity of the input token. The following details the methods for splitting a batch dataset based on these two split granularities.
[0065] If the batch dataset needs to be split according to the granularity of the tasks to be executed, the input data corresponding to the tasks to be executed can be used as task units. The batch dataset can then be split into at least two micro-batch datasets based on the number of task units required in each micro-batch dataset. This means that the input data corresponding to each task to be executed is considered a task unit, and each task unit can be divided into micro-batch datasets based on actual needs.
[0066] In addition, if the batch dataset needs to be split according to the input token granularity, the input data corresponding to at least a portion of each pending task can be split based on the number of input tokens contained in the input data corresponding to each pending task, and the at least two input token sets obtained from the split can be used as task units (each input token set here contains at least one input token). The batch dataset can then be split into at least two micro-batch datasets based on the number of task units required to be included in each micro-batch dataset. This can be understood as using at least a portion of the input tokens contained in the input data corresponding to each pending task as task units, and then, based on actual needs, the entire input data corresponding to the pending task or a portion of the input data corresponding to the pending task can be divided into each micro-batch dataset.
[0067] It should be noted that the number of task units contained in each micro-batch dataset can be different. Of course, in order to maximize the efficiency of executing inference tasks on a large model, the number of task units contained in each micro-batch dataset can also be the same.
[0068] At this time, when the number of input tokens contained in each to-be-executed task is the same, the business platform can select the target splitting granularity based on the number of tasks to be executed and the expected number of micro-batch processing data sets, and split the batch processing data set into at least two micro-batch processing data sets based on the target splitting granularity.
[0069] Specifically, when the number of tasks to be executed is an integer multiple of the number of expected micro-batch processing data sets, the input data corresponding to each task to be executed can be evenly divided into each micro-batch processing data set.
[0070] At this time, the business platform can select the granularity of the tasks to be executed as the target splitting granularity when the number of tasks to be executed is an integer multiple of the number of expected micro-batch processing data sets. According to the target splitting granularity, the input data corresponding to the tasks to be executed is used as the task unit, and then the batch processing data set can be divided into at least two micro-batch processing data sets according to the number of task units required to be included in each micro-batch processing data set.
[0071] The number of task units required in each micro-batch dataset can be determined based on the number of tasks to be executed and the desired number of micro-batch datasets. For example, if the number of tasks to be executed is 10 and the number of micro-batch datasets is 2, the number of task units required in each micro-batch dataset can be 5.
[0072] Furthermore, when the number of tasks to be executed is not an integer multiple of the number of expected micro-batch processing data sets, the business platform can select the input token granularity as the target segmentation granularity, and then can segment the input data corresponding to at least part of the tasks to be executed in each task to be executed according to the target segmentation granularity and the number of each input token in the input data corresponding to each task to be executed, and use the input token set obtained by segmentation and the input data corresponding to the tasks to be executed that have not been segmented as task units. Finally, according to the number of task units required to be included in each micro-batch processing data set, the batch processing data set is divided into at least two micro-batch processing data sets.
[0073] The number of task units required to be included in each of the aforementioned micro-batch data sets may be determined based on the sum of the input token set obtained by segmentation and the number of unsegmented tasks to be executed, as well as the number of desired micro-batch data sets.
[0074] At this time, the business platform may further split at least part of the tasks to be executed according to the input token granularity.
[0075] For example, if the number of tasks to be executed is 5 and the number of micro-batch data sets is 2, then the input data of one of the tasks to be executed can be split into two input token sets, and then all the input tokens contained in the input data of each of the other tasks to be executed can be regarded as a task unit as a whole, and each input token set corresponding to the task to be executed after being split according to the input token granularity can be regarded as a task unit.
[0076] For another example, if there are three tasks to be executed and two micro-batch datasets, the input data for each task can be split into two input token sets. Each input token set can then be treated as a task unit.
[0077] At this time, the business platform can determine the number of task units required to be included in each micro-batch data set based on each input token set, the number of unsplit tasks to be executed, and the expected number of micro-batch data sets, and then split the batch data set into at least two micro-batch data sets based on the number of task units required to be included in each micro-batch data set.
[0078] In actual application scenarios, the number of input tokens contained in each task to be executed may also be different. In this case, the number of input tokens required to be processed when some tasks to be executed is much larger than the number of input tokens required to be processed when other tasks to be executed are executed. As a result, after the tasks to be executed are assigned to different micro-batch data sets according to the number of tasks to be executed, the workload corresponding to different micro-batch data sets may be unbalanced, resulting in the efficiency of large models executing inference tasks cannot be further improved.
[0079] Therefore, when the number of input tokens contained in each task to be executed is different, the business platform can also select a target splitting granularity based on the number of input tokens contained in each task to be executed and the expected number of micro-batch processing data sets, and split the batch processing data set into at least two micro-batch processing data sets according to the target splitting granularity.
[0080] Specifically, if the difference between the number of input tokens in the input data corresponding to each pending task satisfies a preset difference condition, the business platform may use the input token granularity as the target segmentation granularity, determine at least some of the pending tasks from the pending tasks as pending tasks, segment the input data corresponding to the pending tasks based on the target segmentation granularity and the number of input tokens in the input data corresponding to the pending tasks, and use the input token set obtained from the segmentation and the input data of other pending tasks other than the pending tasks as task units. Based on the number of task units required to be included in each micro-batch dataset, the batch dataset is segmented into at least two micro-batch datasets.
[0081] The above-mentioned difference conditions can be set according to actual needs. For example, when it is determined that the range between the number of input tokens in the input data corresponding to each task to be executed exceeds a preset range threshold, it can be considered that the above-mentioned difference conditions are met. At this time, the business platform can filter out the top N tasks to be executed with the largest number of input tokens in the corresponding input data from the tasks to be executed based on the number of input tokens in the input data corresponding to each task to be executed, as the selected tasks to be split. Alternatively, the tasks to be executed whose number of input tokens in the corresponding input data exceeds a preset number threshold can be filtered out from the tasks to be executed, as the selected tasks to be split.
[0082] For another example: when it is determined that the variance between the numbers of input tokens in the input data corresponding to each task to be executed exceeds the preset variance threshold, it can be regarded as meeting the above-mentioned difference condition. At this time, the business platform can filter out the top N tasks to be executed with the largest number of input tokens in the corresponding input data from the tasks to be executed based on the number of input tokens in the input data corresponding to each task to be executed, as the selected tasks to be split.
[0083] It should be noted that for the feedforward neural network operator and Allreduce operator in the large model, there is no dependency between multiple input tokens. At this time, dividing the input data of a single task to be executed into at least two input token sets has no effect on the calculation of the feedforward neural network operator and the Allreduce operator. For the self-attention operator, for each input token, it is necessary to consider the association relationship between the input token and other input tokens before the input token to calculate the attention weight corresponding to the input token, such as Figure 6 shown.
[0084] Figure 6This is a schematic diagram of the decomposition of the calculation process of the self-attention operator provided in an exemplary embodiment.
[0085] Combine Figure 6 As can be seen, during the prefill process of the large model and the subsequent decoding process, when calculating the self-attention operator, the query vector matrix (Q matrix), key vector matrix (K matrix), and value vector matrix (V matrix) required for the self-attention operator calculation can be decomposed into two parts. Assuming the input token sequence is of length n, and decomposition is performed starting at position m, each vector in the vector matrix has dimension d. The first decomposition uses Q1[m, d], K1[m, d], and V1[m, d], while the second decomposition uses Q2[nm, d], K2[n, d], and V2[n, d]. These two parts can then be divided into different micro-batch datasets for concurrent calculation, resulting in A1[m, d] and A2[nm, d]. After this, A1[m, d] and A2[nm, d] can be merged using the Allreduce operator to obtain the final A[n, d].
[0086] It should be noted that the input data corresponding to the task to be executed may include: input sequence data, attention metadata attention_metadata, wherein the input sequence data includes: input identifier sequence input_ids, position identifier sequence position_ids, where each input identifier in the input identifier sequence is used to represent the vocabulary index sequence corresponding to each input token, and each position identifier in the position identifier sequence is used to represent the position of each input token in the entire input token sequence. The attention metadata here is used to represent the position of at least part of the attention tensor matrix that the inference device needs to use when executing the attention operator in the entire attention tensor matrix. The attention tensor matrix here is the Q matrix, K matrix, and V matrix required for self-attention calculation in the above content.
[0087] , 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120,121, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120,121, 109, 110, 122], where each value contained in the input identifier sequence is the unique identifier corresponding to each input token contained in the input text in the vocabulary, and the corresponding position identifier sequence position_ids is position_ids = [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17,18, 19, 20, 21, 22, 23], where each value contained in the position identifier sequence is the position of each input token contained in the input text in the entire input token sequence.
[0088] The above-mentioned attention metadata may include: query_lens: the length of the input sequence corresponding to the task to be executed for generating the query query vector, query_start_loc: the starting position offset of the input sequence corresponding to the task to be executed for generating the query query vector in the entire token sequence, seq_lens: the length of the entire input token sequence of the task to be executed, seq_start_loc: the starting position offset of the entire input token sequence of the task to be executed in the sequence obtained by splicing the input token sequences of all tasks to be executed contained in the batch data set.
[0089] As can be seen from the above content, the above two types of input sequence data are used to represent each input token contained in the input token sequence from different dimensions. By directly splitting the above two types of input sequence data, the input token sequence is split to obtain the split task units.
[0090] The above-mentioned attention metadata is used to mark the position of the attention tensor required when executing the attention operator in the entire attention tensor, so that when executing the large model self-attention operator, at least part of the attention tensor matrix in the entire attention tensor matrix can be obtained from the KV cache according to the attention metadata as the split attention tensor matrix, and the corresponding self-attention calculation can be performed.
[0091] Based on this, the segmentation position information corresponding to at least part of the tasks to be executed in each task to be executed can be determined according to the target segmentation granularity and the number of input tokens in the input data corresponding to each task to be executed. Then, the attention metadata corresponding to each input token set obtained by segmentation can be re-determined according to the segmentation position information and the original attention metadata corresponding to at least part of the tasks to be executed. Here, for each input token set, the attention metadata corresponding to the input token set is used to mark the position in the entire attention tensor matrix of at least part of the attention tensor matrix that needs to be used when performing self-attention calculation for the input token set.
[0092] As can be seen from the above content, the business platform can continue to split the input sequence data and attention metadata in the input data corresponding to each task to be executed contained in the batch data set according to the above method, so as to split the batch data set into two micro-batch data sets, so that the inference device can concurrently process the task units contained in the two micro-batch data sets.
[0093] It should be noted that during the Prifill process of the large model and each Decoder iteration, a new token will be generated as the input token for the next iteration. Therefore, the business platform can re-split the batch data set using the above method at the beginning of each iteration and submit the input data and operator description information corresponding to each task to be executed contained in each micro-batch data set during this iteration to the inference device for concurrent execution.
[0094] In this specification, the execution entity used to implement the large model reasoning optimization method can refer to a designated device such as a server set up in the business platform, or it can refer to a terminal device such as a desktop computer, a laptop computer, etc. For the sake of convenience of description, the large model reasoning optimization method provided in this specification is explained below using the control device as an example of the execution entity.
[0095] S404: Submitting input data and operator description information corresponding to each to-be-executed task contained in each micro-batch data set to an inference device, so that the inference device executes a first-type operator to process the input data corresponding to the to-be-executed task in one micro-batch data set while concurrently executing a second-type operator to process the input data corresponding to the to-be-executed task in another micro-batch data set, wherein the first-type operator and the second-type operator require different hardware resources for execution.
[0096] In this specification, the control device may submit the input data and operator description information corresponding to each task to be executed contained in each micro-batch data set to the inference device, so that the inference device executes a first type of operator to process the input data corresponding to the task to be executed in one micro-batch data set while concurrently executing a second type of operator to process the input data corresponding to the task to be executed in another micro-batch data set.
[0097] In the above content, the hardware resources required for the execution of the first type of operators and the second type of operators are different.
[0098] The operator description information corresponding to the task to be executed may be information required for calling the program code of the operator. The operator here may refer to an operator required for processing the input data corresponding to the task to be executed.
[0099] It should be noted that when performing inference on two micro-batch datasets simultaneously on the same control device, operators within the same execution queue stream will be executed sequentially in the order in which they were submitted to the execution queue, while operators within different execution queues can be executed concurrently in any relative order. Therefore, the control device can submit each task unit contained in each micro-batch dataset and the operator description information required to process each task unit to each preset execution queue, allowing the inference device to alternately obtain the operator description information and task units required to execute different types of operators from each execution queue, thereby concurrently executing different types of operators.
[0100] Specifically, the control device can submit each task unit contained in each micro-batch data set and the operator description information required to process each task unit to the preset execution queues in two ways: Figure 7A 、 Figure 7B As shown, the following combination Figure 7A 、 Figure 7B The two methods are described in detail respectively.
[0101] Figure 7A is a schematic diagram of an execution queue provided in an exemplary embodiment.
[0102] exist Figure 7AIn this example, the control device can, for each micro-batch dataset, submit the operators required to process the task units corresponding to the tasks to be executed contained in the micro-batch dataset in sequence to the execution queue corresponding to the micro-batch dataset according to the order of the task units corresponding to the tasks to be executed contained in the micro-batch dataset. This allows the inference device to alternately obtain the operator description information and task units required to execute different types of operators from each execution queue according to the order of the operators in the execution queue, thereby concurrently executing different types of operators. Each execution queue here corresponds one-to-one to each micro-batch dataset.
[0103] Figure 7B is a schematic diagram of an execution queue provided in another exemplary embodiment.
[0104] exist Figure 7B In this example, the control device can, for each micro-batch dataset, submit different types of operators required to process each task unit corresponding to the tasks to be executed contained in the micro-batch dataset to different execution queues in sequence according to the order of the task units corresponding to the tasks to be executed contained in the micro-batch dataset. This allows the inference device to alternately obtain the operator description information and task units required to execute different types of operators from each execution queue according to the order of the operators in the execution queues, thereby concurrently executing different types of operators. Each execution queue here corresponds one-to-one to each type of operator.
[0105] In actual application scenarios, since different task units may have different sizes, and the time required for the inference device to execute operators for task units of different sizes is also different, the inference device obtains the operator description information and task units required to execute different types of operators from each execution queue according to the order of each operator in the execution queue. In the process of concurrently executing different types of operators, the same type of operators in different execution queues may be executed simultaneously, which leads to competition for hardware resources and thus reduces the efficiency of large models in executing inference tasks.
[0106] For example, if the inference device takes too long to execute the first-type operator a obtained from execution queue A, the second-type operator b obtained from execution queue B, which is executed concurrently with the operator, may have already been executed. If the inference device obtains the first-type operator c from execution queue B again, it will compete with the first-type operator a being executed for hardware resources.
[0107] Based on this, the control device can also adjust the order of each task unit contained in each micro-batch data set according to the size of each task unit contained in the micro-batch data set to obtain an adjusted micro-batch data set, and submit each task unit contained in each adjusted micro-batch data set and the operator description information required to process each task unit to the inference device.
[0108] As can be seen from the above content, the control device can adjust the order of each task unit contained in each micro-batch data set before submitting each task unit contained in each adjusted micro-batch data set and the operator description information required to process each task unit to the inference device, so that the task units in the same position in each micro-batch data set have similar sizes, thereby avoiding the above-mentioned problem of hardware resource competition.
[0109] For example, if the size of task unit a in micro-batch dataset A is 500, the size of task unit b is 200, and the size of task unit c is 400, and the size of task unit d in micro-batch dataset B is 400, the size of task unit e is 200, and the size of task unit f is 500, then the control device may adjust the order of the task units in micro-batch dataset A to task unit b, task unit c, and task unit a based on the sizes of the task units in micro-batch dataset A, and adjust the order of the task units in micro-batch dataset B to task unit e, task unit d, and task unit f based on the sizes of the task units in micro-batch dataset B, so as to avoid the above-mentioned problem of hardware resource competition.
[0110] In addition, when submitting each task unit contained in each micro-batch data set and the operator description information required to process each task unit to the preset execution queues, the control device can also add the event object required for the execution of the operator processing each task unit, so that the inference device can alternately obtain the operator description information and task units required to execute different types of operators from each execution queue to execute different types of operators concurrently.
[0111] The above event objects are used to record the execution status of operators and adjust the execution timing of operators in different execution queues. For example, the event.record event object is used to record an event at a specified position in the current execution queue. The event will be marked as the completion of the operator execution when the execution queue reaches this position. The event.wait event is used to make subsequent operators wait for the specified event to complete before continuing to execute. For ease of understanding, the following is combined Figure 8The above process of fine-grained control of each execution queue through event objects is described in detail.
[0112] Figure 8 It is a schematic diagram of a method for adding an event object provided in an exemplary embodiment.
[0113] exist Figure 8 Each solid arrow in the figure represents an event.record event object, and each dotted arrow represents an event.wait event object. Figure 8 It can be seen that the control device can add an event.record event object after each operator to record the completion of the execution of each operator, and can add an event.wait event object before each operator to control the operator to wait for the completion of the execution of the operator in other execution queues that is executed concurrently with the previous operator of the operator before the operator can be executed.
[0114] For example: Figure 8 As shown in the figure, the An operator of an execution queue will trigger the event.record event object after execution is completed to mark the completion of the execution of the An operator. At this time, due to the event.wait restriction added before the Arn1 operator of another execution queue, the Arn1 operator in the other execution queue will not start execution until it is determined that the An operator has completed execution. Moreover, since the operators in the same execution queue will be executed in sequence strictly according to the order of each operator in the execution queue, when the An operator is executed, the inference device can continue to obtain the Am operator from this execution queue and execute it concurrently with the Arn1 operator that has just started execution due to the event.wait restriction, and so on.
[0115] For ease of understanding, the following describes in detail the reasoning process of the large model in the process of executing each task to be executed in the batch data set according to the above large model reasoning optimization method. Figure 9 shown.
[0116] Figure 9 is a schematic diagram of a large model reasoning process provided in an exemplary embodiment.
[0117] Combine Figure 9 It can be seen that after obtaining the batch data set, the control device can split the batch data set into two micro-batch data sets, and then submit the task units contained in each micro-batch data set to the corresponding execution queue in sequence, so that the inference device can alternately obtain different types of operators from different execution queues for concurrent execution, thereby improving the efficiency of large models in executing inference tasks.
[0118] It should be noted that the aforementioned model inference optimization method splits the batch data set input to the large model into two micro-processing data sets and submits them to two execution queues, allowing a single inference device to concurrently execute the pending tasks in these two micro-processing data sets. The tensor parallel method, on the other hand, splits the parameters of the large model (i.e., the weight parameters of the self-attention layer and the weight parameters of the feedforward neural network layer) into multiple groups, each of which is computed by a different inference device. In other words, the aforementioned model inference optimization method and tensor parallel method split the computational operations during the execution of inference tasks on the large model along two different dimensions (i.e., splitting the input data of the large model and the parameters of the large model, respectively).
[0119] Therefore, the control device can also split the parameters of the large model into different parameter sets, and send each parameter set to a different inference device, and then, for each inference device, submit the input data and operator description information corresponding to each task to be executed contained in each micro-batch processing data set to the inference device, so that the inference device executes the first type of operator to process the input data corresponding to the task to be executed in a micro-batch processing data set while concurrently executing the second type of operator to process the input data corresponding to the task to be executed in another micro-batch processing data set, so as to obtain the sub-execution result corresponding to each task to be executed returned by the inference device, and then, based on the sub-execution result corresponding to each task to be executed returned by each inference device, the execution result corresponding to each task to be executed can be obtained.
[0120] It should be noted that the aforementioned tasks to be performed can be determined based on actual needs, such as image processing, text generation, etc., and this specification does not impose any restrictions thereon. Specifically, if the aforementioned tasks to be performed are image processing tasks, then the input data corresponding to the aforementioned tasks to be performed can be the image data to be processed. If the aforementioned tasks to be performed are text generation tasks, then the input data corresponding to the aforementioned tasks to be performed can be the question text data to be processed.
[0121] As can be seen from the above content, the control device can split the batch data set into at least two micro-batch data sets to perform communication and computation concurrently between the micro-batch data sets, thereby improving hardware resource utilization and reducing inference latency by masking the two. In addition, based on the characteristics of the model, it supports balanced splitting of micro-batch data sets at the granularity of the task to be executed and the granularity of the input token, and adopts a specially designed balanced splitting algorithm to ensure load balancing between each micro-batch data set, further improving resource utilization efficiency. More importantly, the above method can only control the concurrency of the communication process and the computation process of the large model without modifying the original large model structure, thereby maintaining good compatibility with various large models based on the transformer architecture, so as to improve the efficiency of large model inference, reduce latency, and have the versatility to be widely applicable to a variety of model architectures.
[0122] Figure 10 This is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 10 At the hardware level, the device includes a processor 1002, an internal bus 1004, a network interface 1006, a memory 1008, and a non-volatile memory 1010. Of course, it may also include hardware required for other functions. One or more embodiments of this specification can be implemented based on software, such as the processor 1002 reading the corresponding computer program from the non-volatile memory 1010 into the memory 1008 and then running it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0123] Please refer to Figure 11 , the large model inference optimization device can be applied to Figure 11 The device shown in the figure is used to implement the technical solution of this specification. The large model reasoning optimization device may include: An acquisition module 1101 is configured to acquire a batch data set, wherein the batch data set includes input data corresponding to a plurality of tasks to be executed; A splitting module 1102 is configured to split the batch data set into at least two micro-batch data sets; The inference module 1103 is configured to submit input data and operator description information corresponding to each to-be-executed task contained in each micro-batch data set to an inference device, so that the inference device executes a first-type operator to process the input data corresponding to the to-be-executed task in one micro-batch data set while concurrently executing a second-type operator to process the input data corresponding to the to-be-executed task in another micro-batch data set, where the first-type operator and the second-type operator require different hardware resources for execution.
[0124] Optionally, the splitting module 1102 is specifically configured to split the batch data set into at least two micro-batch data sets according to a preset splitting granularity; the splitting granularity includes: a to-be-executed task granularity and an input token granularity.
[0125] Optionally, each micro-batch data set includes the same number of task units; each task unit includes at least part of the input data corresponding to the task to be executed.
[0126] Optionally, the splitting module 1102 is specifically configured to, when the number of input tokens in the input data corresponding to each task to be executed is the same, select a target splitting granularity according to the number of tasks to be executed and the expected number of micro-batch data sets, and split the batch data set into at least two micro-batch data sets according to the target splitting granularity; when the number of input tokens in the input data corresponding to each task to be executed is different, select a target splitting granularity according to the number of input tokens in the input data corresponding to each task to be executed and the expected number of micro-batch data sets, and split the batch data set into at least two micro-batch data sets according to the target splitting granularity.
[0127] Optionally, the splitting module 1102 is specifically configured to, when the number of tasks to be executed is an integer multiple of the number of desired micro-batch data sets, select the granularity of the tasks to be executed as the target splitting granularity; based on the target splitting granularity, use the input data corresponding to the tasks to be executed as task units; and split the batch data set into at least two micro-batch data sets based on the number of task units required to be included in each micro-batch data set.
[0128] Optionally, the splitting module 1102 is specifically configured to, when the number of tasks to be executed is not an integer multiple of the number of desired micro-batch data sets, select the input token granularity as the target splitting granularity; split the input data corresponding to at least part of the tasks to be executed in each task to be executed according to the target splitting granularity and the number of input tokens in the input data corresponding to each task to be executed, and use the input token sets obtained by the splitting as task units; and split the batch data set into at least two micro-batch data sets according to the number of task units required to be included in each micro-batch data set.
[0129] Optionally, the splitting module 1102 is specifically configured to, when a difference between the numbers of input tokens in the input data corresponding to each task to be executed meets a preset difference condition, use the input token granularity as a target splitting granularity, and determine at least part of the tasks to be executed from the tasks to be executed as tasks to be split; split the input data corresponding to the tasks to be split according to the target splitting granularity and the number of input tokens in the input data corresponding to the tasks to be split, and use the input token set obtained by the splitting and the input data of other tasks to be executed except the tasks to be split as task units; and split the batch data set into at least two micro-batch data sets according to the number of task units required to be included in each micro-batch data set.
[0130] Optionally, the input data includes: input sequence data, attention metadata; The segmentation module 1102 is specifically configured to segment the input sequence data corresponding to at least part of the tasks to be executed according to the target segmentation granularity and the number of input tokens in the input data corresponding to each task to be executed, and use the input token sets obtained by segmentation as task units; and Determining segmentation position information corresponding to at least part of the tasks to be executed according to the target segmentation granularity and the number of input tokens in the input data corresponding to each task to be executed; Based on the segmentation position information and the original attention metadata corresponding to at least part of the task to be performed, the attention metadata corresponding to each input token set obtained by segmentation is re-determined.
[0131] Optionally, the inference module 1103 is specifically configured to submit each task unit included in each micro-batch data set and operator description information required for processing each task unit to an inference device.
[0132] Optionally, the reasoning module 1103 is specifically configured to, for each micro-batch data set, adjust the order of each task unit included in the micro-batch data set according to the size of each task unit included in the micro-batch data set to obtain an adjusted micro-batch data set; Each task unit contained in each adjusted micro-batch data set and the operator description information required to process each task unit are submitted to the inference device.
[0133] Optionally, the inference module 1103 is specifically configured to submit each task unit contained in each micro-batch data set and operator description information required for processing each task unit to preset execution queues, so that the inference device alternately obtains operator description information and task units required for executing different types of operators from each execution queue to concurrently execute different types of operators.
[0134] Optionally, the inference module 1103 is specifically configured to submit each task unit contained in each micro-batch data set and the operator description information required for processing each task unit to each preset execution queue, and add an event object required for executing the operator that processes each task unit, so that the inference device alternately obtains the operator description information and task units required for executing different types of operators from each execution queue to concurrently execute different types of operators; the event object is used to record the execution status of the operator and adjust the execution timing of the operators in different execution queues.
[0135] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the method described in any of the above embodiments by running the executable instructions.
[0136] Based on the same concept as the above method, this specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0137] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any of the above embodiments when executed by a processor.
Claims
1. A large model inference optimization method, comprising: Obtaining a batch data set, wherein the batch data set includes input data corresponding to a plurality of tasks to be executed; Splitting the batch data set into at least two micro-batch data sets; Input data and operator description information corresponding to each to-be-executed task contained in each micro-batch data set are submitted to an inference device, so that when the inference device executes a first type of operator to process the input data corresponding to the to-be-executed task in one micro-batch data set, it concurrently executes a second type of operator to process the input data corresponding to the to-be-executed task in another micro-batch data set, and the first type of operator and the second type of operator require different hardware resources for execution.
2. The method according to claim 1, wherein the batch data set is divided into at least two micro-batch data sets, specifically comprising: Splitting the batch data set into at least two micro-batch data sets according to a preset splitting granularity; The segmentation granularity includes: task to be executed granularity and input token granularity.
3. The method of claim 2 , wherein each micro-batch data set comprises the same number of task units; each task unit comprises at least a portion of input data corresponding to the task to be executed.
4. The method according to claim 3, wherein the batch data set is divided into at least two micro-batch data sets according to a preset segmentation granularity, specifically comprising: When the number of input tokens in the input data corresponding to each task to be executed is the same, a target splitting granularity is selected according to the number of tasks to be executed and the expected number of micro-batch data sets, and the batch data set is split into at least two micro-batch data sets according to the target splitting granularity; When the number of input tokens in the input data corresponding to each task to be executed is different, a target splitting granularity is selected based on the number of input tokens in the input data corresponding to each task to be executed and the expected number of micro-batch data sets, and the batch data set is split into at least two micro-batch data sets according to the target splitting granularity.
5. The method according to claim 4, further comprising selecting a target splitting granularity based on the number of tasks to be executed and the desired number of micro-batch datasets, and splitting the batch dataset into at least two micro-batch datasets based on the target splitting granularity, comprising: When the number of tasks to be executed is an integer multiple of the number of desired micro-batch data sets, the granularity of the tasks to be executed is selected as the target segmentation granularity; According to the target segmentation granularity, the input data corresponding to the task to be executed is used as a task unit; According to the number of task units required to be included in each micro-batch data set, the batch data set is divided into at least two micro-batch data sets.
6. The method of claim 4, further comprising: selecting a target splitting granularity based on the number of tasks to be executed and the desired number of micro-batch datasets; and splitting the batch dataset into at least two micro-batch datasets based on the target splitting granularity, specifically comprising: When the number of tasks to be executed is not an integer multiple of the number of desired micro-batch data sets, selecting the input token granularity as the target segmentation granularity; According to the target segmentation granularity and the number of input tokens in the input data corresponding to each task to be executed, the input data corresponding to at least part of each task to be executed is segmented, and the input token sets obtained by segmentation are used as task units; According to the number of task units required to be included in each micro-batch data set, the batch data set is divided into at least two micro-batch data sets.
7. The method of claim 4, further comprising: selecting a target splitting granularity based on the number of input tokens in the input data corresponding to each task to be executed and the desired number of micro-batch datasets; and splitting the batch dataset into at least two micro-batch datasets based on the target splitting granularity, specifically comprising: When a difference between the numbers of input tokens in the input data corresponding to each task to be executed meets a preset difference condition, the input token granularity is used as the target segmentation granularity, and at least part of the tasks to be executed are determined from the tasks to be executed as tasks to be segmented; Split the input data corresponding to the task to be split according to the target splitting granularity and the number of input tokens in the input data corresponding to the task to be split, and use the input token set obtained by the splitting and the input data of other tasks to be executed except the task to be split as task units; According to the number of task units required to be included in each micro-batch data set, the batch data set is divided into at least two micro-batch data sets.
8. The method of claim 6, wherein the input data comprises: Input sequence data, attention metadata; According to the target segmentation granularity and the number of input tokens in the input data corresponding to each task to be executed, the input data corresponding to at least part of each task to be executed is segmented, and the input token sets obtained by segmentation are used as task units, specifically including: According to the target segmentation granularity and the number of input tokens in the input data corresponding to each task to be executed, the input sequence data corresponding to at least part of the tasks to be executed are segmented, and the input token sets obtained by the segmentation are used as task units; and Determining segmentation position information corresponding to at least part of the tasks to be executed according to the target segmentation granularity and the number of input tokens in the input data corresponding to each task to be executed; Based on the segmentation position information and the original attention metadata corresponding to at least part of the task to be performed, the attention metadata corresponding to each input token set obtained by segmentation is re-determined.
9. The method according to any one of claims 5 to 7, wherein the input data and operator description information corresponding to each to-be-executed task contained in each micro-batch data set are submitted to the inference device, specifically comprising: Submit each task unit contained in each micro-batch data set and the operator description information required to process each task unit to the inference device.
10. The method of claim 9, wherein each task unit contained in each micro-batch data set and operator description information required for processing each task unit are submitted to the inference device, specifically comprising: For each micro-batch processing data set, adjusting the order of each task unit contained in the micro-batch processing data set according to the size of each task unit contained in the micro-batch processing data set to obtain an adjusted micro-batch processing data set; Each task unit contained in each adjusted micro-batch data set and the operator description information required to process each task unit are submitted to the inference device.
11. The method of claim 9, wherein submitting each task unit contained in each micro-batch data set and operator description information required for processing each task unit to the inference device specifically comprises: Each task unit contained in each micro-batch data set and the operator description information required to process each task unit are submitted to each preset execution queue, so that the inference device alternately obtains the operator description information and task units required to execute different types of operators from each execution queue to concurrently execute different types of operators.
12. The method of claim 9, comprising submitting each task unit and operator description information required for processing each task unit contained in each micro-batch data set to each preset execution queue, so that the inference device alternately obtains operator description information and task units required for executing different types of operators from each execution queue to concurrently execute different types of operators, specifically comprising: Each task unit contained in each micro-batch data set and the operator description information required to process each task unit are submitted to each preset execution queue, and an event object required for the execution of the operator processing each task unit is added, so that the inference device alternately obtains the operator description information and task units required to execute different types of operators from each execution queue to concurrently execute different types of operators; the event object is used to record the execution status of the operator and adjust the execution timing of the operators in different execution queues.
13. A large model reasoning optimization device, comprising: An acquisition module is used to acquire a batch data set, wherein the batch data set contains input data corresponding to multiple tasks to be executed; A splitting module, configured to split the batch data set into at least two micro-batch data sets; An inference module is configured to submit input data and operator description information corresponding to each to-be-executed task contained in each micro-batch data set to an inference device, so that the inference device executes a first-type operator to process the input data corresponding to the to-be-executed task in one micro-batch data set while concurrently executing a second-type operator to process the input data corresponding to the to-be-executed task in another micro-batch data set, where the first-type operator and the second-type operator require different hardware resources for execution.
14. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 12 by executing the executable instructions.
15. A computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Real-time statistical method, device and equipment based on micro-batch processing and storage medium
CN117149855A
Deep learning reasoning acceleration method based on DNN operator parallelism
CN117196037A
Inference method and device based on multi-batch processing splitting, equipment and medium
CN118536594A
Distributed training micro-batch data determination method and device, equipment and medium
CN118709752A
Parallel training and reasoning adaptation optimization method for new-generation heterogeneous supercomputing large model
CN119783812A
Cited By
Large model reasoning control method and device, equipment and medium
CN121581247A