Task execution method, electronic equipment, medium and program product

By dividing the inference task into batches and synchronously performing post-processing tasks during the calculation process, the problem of waste of computing resources is solved, and the computing speed and efficiency of the language model are improved.

CN120276833AActive Publication Date: 2025-07-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510773133.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

When using language models to perform semantic inference calculations, post-processing calculations need to be performed after the inference calculations are completed, resulting in waste of computing resources and inefficient efficiency.

Method used

Multiple inference tasks are divided into multiple batches, and post-processing sub-tasks are synchronized during the calculation sub-task execution by the calculation unit to improve the utilization rate of computing resources.

Benefits of technology

Without reducing model performance, the computing speed and inference efficiency are improved, especially when processing large-scale data and complex models, the resource utilization is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276833A_ABST
    Figure CN120276833A_ABST
Patent Text Reader

Abstract

The invention provides a task execution method, electronic equipment, a medium and a program product, and relates to the technical field of artificial intelligence. The task execution method comprises the steps that multiple reasoning tasks generated based on the multiple reasoning requests are divided into N batches in response to the received multiple reasoning requests, and each reasoning task comprises a calculation sub-task and a post-processing sub-task to be executed in sequence; the N batches of reasoning tasks are sequentially executed, a task execution result is obtained, and the step of executing the nth batch of reasoning tasks comprises the substeps that the nth batch of calculation subtasks are sent to the calculation unit; and under the condition that the calculation unit executes the nth batch of calculation subtasks, sending the (n + 1) th batch of calculation subtasks to the calculation unit, and synchronously executing the nth batch of post-processing subtasks in the process of executing the (n + 1) th batch of calculation subtasks by the calculation unit. The task execution method can improve the utilization rate of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a task execution method, apparatus, electronic device, medium, and program product. Background Art

[0002] With its powerful computing power and massive data training, language models have shown extensive application potential in multiple fields. When performing semantic reasoning calculations using language models, multiple calculators can be used to collaborate on computing tasks. For example, two calculators can be used to jointly execute a computing task, where one calculator performs the main reasoning calculation and the other calculator performs post-processing calculations.

[0003] In the process of implementing the present invention, it is found that since post-processing calculations need to be executed after the reasoning calculations are completed, one execution method is as follows: after receiving a batch of reasoning tasks, first, one calculator executes the reasoning calculations for all tasks, and then another calculator executes the post-processing calculations for all tasks. Therefore, during the execution of post-processing calculations, one of the calculators is idle and waiting, resulting in waste of computing resources. Summary of the Invention

[0004] In view of the above problems, the present invention provides a task execution method, apparatus, electronic device, medium, and program product.

[0005] According to a first aspect of the present invention, a task execution method is provided. The method includes: in response to receiving a plurality of reasoning requests, dividing a plurality of reasoning tasks generated based on the plurality of reasoning requests into N batches, where each reasoning task includes a calculation subtask and a post-processing subtask to be executed in sequence, and N is an integer greater than 2; sequentially executing the reasoning tasks in N batches to obtain a task execution result, where executing the reasoning tasks in the nth batch includes: sending the nth batch of calculation subtasks to a calculation unit; in the case where the calculation unit has executed the nth batch of calculation subtasks, sending the (n + 1)th batch of calculation subtasks to the calculation unit and, during the process of the calculation unit executing the (n + 1)th batch of calculation subtasks, synchronously executing the nth batch of post-processing subtasks, where n = 1,..., N - 2.

[0006] The second aspect of the present invention provides a task execution device, including: a batching module, configured to divide a plurality of inference tasks generated based on a plurality of inference requests into N batches in response to receiving the plurality of inference requests, where each inference task includes a computing subtask and a post-processing subtask to be executed in sequence, and N is an integer greater than 2; an execution module, configured to sequentially execute the inference tasks of N batches to obtain a task execution result, where executing the inference tasks of the nth batch includes: sending the computing subtasks of the nth batch to a computing unit; in the case where the computing unit has executed the computing subtasks of the nth batch, sending the computing subtasks of the (n + 1)th batch to the computing unit and, during the process of the computing unit executing the computing subtasks of the (n + 1)th batch, synchronously executing the post-processing subtasks of the nth batch, where n = 1, …, N - 2.

[0007] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, where the above-mentioned one or more processors execute the above-mentioned one or more computer programs to implement the steps of the above-mentioned method.

[0008] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above-mentioned method are implemented.

[0009] The fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above-mentioned method are implemented.

[0010] According to the embodiments of the present invention, by dividing the received multiple inference tasks into multiple batches, and sequentially inputting the computing subtasks of multiple batches into the computing unit for calculation, this strategy of separating and calculating multiple tasks can enable the computing unit not to wait during the execution of the post-processing subtasks and synchronously execute the computing subtasks. This method can improve the utilization rate of computing resources. Especially when processing large-scale data and complex models, the computing speed can be increased, and the inference efficiency can be improved without reducing the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Through the following description of the embodiments of the present invention with reference to the drawings, the above-mentioned content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0012] Figure 1 A system architecture diagram of the task execution method according to the embodiments of the present invention is shown;

[0013] Figure 2 A flowchart of the task execution method according to an embodiment of the present invention is shown;

[0014] Figure 3 The schematic diagram of the task execution method in the related art is shown;

[0015] Figure 4 The schematic diagram of the task execution method according to an embodiment of the present invention is shown;

[0016] Figure 5 The flowchart of the task execution method according to another embodiment of the present invention is shown;

[0017] Figure 6 The schematic diagram of the task execution method according to another embodiment of the present invention is shown;

[0018] Figure 7 The structural block diagram of the task execution device according to an embodiment of the present invention is shown;

[0019] Figure 8 The block diagram of the electronic device suitable for implementing the task execution method according to an embodiment of the present invention is shown. Detailed implementation manners

[0020] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0021] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0023] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0024] An embodiment of the present invention provides a task execution method, including: in response to receiving a plurality of inference requests, dividing a plurality of inference tasks generated based on the plurality of inference requests into N batches, where each inference task includes a computing subtask and a post-processing subtask to be executed in sequence, and N is an integer greater than 2; sequentially executing the inference tasks in N batches to obtain a task execution result, where executing the inference tasks in the nth batch includes: sending the computing subtasks of the nth batch to the computing unit; in the case where the computing unit has executed the computing subtasks of the nth batch, sending the computing subtasks of the (n + 1)th batch to the computing unit and, during the process of the computing unit executing the computing subtasks of the (n + 1)th batch, synchronously executing the post-processing subtasks of the nth batch, where n = 1, …, N - 2.

[0025] Figure 1 The system architecture diagram of the task execution method according to the embodiment of the present invention is shown.

[0026] As Figure 1 shown, the system architecture 100 according to this embodiment may include a processing unit 101 and a computing unit 102. Data can be transmitted between the processing unit 101 and the computing unit 102 through a network, and the network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0027] The processing unit 101 and the computing unit 102 can be various types of computing devices. For example, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), etc.

[0028] For example, it can be that the processing unit 101 is a CPU and the computing unit 102 is a GPU. For example, it can also be that the processing unit 101 is a CPU and the computing unit 102 is a CPU. For example, it can also be that the processing unit 101 is a GPU and the computing unit 102 is a GPU.

[0029] In the scenario of the embodiment of the present invention, the processing unit 101 and the computing unit 102 can cooperate with each other to jointly complete the inference tasks of semantic reasoning using a language model, and it is not limited to semantic reasoning tasks. The method of the embodiment of the present invention is applicable to any type of computing task scenario involving multiple computing stages.

[0030] Exemplarily, taking an example where a complete inference task includes pre - processing, computing, and post - processing operations, the cooperation between the processing unit 101 and the computing unit 102 to complete the inference task can be as follows: The processing unit 101 executes pre - processing and / or post - processing operations, and the computing unit 102 focuses on completing the main computing task, that is, mainly executes the forward computing operation. After the processing unit 101 finishes the pre - processing, it sends the results of the pre - processing and the computing task to the computing unit 102. The computing unit 102 executes the forward computing, and after finishing the forward computing, it returns the computing results to the processing unit 101 so that the processing unit 101 can execute the post - processing operation based on the computing results.

[0031] The following will be based on Figure 1 the described architecture, and will describe in detail the task execution method of the embodiments of the invention through Figures 2 to 8 ...

[0032] Figure 2 FIG. shows a flowchart of the task execution method according to an embodiment of the present invention. As Figure 2 shown, the task execution method of this embodiment includes operation S201 to operation S202.

[0033] In operation S201, in response to receiving multiple inference requests, the multiple inference tasks generated based on the multiple inference requests are divided into N batches, where each inference task includes a computing subtask and a post - processing subtask to be executed in sequence, and N is an integer greater than 2;

[0034] In operation S202, the N batches of inference tasks are executed in sequence to obtain the task execution result, where executing the nth batch of inference tasks includes:

[0035] Sending the nth batch of computing subtasks to the computing unit;

[0036] When the computing unit finishes executing the nth batch of computing subtasks, sending the (n + 1)th batch of computing subtasks to the computing unit and, during the process of the computing unit executing the (n + 1)th batch of computing subtasks, synchronously executing the nth batch of post - processing subtasks, where n = 1,..., N - 2.

[0037] According to the embodiments of the present invention, the above - mentioned task execution method can be applied to the scenario of semantic inference using a language model, that is, the above - mentioned inference tasks can be semantic inference tasks. It should be noted that the above - mentioned task execution method is not limited to being applied in the scenario of semantic inference. As long as it is a scenario involving computing tasks that need to be executed in two stages in sequence (the computing in the latter stage depends on the computing results of the former stage), the method of the embodiments of the present invention is applicable.

[0038] Exemplarily, the semantic reasoning task of performing semantic reasoning using a language model includes computational subtasks and post-processing subtasks to be sequentially executed. The computational subtasks mainly complete forward computations, which are the main computational tasks that the language model needs to complete and involve a large amount of computations. In this process, the language model performs a complete forward computation based on the input basic word vectors, generating multiple candidate word vectors and the respective probability distributions of the multiple candidate word vectors. This probability distribution contains the probability values of the next possible generated tokens. After the computational subtasks are completed, the post-processing subtasks are executed. The post-processing subtasks may be to select a target word vector from the multiple candidate word vectors based on a predetermined sampling strategy as the vector of the output token, and convert the target word vector selected based on the probability distribution into text to obtain the semantic reasoning result.

[0039] According to an embodiment of the present invention, it may be to use a processing unit and a computing unit to cooperate to complete the reasoning task. Among them, the processing unit and the computing unit may be various types of computing units, such as including but not limited to CPU, GPU, FPGA, etc. The processing unit and the computing unit may be the same or different. Among them, the computing unit executes the computational subtasks, and the processing unit executes the post-processing subtasks.

[0040] The processing unit and the computing unit adopt different computing units. For example, the processing unit adopts a CPU, and the computing unit adopts a GPU or an FPGA. Among them, the characteristics that the CPU is suitable for executing computational tasks with relatively complex but not particularly large amounts of computations, and the characteristics that the GPU or FPGA is suitable for relatively simple tasks but involving a large amount of computational tasks can be utilized to set the computing unit GPU or FPGA to execute the computational subtasks, and the processing unit CPU to execute the post-processing subtasks.

[0041] Among them, the above task execution method of this embodiment is executed by the processing unit.

[0042] In the above operation S201, in response to receiving multiple reasoning requests (such as multiple semantic reasoning requests initiated by multiple users), the processing unit generating multiple reasoning tasks based on the multiple reasoning requests may be to generate one reasoning task based on one request, and divide the multiple reasoning tasks into N batches.

[0043] To improve the inference efficiency of the language model, in the above method, multiple inference requests are integrated into one batch for unified processing. Specifically, when multiple users simultaneously send inference requests to the model, the system does not immediately start an inference process for each request individually. Instead, these requests are temporarily "queued" and combined into a larger batch according to certain rules. Subsequently, the model processes the input data of the entire batch at once and jointly executes the inference task. The advantage of this approach is to reduce the scheduling overhead. In the single inference request processing mode, each inference requires a series of scheduling operations such as separate task allocation and resource allocation, which themselves consume a certain amount of time and computing resources. By combining multiple requests into one batch, the system only needs to perform one scheduling and can process multiple inference tasks simultaneously, thus saving the time occupied by scheduling. In addition, this batch processing method can also improve the utilization efficiency of computing units. Computing units (such as GPUs) can usually better exert their parallel computing capabilities when processing large-scale data. When multiple inference requests are combined into one batch, the device can make more full use of its computing resources, avoiding resource waste caused by processing a single request, thereby improving the overall inference efficiency.

[0044] Furthermore, the above task execution method can further improve resource utilization by sequentially executing the inference tasks of N batches through the execution logic of the above operation S202.

[0045] Among them, for executing the inference task of any (such as the nth batch) includes: sending the nth batch of computing subtasks to the computing unit so that the computing unit executes the nth batch of computing subtasks; and, in the case where the computing unit has executed the nth batch of computing subtasks, sending the (n + 1)th batch of computing subtasks to the computing unit and, during the process of the computing unit executing the (n + 1)th batch of computing subtasks, synchronously executing the nth batch of post-processing subtasks. Through the task execution method of the embodiments of the present invention, the utilization rate of computing resources can be improved.

[0046] Figure 3 The schematic diagram of the task execution method in the related art is shown. Figure 4 The schematic diagram of the task execution method according to an embodiment of the present invention is shown. The following combines Figure 3 , Figure 4 to illustrate the technical principle by which the method of the embodiments of the present invention can improve the utilization rate of computing resources.

[0047] As Figure 3 shown, taking the processing unit as the CPU and the computing unit as the GPU as an example, in the related art, after receiving a batch of inference tasks, first the GPU executes all the computing tasks. After executing all the computing tasks, the GPU needs to transfer the computing results to the CPU for post-processing, and then the CPU executes all the post-processing tasks.

[0048] In this way, during the process of the CPU executing post-processing tasks, the GPU is in an idle state. Because these post-processing operations mainly involve data transmission and simple logical operations, rather than complex computing tasks. This idle state means that the computing resources of the computing unit are not fully utilized during this period, resulting in a waste of computing resources.

[0049] To improve the utilization rate of computing resources, the method of the embodiment of the present invention divides the received multiple inference tasks into multiple batches, and the computing subtasks of multiple batches are sequentially input into the GPU for computing. By separating and computing multiple tasks, the GPU utilization rate is improved, and the inference efficiency of the language model is further enhanced.

[0050] As Figure 4 shown, after the GPU finishes executing the computing subtasks of the first batch, the CPU executes the post-processing subtasks of the first batch, and at the same time sends the computing subtasks of the second batch to the GPU, so that the post-processing subtasks of the first batch and the computing subtasks of the second batch are executed synchronously. After the computing subtasks of the second batch are completed, the CPU executes the post-processing subtasks of the second batch... In this way, the tasks of each batch are sequentially executed until all tasks are completed.

[0051] It can be seen that by dividing the received multiple inference tasks into multiple batches according to the above method, and the computing subtasks of multiple batches are sequentially input into the computing unit for computing, this strategy of separating and computing multiple tasks can achieve that when the post-processing subtasks are executed, the computing unit does not need to wait and can execute the computing subtasks synchronously. This method can improve the utilization rate of computing resources. Especially when dealing with large-scale data and complex models, the computing speed can be increased, and the inference efficiency is improved without reducing the model performance.

[0052] According to an embodiment of the present invention, in the scenario where the above task execution method is applied to semantic inference using a language model, each inference task further includes a pre-processing subtask whose execution order is before the computing subtask.

[0053] Based on this, the above method further includes: before sending the computing subtasks of the nth batch to the computing unit, executing the pre-processing subtasks of the nth batch. Among them, the pre-processing subtasks mainly include embedding the text input by the user to generate a vector representation, that is, converting it into a data form that can be processed by the model for the model to process.

[0054] Figure 5 shows a flowchart of a task execution method according to another embodiment of the present invention. As Figure 5As shown, after dividing multiple inference tasks into N batches, for any batch of semantic inference tasks, the task execution method of this embodiment includes successively performing preprocessing, calculation, and postprocessing operations; among them, the preprocessing operation is completed by executing the preprocessing subtask, the calculation operation is completed by executing the calculation subtask, and the postprocessing operation is completed by executing the postprocessing subtask.

[0055] Exemplarily, in the scenario of using a language model for semantic inference, the language model will successively perform preprocessing, calculation, and postprocessing operations based on the information input by the user (such as text). Among them, the preprocessing operation is mainly to perform embedding processing on the text input by the user to generate a vector representation. The preprocessing operation can also include data preprocessing on the input text, such as data cleaning (duplicate removal, anomaly removal, missing value filling, etc.). The calculation operation is mainly to perform a complete forward calculation based on the word vectors generated in the previous step to generate multiple candidate word vectors and the probability distribution of each candidate word vector. The postprocessing operation is mainly to select the next token (target word vector) based on a predetermined sampling strategy and convert the target word vector selected based on the probability distribution into a semantic inference result.

[0056] According to an embodiment of the present invention, in the case of using a processing unit and a calculation unit to cooperate to complete the inference task, the task execution method may specifically include: the processing unit performs preprocessing and / or postprocessing operations, and the calculation unit focuses on completing the main calculation task, that is, mainly performing forward calculation operations. After the processing unit finishes the preprocessing, it sends the result of the preprocessing and the calculation task to the calculation unit. The calculation unit performs forward calculation, and after finishing the forward calculation, it returns the calculation result to the processing unit so that the processing unit can perform postprocessing operations based on the calculation result.

[0057] According to an embodiment of the present invention, the above method further includes: during the execution of the (n + 1)-th batch of calculation subtasks by the calculation unit, when the n-th batch of postprocessing subtasks is completed, the (n + 2)-th batch of preprocessing subtasks is executed.

[0058] The calculation unit involves a large amount of computation, so executing the calculation task takes a long time. During the execution of the current batch of calculation subtasks by the calculation unit, after the processing unit finishes the postprocessing subtasks of the previous batch, the processing unit can be used to synchronously process the preprocessing subtasks of the next batch in advance. In this way, after the calculation unit finishes the calculation subtasks of the previous batch, it can start executing the calculation subtasks of the next batch immediately without waiting. The further parallel operation of the two calculation units is realized, and the utilization rate of the computing resources is further improved.

[0059] Figure 6 Shows the schematic diagram of the task execution method according to another embodiment of the present invention. The following combinesFigure 6 The principle that the task execution method of the embodiment of the present invention can further improve resource utilization is described.

[0060] As Figure 6 shown, taking the processing unit as the CPU and the computing unit as the GPU as an example: first, the CPU executes the pre-processing subtask of the first batch; after the pre-processing subtask of the first batch is completed, the GPU executes the computing subtask of the first batch; while the GPU executes the computing subtask of the first batch, the CPU pre-executes the pre-processing subtask of the second batch in advance; after the GPU finishes executing the computing subtask of the first batch, the CPU executes the post-processing subtask of the first batch; at the same time, the GPU executes the computing subtask of the second batch; during the process of the GPU executing the computing subtask of the second batch, after the CPU finishes executing the post-processing subtask of the first batch, the CPU pre-executes the pre-processing subtask of the third batch in advance... In this way, the tasks of each batch are executed in sequence until all tasks are completed. It can be seen that during the process of the GPU executing the computing subtask of the current batch, the CPU is synchronously used to pre-synchronously process the pre-processing subtasks of the next batch in advance. In this way, after the GPU finishes executing the computing subtasks of the previous batch, it does not have to wait and can immediately start executing the computing subtasks of the next batch, improving the utilization rate of computing resources.

[0061] According to the embodiment of the present invention, in the scenario of semantic reasoning using a language model, the reasoning task is used to indicate: calling a predetermined language model to perform semantic reasoning based on the request data carried in the reasoning request to generate a semantic reasoning result, and the process of semantic reasoning includes sequentially executing a pre-processing subtask, a computing subtask, and a post-processing subtask. The language model will perform operations of pre-processing, computing, and post-processing in sequence based on the information input by the user, that is, based on the request data carried in the reasoning request initiated by the user, such as text, voice, image, etc. Semantic reasoning can be used in task scenarios such as performing question answering, translation, recommendation, etc.

[0062] The process of semantic reasoning involves multiple network structures inside the language model, and the multiple network structures of the language model can be a combination of various forms. For example, it can include an embedding layer, multiple attention layers and a mapping layer (multi-layer perceptron), a decoding layer, etc. For example, it can also include an embedding layer, multiple encoders or decoders, etc. For example, it can also include an embedding layer and a multi-layer perceptron, a decoding layer, etc. The embodiment of the present invention does not limit the network architecture of the language model.

[0063] Hereinafter, taking the architecture of the language model including an embedding layer, multiple attention layers and a mapping layer as an example, and taking the execution of the nth batch of reasoning tasks as an example, an exemplary description of the operations of sequentially performing the pre-processing subtask, the computing subtask, and the post-processing subtask is given.

[0064] According to an embodiment of the present invention, performing the nth batch of preprocessing subtasks includes: by performing the nth batch of preprocessing subtasks, calling the embedding layer in a predetermined language model to encode (word embedding) the nth batch of request data, generating basic word vectors, where the nth batch of request data is carried in an inference request for the nth batch of inference tasks.

[0065] The request data can be at least one of text, voice, image, etc. input by the user. In the case where the request data is voice or image, the preprocessing subtasks further include the operations of first converting the voice to text through speech recognition or first recognizing the text in the image through optical character recognition. After preprocessing, each word in the text becomes a vector (a list of numbers).

[0066] According to an embodiment of the present invention, performing the nth batch of computing subtasks includes: the computing unit calls the attention layer and the mapping layer in a predetermined language model by performing the nth batch of computing subtasks, so as to process the basic word vectors and generate a plurality of candidate word vectors and the probability distribution of each of the plurality of candidate word vectors.

[0067] In this structural layer of the language model, an attention layer and a mapping layer form a computing network layer. The input basic word vectors can be processed layer by layer through a plurality of computing network layers using the multi-head attention mechanism. Each computing network layer will perform feature extraction and information interaction on the input data, generating a plurality of candidate word vectors and the probability distribution of each of the plurality of candidate word vectors. This probability distribution contains the probability values of the next possible generated tokens.

[0068] Among them, the operations performed in each computing network layer include the following operations:

[0069] First, use the attention layer to calculate the correlation between every two word vectors to obtain a correlation matrix. Specifically, it includes mapping the basic word vectors to query (Query), key (Key), and value (Value) matrices, calculating the similarity between Query and Key through dot product, and obtaining attention weights after scaling; then using the attention weights to perform weighted summation on Value to generate a new vector representation (including global context information) for each token, that is, generating a plurality of candidate word vectors.

[0070] After that, through the mapping layer, such as a multi-layer perceptron (MLP), perform non-linear transformation and feature enhancement on the output of the attention layer. The specific processing includes: MLP independently processes the representation of each token, and maps the features to a higher-dimensional space through linear transformation (with an activation function in the middle), and then projects back to the original dimension to enhance the expression ability of the model, and outputs the probability distribution of each of the plurality of candidate word vectors.

[0071] According to an embodiment of the present invention, performing the post-processing subtask of the nth batch includes: by performing the post-processing subtask of the nth batch, determining a target word vector from multiple candidate word vectors based on a probability distribution, and converting the target word vector into the semantic inference result of the nth batch.

[0072] After the forward calculation is completed, the post-processing operation can be to determine a target word vector from multiple candidate word vectors based on a probability distribution, or to select the next token based on a predetermined sampling strategy. The sampling strategy can be any strategy, such as random sampling, Top-k sampling, Top-p sampling, etc. The post-processing operation also includes converting the target word vector selected based on the probability distribution into a semantic inference result (i.e., inverse embedding processing) to generate diverse text content.

[0073] In this way, after performing the operations of pre-processing, calculation, and post-processing once, a complete inference task is completed. This inference task will continuously loop in an autoregressive manner, that is, the output of each time is used as the input of the next time (because the language model needs to use the previous context information to predict the next token) until the termination condition is met, such as reaching the preset maximum generation length or generating a specific termination symbol (such as a full stop or a line break). Through this process, the language model gradually constructs a text sequence, considering the semantics and context of the previous text at each step, so as to generate coherent and natural text content.

[0074] According to an embodiment of the present invention, the on-chip storage of the computing unit is pre-divided into N storage areas corresponding to the inference tasks of N batches; the N storage areas are used to store the process data in the execution process of the inference tasks of N batches one by one, and the process data includes input data and output data. The input data is the pre-processing result generated by performing the pre-processing subtask, and the pre-processing result is used as the basic data required for the computing unit to perform the calculation subtask; after the computing unit finishes performing the calculation subtask, a calculation result is generated, and the calculation result is used as the output data.

[0075] Based on this, for any inference task of the nth batch, the above method further includes: when the pre-processing subtask of the nth batch is completed, generating the nth batch of basic data required for the computing unit to perform the nth batch of calculation subtasks, and sending the nth batch of basic data to the computing unit so that the nth batch of basic data is stored in the nth storage area among the N storage areas; reading the nth batch of calculation results generated by the computing unit based on the nth batch of basic data from the nth storage area.

[0076] After the above operations, the basic data and calculation results corresponding to the nth batch are stored in the nth storage area of the computing unit. In this way, the N storage areas store the process data in the execution process of the inference tasks of N batches one by one.

[0077] According to an embodiment of the present invention, in an actual computing environment, memory resources are often limited and scattered. By pre-allocating memory for computing units, conflicts in data reading and writing for different batches are avoided, thus avoiding blocking during the data processing process. As a result, the processing of different batches can be carried out synchronously, improving the computing efficiency. Moreover, by pre-dividing the memory area, memory fragmentation caused by the random dynamic memory allocation method is also avoided, improving the storage utilization rate (memory fragmentation will lead to a decrease in memory utilization rate and an increase in memory access latency). In the case of limited storage space, the method of the embodiment of the present invention has good applicability, improving the memory utilization rate. Since there is no memory fragmentation, all storage space can be fully utilized, reducing the performance loss caused by memory fragmentation, thereby enhancing the computing speed.

[0078] According to an embodiment of the present invention, further, each storage area is pre-divided into a first area and a second area; the nth batch of basic data is stored in the first area of the nth storage area; the nth batch of calculation results is stored in the second area of the nth storage area. That is, on the premise of pre-dividing N storage areas to store the process data of N batches of inference tasks one by one, each storage area is further divided into two independent storage areas to realize the independent storage of input data and output data. This memory division method aims to optimize the data storage and computing processes and improve the operating efficiency of the entire system.

[0079] Taking the case where the processing unit uses a CPU and the computing unit uses a GPU as an example, when the pre - processing result (i.e., the basic data) of the nth batch is transferred from the CPU side to the GPU, the basic data of the nth batch is written into the first area of the nth storage area. The GPU reads the basic data of the nth batch from the first area of the nth storage area for calculation, generates the calculation result of the nth batch, and writes the calculation result of the nth batch into the second area of the nth storage area. The CPU reads the calculation result of the nth batch from the second area of the nth storage area for post - processing calculation of the nth batch. In this process, during the process where the GPU reads the basic data of the nth batch from the first area of the nth storage area for calculation, the first area can be cleared and the basic data of the next batch can be written. After the CPU reads the calculation result of the nth batch from the second area of the nth storage area, the second area can be cleared and the calculation result of the next batch can be written. The whole process will repeat in this way until the inference task is completely finished. This repetitive way of calculation and data processing enables the whole inference process to proceed efficiently and orderly. If the first area and the second area are not pre - divided, during the process where the GPU reads the basic data of the nth batch from the nth storage area for calculation, since the storage area still needs to be occupied to store the calculation result of the nth batch, the storage area cannot be cleared until the calculation result of the nth batch is read, and then the basic data of the next batch can be written. This will prolong the waiting time and increase the conflicts of reading and writing.

[0080] According to the embodiments of the present invention, in order to avoid the conflicts of reading and writing, the input - output memory is pre - allocated. By storing the input data and the calculation results in different memory areas respectively, the blocking in the data - processing process is avoided. At the same time, the preparation and storage of the input data can also be carried out synchronously during the calculation process, reducing the conflicts and waiting time of memory access, optimizing the data - scheduling process, making the input, calculation and output of data more fluent, and reducing unnecessary waiting and delay. At the same time, the data is divided into different batches and the calculation is executed in sequence. The writing and reading of different batches can be operated in parallel, avoiding frequent memory - reading and writing conflicts, reducing the delay of memory access, improving the efficiency of data transmission, and thus indirectly accelerating the calculation speed. Moreover, the computing unit and the first processor can focus on their respective tasks, reducing the frequency of context switching. Since context switching consumes a large amount of time and computing resources, by reducing the overhead of context switching, the calculation process can be made more efficient. It can effectively cover the post - processing time each time and effectively improve the utilization rate of the computing unit.

[0081] According to an embodiment of the present invention, for any nth batch of inference tasks, the nth batch of computing subtasks includes multiple ones. Further, a plurality of computing units can be provided; the nth batch of computing subtasks are executed in parallel by the plurality of computing units. By executing the computing subtasks of the same batch in parallel by the plurality of computing units, the computing efficiency can be further improved.

[0082] According to an embodiment of the present invention, the multiple computing subtasks of the nth batch correspond to multiple inference tasks, and one computing subtask corresponds to one inference task. For example, 10 inference requests are initiated by 10 users, and the 10 inference tasks generated based on these 10 inference requests are divided into two batches, and the method of the embodiment of the present invention is used to sequentially execute the inference tasks of these 2 batches to obtain the task execution results. Among them, for any batch, it includes 5 inference tasks, and these 5 inference tasks correspond to 5 computing subtasks, and these 5 computing subtasks can be executed in parallel by a plurality of computing units.

[0083] Further, each computing subtask is divided into multiple task stages. For example, in the scenario of semantic inference using a language model, executing any computing subtask of the nth batch includes: the computing unit performs calculations by invoking multiple computing network layers (one attention layer and one mapping layer form one computing network layer) in a predetermined language model. Among them, the computing task of one computing network layer corresponds to one task stage. There is a sequential dependency relationship in the execution of multiple task stages. After the calculation of the current task stage (the current computing network layer) is completed, the result needs to be input into the next task stage (the next computing network layer) for calculation, that is, the execution of the subsequent task stage depends on the execution result of the previous task stage. In this way, multiple task stages are sequentially executed in order until the computing subtask is completed.

[0084] According to an embodiment of the present invention, the parallel execution of the multiple computing subtasks of the nth batch by the plurality of computing units may include the following operations:

[0085] The multiple task stages of the current computing subtask among the multiple computing subtasks are sequentially executed by the plurality of computing units, where the ith computing unit executes the ith task stage. For example, the processing unit executes the first task stage. After the first task stage is completed, the computing unit executes the second task stage. After the second task stage is completed, the third computing unit executes the third task stage... where i is a positive integer.

[0086] Among them, after the ith computing unit executes the ith task stage of the current computing subtask, the ith computing unit does not have to wait and can immediately execute the ith task stage of the next computing subtask by the ith computing unit, and so on until all the computing subtasks are completed.

[0087] For example, a certain batch of computing subtasks includes 3, and there are 2 computing units in total. Each computing subtask can include 2 task stages. After the processing unit finishes the first task stage of the first computing subtask, the computing unit executes the second task stage of the first computing subtask; and, after the processing unit finishes the first task stage of the first computing subtask, the processing unit does not have to wait and immediately executes the first task stage of the second computing subtask; after the computing unit finishes the second task stage of the first computing subtask, the computing unit does not have to wait and immediately executes the second task stage of the second computing subtask. After the processing unit finishes executing the first task stage of the second computing subtask, the processing unit does not have to wait and immediately executes the first task stage of the third computing subtask; after the computing unit finishes the second task stage of the second computing subtask, without waiting, the computing unit immediately executes the second task stage of the third computing subtask. In this way, multiple computing units perform parallel computing to complete multiple computing subtasks, which can fully improve the utilization rate of the computing units and speed up the computing speed.

[0088] Based on the above task execution method, the present invention also provides a task execution device. The following will be combined with Figure 7 to describe this device in detail.

[0089] Figure 7 Shows a structural block diagram of a task execution device according to an embodiment of the present invention.

[0090] As Figure 7 shown, the task execution device 700 of this embodiment includes a batching module 701 and an execution module 702.

[0091] Among them, the batching module 701 is used to, in response to receiving multiple inference requests, divide multiple inference tasks generated based on the multiple inference requests into N batches, where each inference task includes a computing subtask and a post-processing subtask to be executed in sequence, and N is an integer greater than 2; in one embodiment, the batching module 701 can be used to execute the operation S201 described above, which will not be elaborated here.

[0092] The execution module 702 is used to sequentially execute the inference tasks of N batches to obtain a task execution result, where executing the inference tasks of the nth batch includes: sending the computing subtasks of the nth batch to the computing unit; in the case where the computing unit finishes executing the computing subtasks of the nth batch, sending the computing subtasks of the (n + 1)th batch to the computing unit and during the process of the computing unit executing the computing subtasks of the (n + 1)th batch, synchronously executing the post-processing subtasks of the nth batch, n = 1,..., N - 2. In one embodiment, the execution module 702 can be used to execute the operation S202 described above, which will not be elaborated here.

[0093] According to an embodiment of the present invention, each inference task further includes a preprocessing subtask whose execution order is before the computing subtask.

[0094] The execution module 702 is further configured to execute the nth batch of preprocessing subtasks before sending the nth batch of computing subtasks to the computing unit.

[0095] According to an embodiment of the present invention, the execution module 702 is further configured to: during the process of the computing unit executing the (n + 1)th batch of computing subtasks, and when the execution of the nth batch of postprocessing subtasks is completed, execute the (n + 2)th batch of preprocessing subtasks.

[0096] According to an embodiment of the present invention, the inference task is used to indicate: calling a predetermined language model to perform semantic inference on the request data carried in the inference request to generate a semantic inference result.

[0097] The execution module 702 is configured to, by executing the nth batch of preprocessing subtasks, call the embedding layer in the predetermined language model to encode the nth batch of request data to generate basic word vectors, where the nth batch of request data is carried in the inference request for the nth batch of inference tasks.

[0098] According to an embodiment of the present invention, the execution module 702 is configured to instruct the computing unit to call the attention layer and the mapping layer in the predetermined language model by executing the nth batch of computing subtasks, so as to process the basic word vectors to generate a plurality of candidate word vectors and the probability distribution of each of the plurality of candidate word vectors.

[0099] According to an embodiment of the present invention, the execution module 702 is configured to, by executing the nth batch of postprocessing subtasks, determine a target word vector from the plurality of candidate word vectors based on the probability distribution, and convert the target word vector into the nth batch of semantic inference results.

[0100] According to an embodiment of the present invention, the on-chip storage of the computing unit is pre-divided into N storage areas corresponding to N batches of inference tasks;

[0101] According to an embodiment of the present invention, the above device further includes a generation module and a reading module.

[0102] The generation module is configured to, when the execution of the nth batch of preprocessing subtasks is completed, generate the nth batch of basic data required for the computing unit to execute the nth batch of computing subtasks, and send the nth batch of basic data to the computing unit, so as to store the nth batch of basic data in the nth storage area among the N storage areas.

[0103] The reading module is configured to read the nth batch of computing results generated by the computing unit based on the nth batch of basic data from the nth storage area.

[0104] According to an embodiment of the present invention, each storage area is pre-divided into a first area and a second area; the nth batch of basic data is stored in the first area of the nth storage area; the nth batch of calculation results is stored in the second area of the nth storage area.

[0105] According to an embodiment of the present invention: the nth batch of calculation subtasks includes a plurality of; the calculation units include a plurality of; the plurality of calculation units execute the nth batch of calculation subtasks in parallel.

[0106] According to an embodiment of the present invention, any plurality of modules among the batching module 701 and the execution module 702 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the batching module 701 and the execution module 702 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Or, at least one of the batching module 701 and the execution module 702 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0107] Figure 8 A block diagram of an electronic device suitable for implementing the task execution method according to an embodiment of the present invention is shown.

[0108] As Figure 8 shown, the electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 can also include on-board memory for caching purposes. The processor 801 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0109] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.

[0110] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.

[0111] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0112] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the above-described ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803.

[0113] An embodiment of the present invention also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the task execution method provided by the embodiment of the present invention.

[0114] When the computer program is executed by the processor 801, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0115] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0116] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or be installed from the removable medium 811. When the computer program is executed by the processor 801, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0117] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0119] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0120] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A task execution method, characterized in that, The method includes: In response to receiving multiple inference requests, dividing multiple inference tasks generated based on the multiple inference requests into N batches, where each of the inference tasks includes a computational subtask and a post-processing subtask to be executed in sequence, and N is an integer greater than 2; Sequentially executing the inference tasks in N batches to obtain task execution results, where executing the inference tasks in the nth batch includes: Sending the nth batch of computational subtasks to the computing unit; When the computing unit has completed the nth batch of computational subtasks, sending the (n + 1)th batch of computational subtasks to the computing unit and, during the process of the computing unit executing the (n + 1)th batch of computational subtasks, synchronously executing the nth batch of post-processing subtasks, where n = 1, …, N - 2.

2. The method according to claim 1, wherein Each of the inference tasks further includes a pre-processing subtask whose execution order is before the computational subtask; The method further includes: Before sending the nth batch of computational subtasks to the computing unit, executing the nth batch of pre-processing subtasks.

3. The method according to claim 2, wherein The method further includes: During the process of the computing unit executing the (n + 1)th batch of computational subtasks, when the nth batch of post-processing subtasks has been completed, executing the (n + 2)th batch of pre-processing subtasks.

4. The method according to claim 2, wherein The inference task is used to indicate: invoking a predetermined language model to perform semantic inference on the request data carried in the inference request to generate a semantic inference result; Executing the nth batch of pre-processing subtasks includes: By executing the nth batch of pre-processing subtasks, invoking the embedding layer in the predetermined language model to encode the nth batch of request data to generate basic word vectors, where the nth batch of request data is carried in the inference request for the nth batch of inference tasks.

5. The method according to claim 4, wherein The nth batch of computational subtasks is executed in the following manner: The computing unit invokes the attention layer and the mapping layer in the predetermined language model by executing the nth batch of computational subtasks to process the basic word vectors, generating multiple candidate word vectors and the probability distribution of each of the multiple candidate word vectors.

6. The method according to claim 5, characterized in that Executing the nth batch of post-processing subtasks includes: By executing the nth batch of post-processing subtasks, determining a target word vector from the multiple candidate word vectors based on the probability distribution and converting the target word vector into the nth batch of semantic inference results.

7. The method according to claim 2, wherein The on-chip memory of the computing unit is pre-divided into N storage areas corresponding to the inference tasks in N batches; The method further includes: When the nth batch of pre-processing subtasks has been completed, generating the nth batch of basic data required for the computing unit to execute the nth batch of computational subtasks and sending the nth batch of basic data to the computing unit so as to store the nth batch of basic data in the nth storage area among the N storage areas; Reading the nth batch of computational results generated by the computing unit based on the nth batch of basic data from the nth storage area.

8. The method according to claim 7, wherein: Each of the storage areas is pre-divided into a first area and a second area; The nth batch of basic data is stored in the first area of the nth storage area; The calculation result of the nth batch is stored in the second area of the nth storage area.

9. The method according to any one of claims 1-8, characterized in that: The nth batch of calculation subtasks includes a plurality of; The calculation units include a plurality of; The plurality of calculation units execute the nth batch of calculation subtasks in parallel.

10. An electronic device, comprising: One or more processors; A memory for storing one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1-9.

11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1-9 are implemented.

12. A computer program product, comprising a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1-9 are implemented.

Citation Information

Patent Citations

  • Deep learning model dynamic batch processing scheduling method and system based on resource adjustment

    CN114217966A

  • Data processing method, data processor, electronic equipment and storage medium

    CN118313458A

  • Machine learning in heterogeneous processing systems

    US20200184369A1