Task execution method, computer device, storage medium and program product

By combining the scheduling configuration information of the device parallel dimension, token prediction parallel dimension and stage parallel dimension, resource allocation information is determined, and the problem of insufficient resource scheduling flexibility in the existing technology is solved, and more efficient resource scheduling is achieved.

CN119718683BActive Publication Date: 2025-06-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510222631.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the prior art, resource scheduling is less flexible and it is difficult to effectively allocate computing resources to improve task processing efficiency.

Method used

By receiving the target request sent by the target device, and combining the scheduling configuration information of the device parallel dimension, token prediction parallel dimension and stage parallel dimension, the device allocation information and resource allocation information are determined to improve the flexibility of resource scheduling.

Benefits of technology

It realizes the combination of scheduling configuration information from multiple dimensions during resource scheduling, improves the flexibility of resource scheduling, and avoids the problem of insufficient flexibility caused by scheduling resources in separate dimensions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718683B_ABST
    Figure CN119718683B_ABST
Patent Text Reader

Abstract

The present application discloses a task execution method, a computer device, a storage medium, and a program product, relating to the technical field of resource allocation, including: a target request may include task information and scheduling configuration information in multiple dimensions. After receiving the target request, first determine device allocation information according to the scheduling configuration information of the device parallel dimension, the task information, and the resource device usage information, and then determine resource allocation information according to the task information, the device allocation information, and the scheduling configuration information of the token prediction parallel dimension and the stage parallel dimension. Finally, obtain an execution result according to the scheduling configuration information of the token prediction parallel dimension, the resource allocation information, and the task information, and feedback it to the target device. During the process of executing the task, device and resource allocation are performed from the scheduling configuration information of each dimension, and the task information itself is further combined during the execution of the task, considering the influence of multiple dimensions on resource scheduling, and solving the problem of low flexibility in the process of resource scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of resource allocation, and particularly to a task execution method, a computer device, a storage medium, and a program product. Background Art

[0002] In recent years, the scale of computer clusters has grown rapidly. At the same time, the scale of model parameters has also increased. Therefore, in order to ensure the task processing efficiency of the model, a pipelining parallel processing method is generally adopted. The processing stage of the model is divided into multiple stages, and multiple stages are run simultaneously in the computer cluster to improve the processing efficiency.

[0003] However, only allocating resources by dividing different stages has low flexibility. Summary of the Invention

[0004] This application provides a task execution method, an apparatus, a computer device, a storage medium, and a program product to at least solve the problem of low flexibility in resource scheduling in related technologies.

[0005] This application provides a task execution method, including:

[0006] Receiving a target request sent by a target device, where the target request includes a scheduling policy and task information, the scheduling policy includes scheduling configuration information in multiple dimensions, and the multiple dimensions include a device parallel dimension, a stage parallel dimension, and a token prediction parallel dimension;

[0007] Determining device allocation information according to the scheduling configuration information in the device parallel dimension, the task information, and the pre-acquired resource device usage information;

[0008] When the scheduling configuration information in the stage parallel dimension is used to indicate phased task execution, determining resource allocation information according to the task information, the device allocation information, the scheduling configuration information in the token prediction parallel dimension, and the scheduling configuration information in the stage parallel dimension;

[0009] Obtaining an execution result according to the scheduling configuration information in the token prediction parallel dimension, the resource allocation information, and the task information, and feeding back the execution result to the target device.

[0010] This application also provides a task execution apparatus, including:

[0011] A receiving module, configured to receive a target request sent by a target device, where the target request includes a scheduling policy and task information, the scheduling policy includes scheduling configuration information in multiple dimensions, and the multiple dimensions include a device parallel dimension, a stage parallel dimension, and a token prediction parallel dimension;

[0012] A determination module, configured to determine device allocation information according to the scheduling configuration information of the device parallel dimension, task information, and pre-acquired resource device usage information; and determine resource allocation information according to the task information, device allocation information, scheduling configuration information of the token prediction parallel dimension, and scheduling configuration information of the phase parallel dimension.

[0013] An acquisition module, configured to obtain an execution result according to the scheduling configuration information of the token prediction parallel dimension, resource allocation information, and task information, and feedback the execution result to the target device.

[0014] This application also provides an electronic device, including: a memory, configured to store a computer program; and a processor, configured to implement the steps of any of the above task execution methods when executing the computer program.

[0015] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above task execution methods are implemented.

[0016] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above task execution methods are implemented.

[0017] Through this application, in the process of resource scheduling, first, device allocation information is determined according to the scheduling configuration information of the device parallel dimension, task information, and resource device usage information. Then, resource allocation information can be determined according to the task information, device allocation information determined in the previous step, and the scheduling configuration information of the token prediction parallel dimension and the phase parallel dimension. In this way, in the process of resource scheduling, by combining the scheduling configuration information of multiple dimensions such as the device parallel dimension, token prediction parallel dimension, and phase parallel dimension, the problem of low flexibility caused by scheduling resources in a single dimension is avoided, and the flexibility of resource scheduling can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic diagram of the architecture of a computer cluster provided by an embodiment of this application;

[0020] Figure 2 It is a schematic flowchart of a task execution method provided by an embodiment of this application;

[0021] Figure 3Schematic diagram of a hierarchical scheduling framework provided by an embodiment of the present application;

[0022] Figure 4 Schematic diagram of the functions of a central scheduler provided by an embodiment of the present application;

[0023] Figure 5 Schematic diagram of the functions of a node scheduler provided by an embodiment of the present application;

[0024] Figure 6 Schematic diagram of the process of a computing resource allocation operation provided by an embodiment of the present application;

[0025] Figure 7 Schematic diagram of the process of a memory resource allocation operation provided by an embodiment of the present application;

[0026] Figure 8 Schematic diagram of the functions of a request scheduler provided by an embodiment of the present application;

[0027] Figure 9 Schematic diagram of a hierarchical scheduling strategy provided by an embodiment of the present application;

[0028] Figure 10 Another schematic diagram of a hierarchical scheduling framework provided by an embodiment of the present application;

[0029] Figure 11 Schematic diagram of the structure of a task execution device provided by an embodiment of the present application;

[0030] Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0033] To enable those skilled in the art of the present technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0034] An embodiment of this application provides a computer cluster. As Figure 1 shown, the computer cluster may include multiple nodes, and each node may include a Central Processing Unit (CPU) and a Graphics Processing Unit (GPU). Among them, the node may be a Bare Metal Server. The central processing unit is used to parse the received request and perform corresponding resource scheduling operations. The graphics processing unit can be used as a resource device to load a model and execute tasks using the model. For example, after using an inference model to execute an inference task, an inference result can be obtained. The client can communicate with the first node (any node) in the computer cluster, and then the first node determines whether to execute the task corresponding to the request from the client by itself according to a preset request processing mechanism, or send the request to other nodes for other nodes to execute the task corresponding to the request from the client. Any node in the computer cluster can perform the task execution method provided by the embodiment of this application, or multiple nodes in the computer cluster can cooperate with each other to perform the task execution method provided by the embodiment of this application.

[0035] In the field of inference model technology, a user can input task content (such as a paragraph, a picture, etc.) on the client, and the client can generate a request according to the content input by the user and send it to the computer cluster. After the computer cluster executes the task, the result is fed back to the client. Since the execution of inference tasks requires a large amount of computing resources, in order to ensure the execution efficiency of inference tasks, the resources in the computer cluster can be scheduled in real time.

[0036] An embodiment of this application provides a task execution method. As Figure 2 shown, the task execution method may include the following steps:

[0037] Step S201, receive a target request sent by a target device.

[0038] Among them, the target device may be a client or the first node. The target request may include a scheduling policy and task information. The scheduling policy may include scheduling configuration information in multiple dimensions, and the multiple dimensions include a device parallel dimension, a stage parallel dimension, and a token prediction parallel dimension.

[0039] The scheduling configuration information in the device parallel dimension may include indication information on whether to perform distributed inference, and / or include the number of nodes, the number of devices, etc.

[0040] The scheduling configuration information of the stage parallel dimension may include indication information on whether to perform tasks in stages. Or, in the case of indicating staged task execution, it may further include the resource allocation ratios for multiple stages. For example, in an inference task, it generally includes a prefill stage and a decode stage. The scheduling configuration information of the stage parallel dimension may include a PD separation ratio, that is, the resource allocation ratios corresponding to the prefill stage and the decode stage respectively.

[0041] The scheduling configuration information of the token prediction parallel dimension may include indication information on whether to perform multi-token prediction (MTP).

[0042] The task information may include model parameters, task content, input length, task type, model identifier, etc. For example, the model parameters may include model size and precision. The task content may be text content, pictures, audio content. The input length may be the length of the task content. For example, in an inference task, the task type may include long text type or short text type. The input length of the task content of the long text type is greater than a preset length threshold, and the input length of the task content of the short text type is less than or equal to the preset length threshold.

[0043] Specifically, the user can, according to their own needs, select a corresponding model in the client, input the corresponding task content, and input a scheduling strategy. In this way, the client can generate a target request based on the selected model, the input task content, and the scheduling strategy, and send it to the first node in the computer cluster. The first node can parse the target request to obtain the task information and the scheduling strategy, or, according to a preset request processing mechanism, send the target request to other nodes, and let other nodes parse the target request to obtain the task information and the scheduling strategy.

[0044] Step S202: Determine device allocation information according to the scheduling configuration information of the device parallel dimension, the task information, and the pre-acquired resource device usage information.

[0045] Among them, the resource device usage information may include the remaining resource amounts of multiple resource devices.

[0046] Specifically, the second node can update in real time the resource usage of each resource device deployed on each node corresponding to the user, that is, it can obtain the resource device usage information in real time. The second node can be any node in the computer cluster. For example, it can be the above-mentioned first node, or it can be the node that receives the target request sent by the first node.

[0047] The second node can determine whether to perform distributed inference according to the indication information of distributed inference included in the scheduling configuration information of the device parallel dimension. If not, it can select a resource to be used from multiple resource devices according to the model parameters and resource device usage information included in the task information. If so, it can select multiple resource devices to be used from multiple resource devices according to the model parameters and resource device usage information included in the task information, and determine the identification information of the resource devices to be used and the identification information of the node where each resource device to be used is located. Among them, the identification information of all the determined resource devices to be used and the identification information of the nodes together constitute the device allocation information.

[0048] Alternatively, the second node can determine whether to perform distributed inference according to the number of devices and the number of nodes included in the scheduling configuration information of the device parallel dimension. For example, when it is determined that both the number of devices and the number of nodes are the first preset value (for example, it can be 1), it is determined to perform distributed inference. Or, when it is determined that the number of devices or the number of nodes is not the first preset value, it is determined not to perform distributed inference. Furthermore, the second node can select the same number of resource devices to be used as the number of devices from multiple resource devices in a corresponding manner (distributed inference or non-distributed inference) according to the number of devices and the number of nodes, the model parameters included in the task information, and the resource device usage information.

[0049] Step S203: Determine the resource allocation information according to the task information, the device allocation information, the scheduling configuration information of the token prediction parallel dimension, and the scheduling configuration information of the phase parallel dimension.

[0050] Specifically, the second node can determine whether the scheduling configuration information of the phase parallel dimension (specifically, the indication information of performing tasks in stages included therein) is used to indicate performing tasks in stages, and perform the resource allocation operation corresponding to the scheduling configuration information of the phase parallel dimension. Specifically, it can be divided into the following two cases:

[0051] Case 1 can include the following steps:

[0052] Step 1: When the scheduling configuration information of the phase parallel dimension is used to indicate performing tasks in stages, obtain the resource allocation ratios of multiple stages.

[0053] Specifically, the second node can determine whether the scheduling configuration information of the phase parallel dimension includes the resource allocation ratios of multiple stages. When it is determined that the scheduling configuration information of the phase parallel dimension includes the resource allocation ratios of multiple stages, the resource allocation ratio of each stage can be extracted therefrom. Or, when it is determined that the scheduling configuration information of the phase parallel dimension does not include the resource allocation ratio, the resource allocation ratio of each stage corresponding to the task type can be obtained according to the task type in the task information.

[0054] For example, in an inference task, multiple stages include a pre-fill stage and a decoding stage. The task type can be a long text type or a short text type. When the task type is a long text type, the resource allocation ratio in the pre-fill stage is greater than that in the decoding stage. Or, when the task type is a short text type, the resource allocation ratio in the pre-fill stage is less than that in the decoding stage.

[0055] Step 2: Determine the resource allocation information based on the task information, device allocation information, resource allocation ratio for each stage, and scheduling configuration information for the token prediction parallel dimension.

[0056] Resources can include various different types of resources, such as computing resources, memory resources, communication resources, etc. Correspondingly, the second node can select multiple elements corresponding to the target resource type from the task information, device allocation information, resource allocation ratio for each stage, scheduling configuration information for the token prediction parallel dimension, and scheduling configuration information for the stage parallel dimension according to the allocation method of the target resource type, and perform an allocation operation on the resources of the target resource type based on the multiple elements corresponding to the target resource type to determine the resource allocation information for the target type. Here, the target resource type is any one of the resource types.

[0057] For example, the allocation methods for different resources can be determined as follows:

[0058] First, for computing resources. Determine the allocation information for computing resources based on the device allocation information and the resource allocation ratio for each stage.

[0059] Step 1: Count the first quantity of the identification information of the resources devices to be used included in the device allocation information.

[0060] Step 2: Determine the number of devices of the resources devices to be used for each stage according to the first quantity and the resource allocation ratio for each stage.

[0061] Step 3: Select, from the resources devices to be used, the same number of resources devices to be used as the number of devices of the resources devices to be used for the first stage as the devices for executing the task of the first stage.

[0062] Here, the first stage is any one of the multiple stages.

[0063] Step 4: jointly determine the allocation information for computing resources based on the identification information and the number of devices of the resources devices to be used for each stage.

[0064] Specifically, the second node can count the first quantity of the identification information of the resources devices to be used included in the device allocation information, that is, count the first quantity of the resources devices to be used.

[0065] Further, when the resource allocation ratio is a percentage, the product of the first number and the resource allocation ratio of the first stage can be used to determine the number of resource devices to be used in the first stage, where the first stage is any one of the multiple stages. Then, the second node can select the same number of resource devices to be used as the number of resource devices to be used in the first stage from the resource devices to be used determined in step S202 as the devices for executing the tasks of the first stage. Finally, the second node can jointly determine the identification information and the number of resource devices to be used in each stage as the allocation information of the computing resources.

[0066] After determining the allocation information of computing resources, the second node can also perform initialization operations on the resource devices to be used. For example, in the inference task, multiple GPUs can be initialized as Prefill GPUs and the remaining GPUs as Decode GPUs according to the number of resource devices to be used in each stage. The initialization operation can include operations of marking GPUs, setting the GPU deployment environment, and deploying the corresponding model on the GPU.

[0067] In some optional implementations, in the process of selecting the equipment for the task of the first stage, the second node can determine the number of resource devices to be used on each node based on the identification information of each resource device to be used and the identification information of the node where it is located, sort each node according to the number of resource devices to be used on each node, and select resource devices to be used from one or more nodes that are equal to the number of resource devices to be used in the first stage as the equipment for performing the task of the first stage according to the sorting of each node. For example, the more nodes there are, the higher they are ranked, and resource devices to be used are selected from each node in turn according to the sorting from front to back until the number of resource devices to be used selected is equal to the number of resource devices to be used in the first stage, then the operation is stopped to determine the equipment for performing the task of the first stage. In this way, the probability that the resources to be used for performing the tasks of the same stage are located on the same node increases, which can reduce the number of cross-node communications and thus reduce the problem of resource waste.

[0068] Second, memory resources: Determine the allocation information of memory resources based on the scheduling configuration information of the token prediction parallel dimension, task information, and resource allocation ratio of each stage.

[0069] Step 1: Determine the amount of memory resources to be used based on the scheduling configuration information and model parameters of the token prediction parallel dimension.

[0070] Step 2: Determine the amount of memory resources to be used in each stage according to the amount of memory resources to be used and the resource allocation ratio of each stage.

[0071] Step 3: jointly determine the amount of memory resources to be used in each stage as the allocation information of memory resources.

[0072] Specifically, the second node can determine whether to perform a multi-token prediction operation based on the indication information of multi-token prediction included in the scheduling configuration information of the token prediction parallel dimension. When the scheduling configuration information of the token prediction parallel dimension is used to indicate multi-token prediction, determine the amount of memory resources to be used according to the first preset multiple and model parameters. Or, when the scheduling configuration information of the token prediction parallel dimension is used to indicate no multi-token prediction, determine the amount of memory resources to be used according to the second preset multiple and model parameters, where the first preset multiple is greater than the second preset multiple. For example, the first preset multiple can be 2, and the second preset multiple can be 1. Since multi-token prediction requires two operations of loading the model to start Dual Pipe for two-way pipelining parallelism, more memory resources can be allocated during the memory resource allocation stage.

[0073] After determining the allocation information of memory resources, the second node can divide memory resources equal in size to the amount of memory resources to be used on its own memory resources or on the memory resources of the nodes where all resources devices to be used are located, and maintain the memory resources corresponding to different stages according to the amount of memory resources to be used in each stage. For example, in an inference task, Prefill KV Cache and Decode KV Cache are maintained separately.

[0074] After determining the allocation information of computing resources and memory resources, the second node can obtain the pre-trained model according to the model identifier and load the pre-trained model onto the resources device to be used. For example, in an inference task, when there are a prefill stage and a decode stage included in multiple stages, the second node can load the prefill model onto the resources device to be used corresponding to the prefill stage and load the decode model onto the resources device to be used corresponding to the decode stage.

[0075] Third, communication resources. Determine the allocation information of communication resources according to the device allocation information and the resource allocation ratio of each stage.

[0076] Step 1: count the second quantity of the identification information of the nodes where the resources devices to be used included in the device allocation information are located.

[0077] Step 2: determine the communication method of each resources device to be used according to the first quantity and the second quantity.

[0078] Step 3: jointly determine the communication method of each resources device to be used and the resource allocation ratio of each stage as the allocation information of communication resources.

[0079] Specifically, the second node may count the second quantity of the identification information of the nodes where the resource devices to be used included in the device allocation information are located. When it is determined that the first quantity or the second quantity is the first preset value (for example, the first preset value is 1), the communication mode of each resource device to be used is determined as the in-node communication mode. Or, when it is determined that the first quantity and the second quantity are not the first preset value, the communication mode of each resource device to be used is determined as the cross-node communication mode. For example, the in-node communication mode may be the NV Link communication mode, and the cross-node communication mode may be the InfiniBand (abbreviated as IB) communication mode. Finally, the second node may jointly determine the communication mode of each resource device to be used and the resource allocation ratio of each stage as the allocation information of the communication resources.

[0080] Alternatively, the second node may also determine the total communication resources to be used for each node according to the quantity of the resource devices to be used on each node and the preset communication resources. The second node may jointly determine the total communication resources to be used for each node, the resource allocation ratio of each stage, and the communication mode of each resource device to be used as the allocation mode of the communication resources.

[0081] In some alternative embodiments, for any node where a resource device to be used is located (taking the third node as an example for illustration), if there are resource devices to be used on the third node that execute the tasks of the pre-filling stage and the decoding stage, the communication resources of each stage on the third node can be obtained according to the third quantity of the resource devices to be used on the third node that execute the tasks of the pre-filling stage, the fourth quantity of the resource devices to be used on the third node that execute the tasks of the decoding stage, the resource allocation ratio of each stage, and the total communication resources to be used for the third node. The second node can then jointly determine the communication resources of each stage on each node and the communication mode of each resource device to be used as the allocation information of the communication resources. In this way, more detailed allocation information can be directly determined, the allocation speed of the communication resources can be accelerated, and the flexibility of the communication resource allocation can be improved.

[0082] After determining the allocation information of communication resources, the second node may set the communication method of each resource device to be used according to the communication method of each resource device to be used included in the allocation information of communication resources, and perform an allocation operation on the communication resources according to the resource allocation ratio of each stage included in the allocation information of communication resources, so that each resource device to be used can share the parameter information in the model after communicating with other resource devices to be used during the process of loading the model to execute tasks. Specifically, the second node may determine the total communication resources to be used for each node according to the number of resource devices to be used on each node and the preset communication resource amount. Furthermore, according to the resource allocation ratio of each stage and the total communication resources to be used, the communication resources to be used for each stage on each node are determined, and the communication resources where each resource device to be used is located are allocated according to the resource allocation ratio of each stage. For example, in the inference task, the KV Cache communicator is initialized according to the PD allocation ratio, that is, it is initialized as the communication resources for the pre-fill stage and the communication resources for the decoding stage.

[0083] In some alternative embodiments, the second node may determine whether the identification information of all the nodes where the resource devices to be used are located is all the same. If so, it may determine that the communication method of the resource devices to be used is the in-node communication method. If not, it may determine that each node only includes one resource device to be used. If so, it may determine that the communication method of the resource device to be used on this node is the cross-node communication method. If not, it may determine that the communication method of the resource device to be used on this node includes the in-node communication method and the cross-node communication method.

[0084] In Case 2, when the scheduling configuration information for the stage parallel dimension is used to indicate that tasks are not executed in stages, the second node can directly obtain the target model according to the model identifier included in the task information (for example, in an inference task, the target model can be a model including a pre-filling function and a decoding function). Furthermore, a target model can be loaded on each resource device to be used respectively, or the target model can be loaded onto all resource devices to be used in a distributed manner. In addition, the second node can determine the amount of memory resources to be used according to the scheduling configuration information and model parameters of the token prediction parallel dimension (the specific processing can refer to the above Case 1). The second node can divide out a memory resource of the same size as the amount of memory resources to be used on its own memory resources, or on the memory resources of the nodes where all resource devices to be used are located, and initialize it as a KV Cache. Also, the second node can count the first quantity of the identification information of the resource devices to be used included in the device allocation information, and the second quantity of the identification information of the nodes where the resource devices to be used are located included in the device allocation information, and determine the communication method for each resource device to be used according to the first quantity and the second quantity. Finally, the communication methods of all resource devices to be used can be jointly determined as the allocation information of communication resources. In this way, in the case of not executing tasks in stages, resource allocation operations can be performed from the token parallel dimension and the device allocation information determined through the device parallel dimension, with relatively high flexibility.

[0085] The allocation information of computing resources, the allocation information of memory resources, and the allocation information of communication resources determined through the above operations together constitute the resource allocation information.

[0086] Step S204, obtain the execution result according to the scheduling configuration information, resource allocation information, and task information of the token prediction parallel dimension, and feedback it to the target device.

[0087] Specifically, the second node can not only improve the task execution efficiency from the resource allocation dimension, but also improve the task execution efficiency from the dimension of task division. For example, when the target request is an inference request, executing the task corresponding to the target request can include:

[0088] Step 1, when it is determined to divide the task content into chunks according to the input length and the preset pre-filling length, determine the chunk size according to the input length and the number of resource devices to be used in the pre-filling stage.

[0089] Step 2, divide the task content into multiple chunks according to the chunk size.

[0090] Step 3: According to the identification information of the resource device to be used in the pre-filling stage, allocate a chunk to a resource device to be used in the pre-filling stage for pre-filling the chunk, and obtain a pre-filling processing result corresponding to the chunk.

[0091] Step 4: Obtain a decoding model corresponding to the scheduling configuration information of the token prediction parallel dimension.

[0092] Step 5: According to the identification information of the resource device to be used in the decoding stage, load the decoding model on the resource device to be used in the decoding stage, and use the decoding model to perform decoding processing on the pre-filling processing result corresponding to each chunk to obtain an execution result.

[0093] Specifically, the second node can determine whether the input length is greater than a preset pre-filling length (PrefillLength). If not, the task content can be determined as a chunk. If so, it is determined that chunking operation needs to be performed, and the ratio of the input length to the number of resource devices to be used in the pre-filling stage can be determined as the chunk size. Further, the second node can chunk the task content according to the chunk size to obtain multiple chunks. The second node can allocate a chunk to a resource device to be used in the pre-filling stage according to the identification information of the resource device to be used in the pre-filling stage. In this way, the resource device can perform pre-filling operation according to the loaded pre-filling model and obtain a pre-filling processing result corresponding to each chunk. When it is determined not to perform multi-token prediction according to the scheduling configuration information of the token prediction parallel dimension, the second node can use the decoding model loaded on the resource device to be used in the decoding stage to process the pre-filling processing result corresponding to each chunk to obtain an execution result. Or, when it is determined to perform multi-token prediction according to the scheduling configuration information of the token prediction parallel dimension, the second node can train multiple decoding models in real time using the current training method and deploy each decoding model to a resource device to be used in the decoding stage respectively to process the pre-filling processing result corresponding to each chunk to obtain an execution result. Finally, when the second node is the first node, the second node can send the execution result to the client. Or, when the second node is not the first node, the second node can send the execution result to the first node, and the first node feeds it back to the client.

[0094] Since multi-token prediction requires real-time training of multiple decoding models, this solution also obtains the corresponding decoding model according to the token parallel dimension, which is relatively flexible. And in the case of multi-token prediction, it can meet the need of decoding multiple words simultaneously, improve the decoding efficiency, that is, it can improve the task execution efficiency.

[0095] During the execution of the task, the second node can use the memory resources in the pre-filling stage to cache the model weights shared by the resource devices to be used in the pre-filling stage in the pre-filling stage (e.g., embeddings and heads), and use the memory resources in the decoding stage to cache the model weights shared by the resource devices to be used in the decoding stage, so as to execute the tasks in each stage in parallel. Also, the second node can use the communication resources in the pre-filling stage and the communication methods of the resource devices to be used in the pre-filling stage to share the model weights in the memory resource cache in the pre-filling stage to other resource devices to be used in the pre-filling stage, and use the communication resources in the decoding stage and the communication methods of the resource devices to be used in the decoding stage to share the model weights in the memory resource cache in the decoding stage to other resource devices to be used in the decoding stage.

[0096] In some alternative embodiments, to further improve the task execution efficiency, the second node can count the task execution efficiency during each task execution, and record the task information, adjustment strategy, and task execution efficiency of each task, so as to periodically perform the following steps:

[0097] Step 1: Obtain the task execution efficiencies corresponding to multiple tasks respectively, as well as the task types and scheduling policies corresponding to each task.

[0098] Step 2: Determine the scheduling policy corresponding to the task with the highest task execution efficiency among the tasks of the target task type as the scheduling policy corresponding to the target task type.

[0099] Wherein, the target task type is any task type.

[0100] Step 3: Send each task type and the optimal scheduling policy corresponding to each task type to the target device to instruct the target object to adjust the scheduling policy.

[0101] Wherein, the target object can be a user or the client where the user is located. The task corresponding to the target request can be one of the above multiple tasks.

[0102] In this way, when the target object needs to execute a task of the target task type, inputting the optimal scheduling policy corresponding to the target task type can greatly improve the task execution efficiency.

[0103] In some alternative embodiments, to further enhance the flexibility of resource allocation, the second node can perform the following steps:

[0104] Step 1: Monitor the processing progress in the pre-filling stage.

[0105] Step 2: When it is monitored that the processing progress is equal to the preset processing progress, obtain the resource ratio adjustment method corresponding to the preset processing progress and the latest task execution information.

[0106] Among them, there can be multiple preset processing progress values. The resource ratio adjustment method includes the percentage increase or decrease in each stage. The latest task execution information may include information about the chunks that have not completed pre-filling processing and information about the pre-filling processing results that have not completed decoding processing.

[0107] Step 3: According to the resource ratio adjustment method corresponding to the preset processing progress, adjust the resource allocation ratio of each stage to obtain the updated resource allocation ratio of each stage.

[0108] Step 4: According to the updated resource allocation ratio of each stage, reallocate each type of resource to obtain the resource allocation amount of each type corresponding to each stage.

[0109] Among them, the resource allocation amount of each stage corresponding to computing resources is the number of resource devices to be used in each stage, the resource allocation amount of each stage corresponding to memory resources is the amount of memory resources to be used in each stage, and the resource allocation amount of each stage corresponding to communication resources is the amount of communication resources to be used in each stage.

[0110] Step 5: According to the latest task execution information and the resource allocation amount of each type corresponding to each stage, continue to execute the task to obtain the execution result.

[0111] Since in the process of executing the inference task, generally the tasks in the pre-filling stage are executed first. After obtaining some preprocessing results, the tasks in the pre-filling stage and the decoding stage can be executed in parallel. As the task execution progresses, the resources required for the tasks in the pre-filling stage become less and less, and the resources required for the tasks in the decoding stage become more and more. Therefore, by monitoring the processing progress of the pre-filling stage in real time, when the processing progress of the pre-filling stage reaches the preset processing progress, the resource allocation ratio of each stage can be adjusted according to the corresponding resource ratio adjustment method (for example, reducing the resource allocation ratio of the pre-filling stage and increasing the resource allocation ratio of the decoding stage), and according to the re-determined resource allocation ratio of each stage, re-determine the resource allocation amount of each stage under each type of resource. At the same time, the second node can use the resource allocation amount of each type corresponding to each stage to continue to execute the tasks of each stage. In this way, during the task execution process, resources can be scheduled in real time according to the actual load, which can improve the task execution efficiency and has high flexibility.

[0112] In the task execution method according to the embodiments of the present application, during the process of resource scheduling, first, according to the scheduling configuration information of the device parallel dimension, task information, and resource device usage information, the device allocation information is determined. Then, according to the task information, the device allocation information determined in the previous step, and the scheduling configuration information of the token prediction parallel dimension and the stage parallel dimension, the resource allocation information can be determined. In this way, during the resource scheduling process, by combining the scheduling configuration information of multiple dimensions such as the device parallel dimension, the token prediction parallel dimension, and the stage parallel dimension, the problem of low flexibility caused by scheduling resources in a single dimension is avoided, and the flexibility of resource scheduling can be improved.

[0113] Moreover, when the related technology allocates resources by dividing stages, the resource ratio is usually fixed. However, in fact, when executing certain tasks, the stage division is not very useful for improving the task execution efficiency, but instead reduces the task execution efficiency. That is, in any case, the method of resource scheduling and task execution by using the stage division method has low flexibility. Therefore, in this solution, only when performing tasks in stages, the scheduling configuration information of the stage parallel dimension is used to determine the resource allocation information. And this solution can obtain the resource allocation ratios of corresponding multiple stages according to the task type, that is, the task type is considered in resource scheduling, which is more flexible.

[0114] Currently, in some solutions, a preset block size is used to block the task content. However, in this solution, it is determined whether to perform blocking according to the actual situation. And when it is determined to perform blocking, the block size is determined according to the number of resource devices to be used and the input length. During the blocking process, the scheduling configuration information of the device parallel dimension and the characteristics of the task information itself are combined, that is, the task execution process is considered from multiple aspects, greatly improving the flexibility and task execution efficiency.

[0115] The above task execution method will be described below with a specific example. A hierarchical scheduling framework provided by the embodiments of the present application can be as Figure 3 shown, including a Central Scheduler, a Node Scheduler, a Request Scheduler, a Compute Scheduler, a Cache Scheduler, a Communication Scheduler, a Sequence Scheduler, a Pre-Fill Decoding Scheduler (PD Scheduler), and a Token Scheduler. Each of the above schedulers is a software component and is deployed on each node of the computer cluster. In this example, the resource device is a GPU.

[0116] Among them, the central scheduler is a first-level scheduler, the node scheduler and the request scheduler are second-level schedulers, and the computing scheduler, the cache scheduler, the communication scheduler, the sequence scheduler, the pre-fill decoding scheduler, and the word scheduler are third-level schedulers. The first-level scheduler is responsible for directly managing the second-level schedulers and obtaining the allocation information generated by the third-level schedulers. The second-level schedulers are responsible for directly managing the third-level schedulers.

[0117] As Figure 4 shown, the main functions of the central scheduler may include: when the user initiates an inference request each time, initializing the node scheduler and the request scheduler, calling the node scheduler for scheduling, calling the request scheduler for scheduling, and executing the final scheduling result, etc.

[0118] As Figure 5 shown, the main functions of the node scheduler may include initializing the computing scheduler, the cache scheduler, and the communication scheduler; calling the computing scheduler for computing resource allocation; calling the cache scheduler for memory resource allocation; calling the communication scheduler for communication resource allocation; and returning the final scheduling result to the central scheduler. As Figure 6 shown, the computing scheduler may perform the allocation operation of computing resources according to step S203 and determine and record the allocation information of computing resources. As Figure 7 shown, the cache scheduler may perform the allocation operation of memory resources according to step S203 and determine and record the allocation information of computing resources.

[0119] As Figure 8 shown, the main functions of the request scheduler may include initializing the sequence scheduler, the pre-fill decoding scheduler, and the word scheduler; calling the sequence scheduler to perform chunking operations on the task content; calling the pre-fill decoding scheduler to adjust the resource allocation ratio in the pre-fill stage and the decoding stage; calling the word scheduler to train the decoding model, and returning the final scheduling result to the central scheduler.

[0120] The above operations of initializing the scheduler may include initializing the configuration parameters in the scheduler. For example, initializing the configuration parameters to preset values.

[0121] In summary, under the above hierarchical scheduling framework, the hierarchical scheduling strategy may be as Figure 9 shown, and resource scheduling can be performed from multiple dimensions such as computing resources (graphics processors), memory resources (caches), communication resources, chunking, stages, tokens (Tokens), etc., with relatively high flexibility. Moreover, through multi-layer hierarchical scheduling, it is beneficial to the management of different types of resources and the management of the task execution process.

[0122] In some alternative embodiments, the above hierarchical scheduling framework may further include an execution module (a software component), asFigure 10 As shown. The execution module may include a Data Generator, a Model Generator, and a Data Collector. Among them, the Data Generator can be used to generate virtual input data. The Model Generator can be used to generate virtual models, including the decoding models mentioned in the word scheduler. Furthermore, by copying the user model weights from the cache and the virtual models generated by the Data Generator, multiple decoding models can be trained. Then, the multiple trained decoding models can be used to decode multiple words simultaneously. After performing model inference (i.e., executing the task), the final execution result can be obtained. The Data Collector can be used to collect inference data (e.g., task execution efficiency), parameter data (e.g., task type, adjustment strategy, etc.), GPU data (e.g., allocation information of computing resources), and communication data (e.g., allocation data of communication resources) generated during the model inference process.

[0123] Based on the collected data, the Data Collector can determine the optimal scheduling strategy corresponding to each task type for subsequent task scheduling processes, further improving the task execution efficiency and the rationality of resource allocation.

[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0125] The embodiments of the present application also provide a task execution device, as Figure 11 shown. The device may include:

[0126] A receiving module 1110, configured to receive a target request sent by a target device, where the target request includes a scheduling strategy and task information, and the scheduling strategy includes scheduling configuration information in multiple dimensions, and the multiple dimensions include a device parallel dimension, a stage parallel dimension, and a token prediction parallel dimension;

[0127] A determining module 1120, configured to determine device allocation information according to the scheduling configuration information in the device parallel dimension, the task information, and the pre-acquired resource device usage information; and determine resource allocation information according to the task information, the device allocation information, the scheduling configuration information in the token prediction parallel dimension, and the scheduling configuration information in the stage parallel dimension;

[0128] An obtaining module 1130, configured to obtain an execution result according to the scheduling configuration information in the token prediction parallel dimension, the resource allocation information, and the task information, and feedback it to the target device.

[0129] In some alternative embodiments, the determining module 1120 is specifically configured to:

[0130] When the scheduling configuration information of the stage parallel dimension is used to indicate phased execution of tasks, obtain the resource allocation ratios of multiple stages;

[0131] Determine the resource allocation information according to the task information, device allocation information, resource allocation ratio of each stage, and scheduling configuration information of the token prediction parallel dimension.

[0132] In some alternative embodiments, the determining module 1120 is specifically configured to:

[0133] Determine the allocation information of computing resources according to the device allocation information and the resource allocation ratio of each stage;

[0134] Determine the allocation information of memory resources according to the scheduling configuration information of the token prediction parallel dimension, task information, and resource allocation ratio of each stage;

[0135] Determine the allocation information of communication resources according to the device allocation information and the resource allocation ratio of each stage;

[0136] Among them, the allocation information of computing resources, the allocation information of memory resources, and the allocation information of communication resources constitute the resource allocation information.

[0137] In some alternative embodiments, the device allocation information includes the identification information of the resource devices to be used and the identification information of the nodes where the resource devices to be used are located; the determining module 1120 is specifically configured to:

[0138] Count the first quantity of the identification information of the resource devices to be used included in the device allocation information;

[0139] Determine the number of resource devices to be used in each stage according to the first quantity and the resource allocation ratio of each stage;

[0140] Select, from the resource devices to be used, the same number of resource devices to be used as the number of resource devices to be used in the first stage as the devices for executing the tasks of the first stage, where the first stage is any one of the multiple stages;

[0141] Jointly determine the identification information and the number of devices of the resource devices to be used in each stage as the allocation information of computing resources.

[0142] In some alternative embodiments, the task information includes model parameters; the determining module 1120 is specifically configured to:

[0143] Determine the amount of memory resources to be used according to the scheduling configuration information of the token prediction parallel dimension and the model parameters;

[0144] Determine the amount of memory resources to be used in each stage according to the amount of memory resources to be used and the resource allocation ratio of each stage.

[0145] Collectively determine the amount of memory resources to be used in each stage as the allocation information of memory resources.

[0146] In some alternative embodiments, the determining module 1120 is specifically configured to:

[0147] Count the second quantity of the identification information of the nodes where the resource devices to be used included in the device allocation information are located;

[0148] Determine the communication mode of each resource device to be used according to the first quantity and the second quantity;

[0149] Collectively determine the communication mode of each resource device to be used and the resource allocation ratio of each stage as the allocation information of communication resources.

[0150] In some alternative embodiments, the multiple stages include a pre-filling stage and a decoding stage, and the task information includes task content and input length; the obtaining module 1130 is specifically configured to:

[0151] When it is determined to divide the task content according to the input length and a preset pre-filling length, determine the chunk size according to the input length and the number of device resources to be used in the pre-filling stage;

[0152] Divide the task content into multiple chunks according to the chunk size;

[0153] Allocate a chunk to a resource device to be used in the pre-filling stage according to the identification information of the resource device to be used in the pre-filling stage, so as to perform pre-filling processing on the chunk to obtain a pre-filling processing result corresponding to the chunk;

[0154] Obtain a decoding model corresponding to the scheduling configuration information of the token prediction parallel dimension;

[0155] Load the decoding model on the resource device to be used in the decoding stage according to the identification information of the resource device to be used in the decoding stage, and use the decoding model to perform decoding processing on the pre-filling processing result corresponding to each chunk to obtain an execution result.

[0156] In some alternative embodiments, the apparatus further includes an adjustment module 1140, configured to:

[0157] Obtain the task execution efficiency corresponding to multiple tasks respectively, and the task type and scheduling policy corresponding to each task;

[0158] Determine the scheduling policy corresponding to the task with the highest task execution efficiency among the tasks of the target task type as the optimal scheduling policy corresponding to the target task type, where the target task type is any task type;

[0159] Send each task type and the optimal scheduling policy corresponding to each task type to the target device to instruct the target object to adjust the scheduling policy.

[0160] In some alternative embodiments, the determining module 1120 is specifically configured to:

[0161] When the scheduling configuration information of the token prediction parallel dimension is used to indicate multi-token prediction, determine the amount of memory resources to be used according to the first preset multiple and the model parameters;

[0162] Or,

[0163] When the scheduling configuration information of the token prediction parallel dimension is used to indicate no multi-token prediction, determine the amount of memory resources to be used according to the second preset multiple and the model parameters, where the first preset multiple is greater than the second preset multiple.

[0164] In some alternative embodiments, the determining module 1120 is specifically configured to:

[0165] Determine whether the scheduling configuration information of the stage parallel dimension includes the resource allocation ratios of multiple stages;

[0166] When it is determined that the scheduling configuration information of the stage parallel dimension includes the resource allocation ratios of multiple stages, extract the resource allocation ratio of each stage therefrom;

[0167] Or,

[0168] When it is determined that the scheduling configuration information of the stage parallel dimension does not include the resource allocation ratio, obtain the resource allocation ratio of each stage corresponding to the task type according to the task type in the task information.

[0169] In some alternative embodiments, the multiple stages include a pre-filling stage and a decoding stage, the task type is a long text type or a short text type. When the task type is a long text type, the resource allocation ratio of the pre-filling stage is greater than that of the decoding stage, or when the task type is a short text type, the resource allocation ratio of the pre-filling stage is less than that of the decoding stage.

[0170] In some alternative embodiments, the determining module 1120 is specifically configured to:

[0171] When it is determined that the first quantity or the second quantity is the first preset value, determine the communication mode of each resource device to be used as the in-node communication mode;

[0172] Alternatively, when it is determined that the first quantity and the second quantity are not the first preset values, the communication mode of each resource to be used is determined as the cross-node communication mode.

[0173] For the description of the features in the corresponding embodiments of the task execution device, reference may be made to the relevant descriptions in the corresponding embodiments of the task execution method, which will not be elaborated here one by one.

[0174] An embodiment of the present application also provides an electronic device, as Figure 12 shown, which may include a processor 10 and a memory 20. A computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any one of the above-described embodiments of the task execution method. The electronic device may be the above-mentioned node.

[0175] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-described embodiments of the task execution method when running.

[0176] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.

[0177] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described embodiments of the task execution method.

[0178] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described embodiments of the task execution method.

[0179] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0180] The above has introduced in detail a task execution method, device, computer device, storage medium and program product provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A task execution method, characterized in that: include: Receive a target request sent by a target device, wherein the target request includes a scheduling strategy and task information, the scheduling strategy includes scheduling configuration information of multiple dimensions, and the multiple dimensions include a device parallel dimension, a stage parallel dimension, and a token prediction parallel dimension; Determine device allocation information according to the scheduling configuration information of the device parallel dimension, the task information and the pre-acquired resource device usage information; When the scheduling configuration information of the stage parallel dimension is used to indicate to execute the task in stages, obtaining resource allocation ratios of multiple stages; Determine resource allocation information according to the task information, the device allocation information, the resource allocation ratio of each stage, and the scheduling configuration information of the token prediction parallel dimension; According to the scheduling configuration information of the token prediction parallel dimension, the resource allocation information, and the task information, an execution result is obtained and fed back to the target device.

2. The task execution method according to claim 1, characterized in that: Determining the resource allocation information according to the task information, the device allocation information, the resource allocation ratio of each stage, and the scheduling configuration information of the token prediction parallel dimension includes: Determining computing resource allocation information according to the device allocation information and the resource allocation ratio of each stage; Determine memory resource allocation information according to the scheduling configuration information of the token prediction parallel dimension, the task information and the resource allocation ratio of each stage; Determining communication resource allocation information according to the device allocation information and the resource allocation ratio in each stage; The allocation information of the computing resources, the allocation information of the memory resources, and the allocation information of the communication resources constitute the resource allocation information.

3. The task execution method according to claim 2, characterized in that: The device allocation information includes identification information of the resource device to be used and identification information of the node where the resource device to be used is located; the determining the allocation information of the computing resources according to the device allocation information and the resource allocation ratio of each stage includes: Counting a first quantity of identification information of resource devices to be used included in the device allocation information; Determine the number of resource devices to be used in each stage according to the first number and the resource allocation ratio of each stage; Selecting, from the resource devices to be used, a number of resource devices to be used that is equal to the number of resource devices to be used in the first stage as devices for performing the tasks of the first stage, wherein the first stage is any one of the multiple stages; The identification information and the number of resource devices to be used in each phase are jointly determined as the allocation information of the computing resources.

4. The task execution method according to claim 2, characterized in that: The task information includes model parameters; the scheduling configuration information of the token prediction parallel dimension, the task information and the resource allocation ratio of each stage are used to determine the allocation information of the memory resources, including: Determining the amount of memory resources to be used according to the scheduling configuration information of the token prediction parallel dimension and the model parameters; Determine the amount of memory resources to be used in each stage according to the amount of memory resources to be used and the resource allocation ratio in each stage; The amount of memory resources to be used in each stage is collectively determined as the allocation information of the memory resources.

5. The task execution method according to claim 3, characterized in that: The determining, according to the device allocation information and the resource allocation ratio of each stage, the communication resource allocation information comprises: Counting a second amount of identification information of the node where the resource device to be used is located included in the device allocation information; Determining a communication mode of each resource device to be used according to the first quantity and the second quantity; The communication mode of each resource device to be used and the resource allocation ratio in each stage are jointly determined as the allocation information of the communication resources.

6. The task execution method according to claim 3, characterized in that: The multiple stages include a pre-filling stage and a decoding stage, the task information includes task content and input length; the scheduling configuration information, the resource allocation information, and the task information predicted according to the token parallel dimension, obtaining the execution result and feeding it back to the target device, including: When determining to divide the task content into blocks according to the input length and the preset pre-filling length, determining the block size according to the input length and the number of resource devices to be used in the pre-filling stage; Dividing the task content into a plurality of blocks according to the block size; According to the identification information of the resource device to be used in the pre-filling stage, one of the blocks is allocated to one of the resource devices to be used in the pre-filling stage, so as to perform pre-filling processing on the block and obtain a pre-filling processing result corresponding to the block; Obtaining a decoding model corresponding to the scheduling configuration information of the token prediction parallel dimension; According to the identification information of the resource device to be used in the decoding stage, the decoding model is loaded on the resource device to be used in the decoding stage, and the pre-filled processing result corresponding to each block is decoded using the decoding model to obtain the execution result.

7. The task execution method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtaining the task execution efficiencies corresponding to the multiple tasks, as well as the task types and scheduling strategies corresponding to each of the tasks; Determine the scheduling strategy corresponding to the task with the highest task execution efficiency among the tasks of the target task type as the optimal scheduling strategy corresponding to the target task type, wherein the target task type is any task type; Each of the task types and the optimal scheduling strategy corresponding to each of the task types are sent to the target device to instruct the target object to adjust the scheduling strategy.

8. The task execution method according to claim 4, characterized in that: The determining the amount of memory resources to be used according to the scheduling configuration information of the token prediction parallel dimension and the model parameters includes: When the scheduling configuration information of the token prediction parallel dimension is used to indicate multi-token prediction, determining the amount of memory resources to be used according to the first preset multiple and the model parameters; or, When the scheduling configuration information of the token prediction parallel dimension is used to indicate that multi-token prediction is not performed, the amount of memory resources to be used is determined according to a second preset multiple and the model parameters, wherein the first preset multiple is greater than the second preset multiple.

9. The task execution method according to any one of claims 1 to 5, characterized in that: When the scheduling configuration information of the stage parallel dimension is used to indicate to execute the task in stages, obtaining resource allocation ratios of multiple stages includes: Determining whether the scheduling configuration information of the stage parallel dimension includes the resource allocation ratios of multiple stages; When it is determined that the scheduling configuration information of the stage parallel dimension includes the resource allocation ratios of multiple stages, extracting the resource allocation ratio of each stage therefrom; or, When it is determined that the scheduling configuration information of the stage parallel dimension does not include the resource allocation ratio, the resource allocation ratio of each stage corresponding to the task type is acquired according to the task type in the task information.

10. The task execution method according to claim 9, characterized in that: The multiple stages include a pre-filling stage and a decoding stage, and the task type is a long text type or a short text type. When the task type is a long text type, the resource allocation ratio of the pre-filling stage is greater than the resource allocation ratio of the decoding stage, or, when the task type is a short text type, the resource allocation ratio of the pre-filling stage is less than the resource allocation ratio of the decoding stage.

11. The task execution method according to claim 5, characterized in that: The determining, according to the first quantity and the second quantity, a communication mode of each resource device to be used includes: When it is determined that the first number or the second number is a first preset value, determining the communication mode of each of the resource devices to be used as an intra-node communication mode; Alternatively, when it is determined that the first number and the second number are not the first preset value, the communication mode of each of the resources to be used is determined to be an inter-node communication mode.

12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the task execution method as claimed in any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the task execution method according to any one of claims 1 to 11.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the task execution method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Task scheduling method and device, computer equipment and storage medium

    CN113238848A