Scheduling method and system of inference model, electronic equipment and storage medium
By processing on the premise that the local decoding node supports decoding, and forwarding it to the remote node when it is not supported, combined with the central load balancing manager to optimize scheduling, the problem of low inference efficiency in the existing technology is solved, and more efficient load balancing and inference performance is achieved.
Patent Information
- Application Number
- CN202510408403.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-12
AI Technical Summary
In the inference process of large language models, especially under pre-filling and decoding separation architectures, existing scheduling systems cannot achieve efficient load balancing, resulting in high node performance requirements and low inference efficiency.
By performing decoding processing when supported by the local decoding node and forwarding to the remote decoding node when not supported, cross-machine scheduling is reduced, local resources are preferred for decoding, and global scheduling optimization is combined with the central load balancing manager.
Improve inference efficiency, reduce cross-machine scheduling, ensure normal completion of the decoding stage, and improve overall inference performance.
Smart Images

Figure CN120469773A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of model reasoning technology, and in particular to a scheduling method and system for a reasoning model, an electronic device, and a storage medium. Background Art
[0002] In the reasoning process of the Large Language Model (LLM), there are two core stages: the prefill stage and the decoding stage. Among them, the prefill stage performs computationally intensive tasks, while the decoding stage performs storage-intensive tasks. Therefore, if the prefill processing and decoding processing are performed on the same node successively, the node will need to have high performance in both computing and storage. Based on this, in order to perform reasoning more efficiently, an architecture with separated prefilling and decoding is proposed. Correspondingly, a scheduling system is also provided, which schedules the prefill nodes and decoding nodes separately on the architecture with separated prefilling and decoding to achieve more efficient reasoning.
[0003] However, the existing scheduling system can only implement a crude and simple load balancing mechanism. When facing complex processing tasks under a distributed architecture with multi-node collaboration, more efficient scheduling is still needed to improve inference efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a scheduling method and system for an inference model, an electronic device, and a storage medium, which at least help to improve the efficiency of completing inference.
[0005] According to some embodiments of the present application, a first aspect of the embodiments of the present application provides a scheduling method for an inference model, which is applied to a sub-load balancing manager, and the method includes: receiving a pre-filling result generated by a local pre-filling node; wherein the local pre-filling node is a pre-filling node deployed on the same machine as the sub-load balancing manager; detecting whether a local decoding node supports decoding processing of the pre-filling result; wherein the local decoding node is a decoding node deployed on the same machine as the sub-load balancing manager; if the local decoding node supports decoding processing of the pre-filling result, scheduling the local decoding node to decode the pre-filling result; if the local decoding node does not support decoding processing of the pre-filling result, forwarding the pre-filling result to other sub-load balancing managers, so as to schedule a remote decoding node through the other sub-load balancing managers to decode the pre-filling result; wherein the remote decoding node is a decoding node deployed on a different machine from the sub-load balancing manager.
[0006] In some embodiments, forwarding the pre-filled result to other sub-load balancing managers includes: sending a decoding request message to the central load balancing manager, for the central load balancing manager to select one as the target sub-load balancing manager from all the other sub-load balancers based on the decoding request message, and return the address information of the target sub-load balancing manager; receiving the address information returned by the central load balancing manager; forwarding the pre-filled result to the location indicated by the address information.
[0007] In some embodiments, the scheduling of the local decoding node to decode the pre-filled result includes: scheduling the idle local decoding node to decode the pre-filled result; or, scheduling the local decoding node with the most idle resources to decode the pre-filled result; or, scheduling the local decoding node with the lowest resource utilization to decode the pre-filled result.
[0008] In some embodiments, when supporting decoding processing of the pre-filled result, the local decoding node satisfies at least one of the following conditions: there is an idle local decoding node, there is a local decoding node whose resource utilization is less than a first threshold, there is a local decoding node whose idle resource amount is greater than a second threshold, there is a local decoding node whose resource utilization is less than the resource utilization of the local decoding nodes corresponding to other sub-load balancing managers, and there is a local decoding node whose resource amount is less than the resource amount of the local decoding nodes corresponding to other sub-load balancing managers; wherein the first threshold and / or the second threshold is: a preset value, or a parameter associated with the data volume of the pre-filled result.
[0009] According to some embodiments of the present application, the second aspect of the embodiments of the present application also provides a scheduling method for an inference model, which is applied to a central load balancing manager, the method comprising: receiving an inference request sent by a user terminal; forwarding the inference request to a sub-load balancing manager, for the sub-load balancing manager that receives the inference request, if the local decoding node supports decoding, scheduling the local decoding node, and decoding the pre-filled result obtained by scheduling the local pre-filled node to process the inference request, or, if the local decoding node does not support decoding the pre-filled result, forwarding the pre-filled result to other sub-load balancing managers, so that the remote decoding node is scheduled by the other sub-load balancing managers to decode the pre-filled result; wherein, the local pre-filled node is a pre-filled node deployed on the same machine as the sub-load balancing manager, the local decoding node is a decoding node deployed on the same machine as the sub-load balancing manager, and the remote decoding node is a decoding node deployed on a different machine from the sub-load balancing manager.
[0010] In some embodiments, forwarding the inference request to a sub-load balancing manager includes: determining a target load balancing manager from all the sub-load balancing managers based on the status information of the local pre-filled nodes and the status information of the local decoding nodes of each sub-load balancing manager; wherein the status information is used to indicate at least one of the following node conditions: load condition, queue condition, working status; and forwarding the inference request to the target load balancing manager.
[0011] In some embodiments, before determining the target load balancing manager from all the sub-load balancing managers based on the status information of the local pre-filling nodes and the status information of the local decoding nodes of each of the sub-load balancing managers, the method further includes: receiving the status information reported by the local pre-filling nodes and the local decoding nodes of each of the sub-load balancing managers.
[0012] According to some embodiments of the present application, the third aspect of the embodiments of the present application also provides a scheduling system, including: a sub-load balancing manager and a central load balancing manager; wherein, different first load balancing managers are deployed on different machines, and each machine deployed by the first load balancing manager is also deployed with at least one pre-filled node and at least one decoding node, the sub-load balancing manager is used to execute the scheduling method of the inference model as described in any one of the first aspects, and the central load balancing manager is used to execute the scheduling method of the inference model as described in any one of the second aspects.
[0013] According to some embodiments of the present application, a fourth aspect of the embodiments of the present application also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the scheduling method of the inference model as described in any one of the first aspects, or to execute the scheduling method of the inference model as described in any one of the second aspects.
[0014] According to some embodiments of the present application, the fifth aspect of the embodiments of the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the scheduling method of the inference model as described in any one of the first aspects, or implements the scheduling method of the inference model as described in any one of the second aspects.
[0015] The technical solution provided by the embodiments of the present application has at least the following advantages:
[0016] After obtaining the pre-filled result generated by the local pre-filled node, the system detects whether the local decoding node supports decoding the pre-filled result. If the local decoding node supports decoding the pre-filled result, the local decoding node is scheduled to decode the pre-filled result. If the local decoding node does not support decoding the pre-filled result, the pre-filled result is forwarded to other sub-load balancing managers, which then schedule appropriate remote decoding nodes to decode the pre-filled result generated by the local decoding node. This ensures that the decoding phase after the pre-filling phase is completed normally, giving priority to using local decoding nodes for decoding. Returning to the scheduling center for global scheduling is no longer a necessary step for inference, and cross-machine scheduling is reduced, thereby improving inference efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0018] Figure 1 This is a flowchart of a scheduling method for an inference model provided in one embodiment of the present application;
[0019] Figure 2 is a flowchart of a scheduling method of an inference model provided in another embodiment of the present application;
[0020] Figure 3is a flowchart of a scheduling method of an inference model provided in another embodiment of the present application;
[0021] Figure 4 is a flowchart of a scheduling method of an inference model provided in another embodiment of the present application;
[0022] Figure 5 is a flowchart of a scheduling method of an inference model provided in another embodiment of the present application;
[0023] Figure 6 is a structural diagram of a scheduling system and its corresponding reasoning model nodes provided in another embodiment of the present application;
[0024] Figure 7 is a structural diagram of a scheduling system and its corresponding reasoning model nodes provided in another embodiment of the present application;
[0025] Figure 8 It is a structural diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0026] As can be seen from the background technology, the inference efficiency of the existing scheduling system combined with the architecture of pre-filling and decoding separation still cannot meet the requirements.
[0027] After analysis, it was found that there are at least the following problems with the existing scheduling system and the architecture with separated pre-filling and decoding: before the pre-filling stage and the decoding stage start to execute tasks, global scheduling is required on the scheduling center. Specifically, after receiving the inference request sent by the user, the scheduling system schedules the pre-filling node from the pre-filling node cluster according to a certain strategy to perform pre-filling processing to obtain the returned pre-filling result (the first token and key-value pair vector); then, the scheduling system schedules the decoding node from the decoding node cluster according to a certain strategy to decode the pre-filling result to obtain the returned inference result and feed it back to the user. In other words, there is at least room for further improvement in inference efficiency in: reducing the global scheduling required on the scheduling center.
[0028] Based on this, the embodiments of the present application provide a scheduling method and system for an inference model, an electronic device, and a storage medium. By, when permitted, continuing to pre-fill the local pre-filled node on the local decoding node and then continuing to decode the pre-filled result obtained, it is no longer necessary to return to the scheduling center for global decoding scheduling, thereby reducing the global scheduling required by the scheduling center and improving the efficiency of completing inference. In addition, through the scheduling policy provided by the sub-load balancing manager, the local decoding node can be scheduled for decoding first, so that the pre-filling processing and decoding processing in the same inference process can be completed on the same machine (i.e., the machine where the sub-load balancer is located), which reduces cross-machine scheduling and can improve inference efficiency.
[0029] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, each embodiment of the present application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will appreciate that many technical details are provided in each embodiment of the present application to help readers better understand the present application. However, even without these technical details and various variations and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0030] The following embodiments are divided for the convenience of description and should not constitute any limitation on the specific implementation of the present application. The various embodiments can be combined with each other and referenced to each other without contradiction.
[0031] On one hand, embodiments of the present application provide a method for scheduling an inference model, applicable to a sub-load balancing manager. It should be noted that while the present embodiments are intended for scheduling inference models, considering that inference primarily comprises a pre-population phase and a decoding phase, i.e., an inference model can be primarily viewed as a pre-population phase and a decoding node, the present embodiments primarily describe scheduling the inference model by scheduling the pre-population phase and decoding nodes. This will be explained below in conjunction with the processes provided in different embodiments.
[0032] In some embodiments, as Figure 1 As shown, when applied to a sub-load balancing manager (LM), the scheduling method of the inference model includes at least the following steps:
[0033] Step 101: Receive a pre-filling result generated by a local pre-filling node; wherein the local pre-filling node is a pre-filling node deployed on the same machine as the sub-load balancing manager.
[0034] Step 102: Check whether the local decoding node supports decoding the pre-filled result; the local decoding node is a decoding node deployed on the same machine as the sub-load balancing manager. If yes, go to step 103; if not, go to step 104.
[0035] Step 103: Schedule the local decoding node to decode the pre-filled result.
[0036] Step 104: forward the pre-filled result to other sub-load balancing managers, so that other sub-load balancing managers schedule remote decoding nodes to decode the pre-filled result; the remote decoding node is a decoding node deployed on a different machine from the sub-load balancing manager.
[0037] In this way, after obtaining the pre-filled result generated by the local pre-filled node, by detecting whether the local decoding node supports decoding the pre-filled result, if the local decoding node supports decoding the pre-filled result, the local decoding node is scheduled to decode the pre-filled result; if the local decoding node does not support decoding the pre-filled result, the pre-filled result is forwarded to other sub-load balancing managers, so that other sub-load balancing managers can schedule appropriate remote decoding nodes to decode the pre-filled result generated by the local decoding node. In this way, while ensuring that the decoding stage after the pre-filling stage is completed normally, the local decoding node is preferentially used for decoding processing. Returning to the scheduling center for global scheduling is no longer a necessary step for inference, and cross-machine scheduling is reduced, achieving the effect of improving inference efficiency.
[0038] To facilitate better understanding of those skilled in the art Figure 1 The embodiment shown is described below and its steps are explained.
[0039] In step 101, the pre-filling result generated by the local pre-filling node is received; wherein, the local pre-filling node is a pre-filling node deployed on the same machine as the sub-load balancing manager. In the embodiment of the present application, there is a communication path between the sub-load balancing manager and the local pre-filling node, so the pre-filling result of the local pre-filling node can be sent and received by the sub-load balancing manager. The pre-filling result is the output obtained by the local pre-filling node after pre-filling the input (input or prompt) of the request inference, which usually includes the first token and the key-value pair vector. Among them, the embodiment of the present application does not limit the input of the local pre-filling node. It can be the inference request itself initiated by the user, or it can be the inference request combined with the token, key-value pair vector, etc. obtained by historical pre-filling processing as a reference, etc., which will not be listed one by one here.
[0040] It should be noted that the embodiments of the present application do not limit the manner in which the local pre-filling node triggers the pre-filling process and obtains the pre-filling result. In some embodiments, a central load balancing manager can be configured that is located above the sub-load balancing manager and can support scheduling between sub-load balancing managers (i.e., global scheduling of pre-filling nodes and decoding nodes under the inference model is achieved through the scheduling of the sub-load balancing manager). On this basis, the inference request can first be received by the central load balancing manager, and then the central load balancing manager selects a pre-filling node based on the status of several pre-filling nodes in the inference model, and forwards the inference request (or the inference request combined with the token, key-value pair vector, etc. obtained by the historical pre-filling process as a reference) to the corresponding sub-load balancing manager, which is then forwarded by the sub-load balancing manager to the selected pre-filling node for pre-filling processing to obtain the pre-filling result. In this way, the management of the second stage can also be achieved through the sub-load balancing manager, which is conducive to more stable, reliable and unified scheduling. Of course, the above is only an example. In some embodiments, the central load balancing manager can also directly establish a connection with the pre-filling node, so that the inference request (or the inference request combined with the token, key-value pair vector, etc. obtained by historical pre-filling processing as a reference, etc.) can be forwarded directly from the central load balancing manager to the pre-filling node. In this way, there is no need for scheduling and forwarding by the sub-load balancing manager, and the pre-filling processing can be completed more efficiently. They will not be listed one by one here.
[0041] In addition, in the case where the sub-load balancing manager needs to decide on a local pre-filled node to process an inference request, the embodiments of the present application do not limit the scheduling strategy of the sub-load balancing manager. In some embodiments, the sub-load balancing manager may give priority to selecting a local pre-filled node with a shorter request queue to process the inference request, thereby optimizing resource allocation. Among them, the request queue temporarily stores inference requests waiting to be processed by the node where it is located. The length of the request queue can be expressed based on the number of requests in the queue, or based on the total length of the user input (input or prompt) carried by the requests in the queue. In some embodiments, in order to make better scheduling decisions, the sub-load balancing manager can also maintain a corresponding process sharing list for the local pre-filled node, and support receiving status information of the local decoding node to dynamically adjust the scheduling strategy. Of course, the above is only an example, and the scheduling strategy can also be the one with the largest amount of idle resources, etc., which will not be listed here one by one.
[0042] In step 102, a check is performed to determine whether the local decoding node supports decoding the pre-populated result. The local decoding node is a decoding node deployed on the same machine as the sub-load balancing manager. This embodiment of the present application does not limit the determination of whether the local decoding node supports decoding the pre-populated result; the determination can be flexibly configured based on application scenarios, user needs, device capabilities, and other factors.
[0043] In some embodiments, when supporting decoding of pre-populated results, a local decoding node satisfies at least one of the following conditions: an idle local decoding node, a local decoding node with a resource utilization less than a first threshold, a local decoding node with an idle resource amount greater than a second threshold, a local decoding node with a resource utilization less than that of local decoding nodes corresponding to other sub-load balancing managers, or a local decoding node with a resource amount less than that of local decoding nodes corresponding to other sub-load balancing managers. The first threshold and / or the second threshold are preset values or parameters associated with the data volume of the pre-populated results. Therefore, whether a local decoding node supports decoding of pre-populated results can be determined by determining whether the above conditions are met.
[0044] It should be noted that the above is merely an example, primarily based on the resource capabilities of the current sub-load balancing manager itself, or its relative resource capabilities relative to other sub-load balancing managers, thereby facilitating more real-time and efficient decoding and inference. However, in some embodiments, other conditions can be set to determine whether the local decoding node supports decoding of pre-populated results, which are not listed here.
[0045] In step 103, a local decoding node is scheduled to decode the pre-filled result. Since the local decoding node is scheduled, data transfer and communication between the completion of the pre-filling process and the start of the decoding process are faster, more stable, and more efficient.
[0046] It should be noted that the local decoding nodes corresponding to the sub-load balancing manager may exist in various situations. Therefore, the scheduling of local decoding nodes can be implemented in various ways. For example, in some embodiments, an idle local decoding node can be scheduled to decode the pre-populated result. In this way, since the idle local decoding node is scheduled, the pre-populated result can be decoded immediately without waiting, and decoding can be performed with ample resources. In another example, in some embodiments, the local decoding node with the most idle resources is scheduled to decode the pre-populated result. In this way, decoding can be performed with the most ample resources, which is more efficient. In another example, in some embodiments, the local decoding node with the lowest resource utilization is scheduled to decode the pre-populated result. In this way, decoding can be performed more centrally, which is more efficient. Specifically, in some embodiments, the num_token_use_percent parameter is used as the basis for selecting a decoding node, and the local decoding node with the smallest num_token_use_percent value is selected through the NVIDIA Collective Communication Library (NCCL) communication method to perform decoding. Of course, the above are only examples. In some embodiments, other strategies may be used to schedule local decoding nodes, which will not be listed here one by one.
[0047] It should be noted that the specific decoding process has been recorded in the relevant technology and will not be described in detail here.
[0048] In step 104, the pre-populated result is forwarded to another sub-load balancing manager, which then schedules a remote decoding node to decode the pre-populated result. A remote decoding node is a decoding node deployed on a different machine from the sub-load balancing manager. Since the local decoding node does not support decoding the pre-populated result, decoding is forwarded to another sub-load balancing manager to ensure proper inference and user experience.
[0049] It should be noted that the term "remote" in step 104 refers to the remote decoding node relative to the current sub-load balancing manager, meaning that the decoding of the pre-filled result is performed by a decoding node on another machine. However, from the perspective of other sub-load balancing managers that receive the forwarded pre-filled result, they schedule local decoding nodes to decode the pre-filled result. Taking sub-load balancing managers a1 and b1 as examples, assuming that sub-load balancing manager a1, pre-filling node a2, and decoding node a3 are deployed on machine A, and sub-load balancing manager b1, pre-filling node b2, and decoding node b3 are deployed on machine B, then: pre-filling node a2 and decoding node a3 are both the local pre-filling node and local decoding node of sub-load balancing manager a1, and the remote pre-filling node and remote decoding node of sub-load balancing manager b1, respectively; pre-filling node b2 and decoding node b3 are both the remote pre-filling node and remote decoding node of sub-load balancing manager a1, and the local pre-filling node and local decoding node of sub-load balancing manager b1, respectively. When the sub-load balancing manager a1 receives the pre-filling result of the pre-filling node a2 and the decoding node a3 does not support the decoding processing of the pre-filling result, the sub-load balancing manager a1 can forward the pre-filling result of the pre-filling node a2 to the sub-load balancing manager b1, and the sub-load balancing manager b1 calls the decoding node b3 to decode the pre-filling result of the pre-filling node a2.
[0050] It should also be noted that the above description of the inference scheduling process provided by sub-load balancing manager a1 and sub-load balancing manager b1 is merely illustrative and does not imply that there is only one pre-population node and one decoding node on the same machine as the sub-load balancing manager and other sub-load balancing managers. In some embodiments, if multiple decoding nodes are deployed on the machine where other sub-load balancing managers reside, the other sub-load balancing managers can use a load balancing method to schedule the local decoding nodes of the other sub-load balancing managers (the remote decoding nodes of the current sub-load balancing manager), which is more conducive to improving the inference effect.
[0051] It is understandable that in Figure 1 Based on the illustrated embodiment, a central load balancing manager can also be used to provide global load balancing. This load balancing can involve load balancing of inference requests or the aforementioned load balancing when forwarding pre-populated results for decoding processing, thereby further improving efficiency. The following uses load balancing when forwarding pre-populated results for decoding processing as an example to illustrate.
[0052] In some embodiments, as Figure 2As shown, when applied to a sub-load balancing manager, the scheduling method of the inference model includes at least the following steps:
[0053] Step 201: receiving a pre-filling result generated by a local pre-filling node; wherein the local pre-filling node is a pre-filling node deployed on the same machine as the sub-load balancing manager.
[0054] Step 202: Check whether the local decoding node supports decoding the pre-filled result; the local decoding node is a decoding node deployed on the same machine as the sub-load balancing manager. If yes, go to step 203; if not, go to step 204.
[0055] Step 203: Schedule the local decoding node to decode the pre-filled result.
[0056] Step 204: Send a decoding request message to the central load balancing manager, so that the central load balancing manager selects a target sub-load balancing manager from all other sub-load balancers according to the decoding request message, and returns the address information of the target sub-load balancing manager.
[0057] Step 205: Receive the address information returned by the central load balancing manager.
[0058] Step 206: forward the pre-filled result to the location indicated by the address information, so that other sub-load balancing managers can schedule remote decoding nodes to perform decoding processing on the pre-filled result.
[0059] In this way, based on the above-mentioned embodiment, a central load balancing manager is further introduced, which is responsible for determining a decoding node suitable for performing decoding processing for the current sub-load balancing manager globally when the pre-filled result cannot be processed by the local decoding node of the current sub-load balancing manager, thereby improving efficiency through global scheduling.
[0060] To facilitate better understanding of those skilled in the art Figure 2 The steps of the embodiment shown are explained below. Steps 201 to 203 are substantially the same as steps 101 to 103 of the aforementioned embodiment, with the main difference being that step 104 is implemented via steps 204 to 206. Therefore, steps 204 to 206 will be explained below, and steps 201 to 203 will not be described again.
[0061] In step 204, a decoding request message is sent to the central load balancing manager. This embodiment of the present application does not limit the decoding request message; it can be any message that can trigger the central load balancing manager to initiate the following actions: selecting a target sub-load balancing manager from all other sub-load balancers and returning the address information of the target sub-load balancing manager. For example, the decoding request message can be a notification message or request containing a specified string, which will not be listed here.
[0062] In step 205, the address information returned by the central load balancing manager is received. This embodiment of the present application does not limit the address information of other sub-load balancing managers. It can be any information that can indicate other sub-load balancing managers, such as an IP address or a unique identifier, which will not be listed here.
[0063] It should be noted that the address information of the other sub-load balancing managers can only include the address information of the load balancing manager. In this way, the decoding processing can be performed efficiently and with high quality by scheduling on the other load balancing managers again. The address information has a small amount of data and is easier to transmit. It can also include the address information of its local decoding node on the basis of the address information of the load balancing manager. In this way, there is no need to use other sub-load balancing managers for scheduling, which is more efficient.
[0064] Step 206: Forward the pre-population result to the location indicated by the address information. This allows the decoding nodes of other sub-load balancing managers to be scheduled to decode the pre-population result. In this embodiment of the present application, the current sub-load balancing manager may have a communication connection with the other sub-load balancing managers. Therefore, the pre-population result can be sent from the current sub-load balancing manager to the other sub-load balancing managers.
[0065] It should be noted that Figure 2 This is only an example of an implementation method for forwarding pre-filled results. In some embodiments, the pre-filled results can also be forwarded to the central load balancing manager, and then sent by the central load balancing manager to other corresponding sub-load balancing managers. In this way, there is no need to maintain the communication connection between the sub-load balancing managers, making the maintenance of the scheduling system easier and less costly, etc., which will not be listed one by one here.
[0066] Correspondingly, in conjunction with the work of the sub-load balancing manager, the second aspect of the embodiment of the present application further provides a scheduling method of an inference model, which is applied to the central load balancing manager. In some embodiments, Figure 3 As shown, the scheduling method process of the inference model applied to the central load balancing manager includes at least the following steps:
[0067] Step 301: Receive an inference request sent by a user terminal.
[0068] Step 302: forward the inference request to a sub-load balancing manager. The sub-load balancing manager that receives the inference request schedules the local decoding node to decode the pre-filled result obtained by scheduling the local pre-filled node to process the inference request if the local decoding node supports decoding. Alternatively, if the local decoding node does not support decoding of the pre-filled result, the pre-filled result is forwarded to other sub-load balancing managers so that the remote decoding nodes are scheduled by other sub-load balancing managers to decode the pre-filled result.
[0069] Among them, the local pre-filling node is the pre-filling node deployed on the same machine as the sub-load balancing manager, the local decoding node is the decoding node deployed on the same machine as the sub-load balancing manager, and the remote decoding node is the decoding node deployed on a different machine from the sub-load balancing manager.
[0070] In this way, after the inference request is forwarded by the central load balancing manager to the corresponding sub-load balancing manager, if the local decoding node supports decoding the pre-filled result, the local decoding node can be scheduled to decode the pre-filled result. If the local decoding node does not support decoding the pre-filled result, the pre-filled result will be forwarded to other sub-load balancing managers, which will then schedule appropriate remote decoding nodes to decode the pre-filled result generated by the local pre-filled node. This allows the local decoding node to be prioritized for decoding while ensuring the normal completion of the decoding phase after the pre-filling phase. Returning to the scheduling center for global scheduling is no longer a necessary step for inference, and cross-machine scheduling is reduced, achieving the effect of improving inference efficiency.
[0071] It should be noted that the forwarding of inference requests in the above embodiments does not necessarily only forward inference requests. In some embodiments, when the central load balancing manager forwards the inference request to the sub-load balancing manager, it can also simultaneously forward the corresponding key-value pair vector obtained historically (the key-value pair vector has a certain correlation with the current inference request, and the key-value pair vector may not be the complete key-value pair vector obtained historically. The specific length of the key-value pair vector depends on the strength of the correlation between the current inference request and the historically processed inference request), etc., which will not be elaborated here.
[0072] It is also necessary to point out that it is not difficult to find that Figure 3 The embodiment shown is Figure 1 The method embodiment corresponding to the embodiment shown, Figure 3 The illustrated embodiment can be used with Figure 1 The illustrated embodiments are implemented in conjunction with each other. Figure 1The relevant technical details mentioned in the embodiment shown are Figure 3 The embodiments shown are still valid and will not be described here in order to reduce repetition. Figure 3 The relevant technical details mentioned in the embodiment shown can also be applied to Figure 1 The method embodiment shown will not be described in detail here.
[0073] As mentioned above, from the perspective of improving efficiency, it is hoped that the pre-filling processing and decoding processing of the same inference request can be completed on the same machine as much as possible. To this end, in some embodiments, such as Figure 4 As shown, the scheduling method process of the inference model applied to the central load balancing manager may also include at least the following steps:
[0074] Step 401: Receive an inference request sent by a user terminal.
[0075] Step 402, determining the target load balancing manager from all sub-load balancing managers based on the status information of the local pre-filled nodes and the status information of the local decoding nodes of each sub-load balancing manager; wherein the status information is used to indicate at least one of the following node conditions: load condition, queue condition, and working status.
[0076] Step 403: forward the inference request to the target load balancing manager, so that the sub-load balancing manager that receives the inference request can schedule the local decoding node to decode the pre-filled result obtained by scheduling the local pre-filled node to process the inference request if the local decoding node supports decoding, or forward the pre-filled result to other sub-load balancing managers to schedule remote decoding nodes to decode the pre-filled result if the local decoding node does not support decoding of the pre-filled result.
[0077] In this way, on the basis of the aforementioned embodiments, load balancing scheduling of inference requests is further provided. In particular, when the scheduled inference request will first enter the pre-filling process, not only the status information of the local pre-filling node is considered when determining the target load balancing manager, but also the status information of the local decoding node is considered. When selecting the scheduling inference request to enter the pre-filling stage, not only the pre-filling stage is considered, but also the subsequent decoding stage that may occur on the same machine can be combined, which is conducive to executing the pre-filling processing and decoding processing of the same inference request on the same machine in sequence, further avoiding problems such as low efficiency caused by cross-machine operations.
[0078] To facilitate understanding of the above embodiment, the steps are explained below. Step 401 is substantially the same as step 301 of the above embodiment, and step 403 is substantially the same as step 302 of the above embodiment. The main difference is that in step 403, the sub-load balancing manager that will receive the inference request is further clarified as the target load balancing manager determined for processing in step 402. A detailed description of each step will not be repeated here.
[0079] In step 402, a target load balancing manager is determined from all sub-load balancing managers based on the status information of the local pre-populated nodes and the status information of the local decoding nodes of each sub-load balancing manager. The status information indicates at least one of the following node conditions: load condition, queue condition, and operating status. In the embodiments of the present application, load condition, queue condition, and operating status are examples. The status information may include any information that can provide a reference for executing policies such as load balancing, and is not listed here.
[0080] It should be noted that the embodiment of the present application emphasizes that when selecting the target load balancing manager, the subsequent pre-filling and decoding processing of the same inference request on the same machine is considered, but it does not mean that only the subsequent pre-filling and decoding processing of the same inference request on the same machine is considered. In other words, the embodiment of the present application does not limit the target load balancing manager that is finally determined. In some embodiments, the target load balancing manager can be the sub-load balancing manager with the highest probability of performing the pre-filling and decoding processing of the same inference request on the same machine. This can maximize the guarantee of the subsequent pre-filling and decoding processing of the same inference request on the same machine and avoid problems such as low efficiency caused by cross-machine operations to the greatest extent. In some embodiments, the target load balancing manager can be the sub-load balancing manager with the highest probability of performing the pre-filling and decoding processing of the same inference request on the same machine in which the amount of idle resources exceeds the amount of resources required to respond to the inference request. In this way, the normal response of the subsequent inference process is guaranteed while the problem of low efficiency is avoided as much as possible. These problems will not be listed here one by one.
[0081] It is also understandable that in order to make better decisions, it is desirable to obtain more accurate status information. Based on this, in some embodiments, such as Figure 5 As shown, when applied to a sub-load balancing manager, the scheduling method of the inference model includes at least the following steps:
[0082] Step 501: Receive status information reported by the local pre-filling node and the local decoding node of each sub-load balancing manager.
[0083] Step 502: Receive an inference request sent by a user terminal.
[0084] Step 503, determine the target load balancing manager from all sub-load balancing managers based on the status information of the local pre-filled nodes and the status information of the local decoding nodes of each sub-load balancing manager; wherein the status information is used to indicate at least one of the following node conditions: load condition, queue condition, and working status.
[0085] Step 504: forward the inference request to the target load balancing manager, so that the sub-load balancing manager that receives the inference request can schedule the local decoding node to decode the pre-filled result obtained by scheduling the local pre-filled node to process the inference request if the local decoding node supports decoding, or forward the pre-filled result to other sub-load balancing managers to schedule remote decoding nodes to decode the pre-filled result if the local decoding node does not support decoding of the pre-filled result.
[0086] On the basis of the above embodiment, the reporting of status information is further combined so that the central load balancing manager can perform global scheduling, thereby further improving efficiency.
[0087] To facilitate those skilled in the art to better understand the above embodiment, the steps are explained below. Among them, steps 502 to 504 are substantially the same as steps 401 to 403 of the above embodiment, and are not described in detail here.
[0088] In step 501, the status information reported by the local pre-filling node and the local decoding node of each sub-load balancing manager is received. The embodiment of the present application does not limit the reporting of status information. It can be reported when the queue of the local pre-filling node and / or the local decoding node changes, or it can be reported in real time, or it can be reported in an asynchronous manner, etc., which will not be described in detail here. In addition, the embodiment of the present application does not limit the communication method between the sub-load balancing manager, the central load balancing manager, and the pre-filling node and the decoding node. For example, it can be communicated through the Hypertext Transfer Protocol (HTTP) or the Google Remote Procedure Call (gRPC) protocol to ensure the stability and efficiency of the system under high load conditions, etc., which will not be listed here one by one.
[0089] It should be noted that if the status information is reported periodically, the reporting of the status information can be regarded as a way to implement heartbeat detection, so that the central load balancing manager can promptly perceive whether the corresponding node is online, so as to implement model management through the central load balancing manager. Of course, the heartbeat mechanism can also be directly configured, and the sub-load balancing manager or the central load balancing manager periodically sends heartbeat packets to the pre-filled node and the decoding node. If the heartbeat packet is not received within the specified time, the corresponding node is marked as inactive, thereby realizing node health status monitoring. I will not list them one by one here.
[0090] It should also be noted that in the above embodiment, the pre-filling nodes and decoding nodes of the inference model, as local pre-filling nodes and local decoding nodes of the corresponding sub-load balancing managers, can directly report status information to the central load balancing manager. However, this is merely an example. In some embodiments, the pre-filling nodes and decoding nodes of the inference model can also report status information to the central load balancing manager via the sub-load balancing manager. In other words, the pre-filling nodes and decoding nodes of the inference model first report status information to the corresponding sub-load balancing manager, which then reports the received status information to the central load balancing manager. In this way, the status information will pass through the sub-load balancing manager, and the sub-load balancing manager will subsequently make relevant decisions. For example, the sub-load balancing manager can make decisions based on this status information, such as detecting whether the local decoding node supports decoding the pre-filled result and scheduling the decoding node. Various ways of reporting status information will not be listed here.
[0091] Therefore, corresponding to the above, when the central load balancing manager directly receives the status information reported by the pre-filling node and decoding node of the inference model, the central load balancing manager can specify the pre-filling node at the same time when forwarding the inference request, so that the corresponding sub-load balancing manager can directly forward the inference request to the pre-filling node specified by the central load balancing manager after receiving the inference request. In this way, the scheduling workload of the sub-load balancing manager can be reduced, its work efficiency can be improved, and the user experience can be improved. When the central load balancing manager indirectly receives the status information reported by the pre-filling node and decoding node of the inference model through the sub-load balancing manager, when forwarding the inference request, the central load balancing manager can not specify the pre-filling node, but the sub-load balancing manager can schedule the received inference request between its corresponding local pre-filling node. In this way, the amount of information that needs to be transmitted can be reduced, the communication overhead can be reduced, etc., which will not be listed one by one here.
[0092] It can be seen that determining the target load balancing manager is only an example of the application of status information in the scheduling process of the pre-filling stage. In some embodiments, the status information can also be used for scheduling in the decoding stage when the local decoding node does not support decoding processing of the pre-filling results. These will not be listed here one by one.
[0093] It is not difficult to find that the embodiments involved in the second aspect are method embodiments corresponding to the embodiments involved in the first aspect, and the embodiments involved in the second aspect can be implemented in conjunction with the embodiments involved in the first aspect. The relevant technical details mentioned in the embodiments involved in the first aspect are still valid in the embodiments involved in the second aspect. In order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in the embodiments involved in the second aspect can also be applied to the method embodiments involved in the first aspect and are not repeated here.
[0094] It should be noted that while the above method embodiments limit the scheduling strategies within pre-population nodes and decoding nodes, this does not mean that other scheduling strategies cannot be employed. For example, in some embodiments, multi-process queues can be set up on decoding nodes to store pending data, thereby addressing issues such as uneven resource allocation and task congestion that may arise on decoding nodes, enhancing the system's processing power and response speed, and ensuring higher decoding efficiency.
[0095] The steps of the various methods above are divided only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this patent.
[0096] The third aspect of the embodiment of the present application further provides a scheduling system, such as Figure 6 、 Figure 7 As shown, the scheduling system includes: a sub-load balancing manager and a central load balancing manager.
[0097] Among them, different central load balancing managers deploy different machines, and each central load balancing manager deploys at least one pre-filled node and at least one decoding node. The sub-load balancing manager is used to execute the scheduling method of the inference model as described in any embodiment of the first aspect, and the central load balancing manager is used to execute the scheduling method of the inference model as described in any embodiment of the second aspect.
[0098] It should be noted that the embodiments of the present application do not limit the number of central load balancing managers. A scheduling system may have 1, 2, 3, 32, etc. central load balancing managers. Furthermore, the embodiments of the present application do not limit the number of pre-filling nodes and decoding nodes corresponding to a single central load balancing manager. The number of pre-filling nodes and decoding nodes may be the same or different. The specific number of pre-filling nodes and / or decoding nodes may be 1, 2, 3, 14, etc. This will not be further elaborated here.
[0099] It should also be noted that Figure 6 and Figure 7 The main difference lies in the relationship between the inference model (including pre-filled nodes and decoding nodes) and the central load balancing manager. Specifically, Figure 6 There is a connection between the pre-filling node and the decoding node and the central load balancing manager, so it can support the pre-filling node and the decoding node in the above method embodiment to directly report status information to the central load balancing manager; Figure 7 There is no connection between the pre-filling node and the decoding node and the central load balancing manager. Therefore, the pre-filling node and the decoding node described in the aforementioned method embodiment can support reporting status information indirectly to the central load balancing manager through the sub-load balancing manager, etc., which will not be repeated here.
[0100] In this way, the above-mentioned scheduling system realizes a multi-machine and multi-load balancing manager collaborative architecture to improve the scalability and fault tolerance of the system. It can support the fault tolerance and expansion of the pre-filled clusters and decoding machines composed of the pre-filled nodes and decoding nodes in the inference model respectively. It can adapt to application scenarios of different scales and complexities, and can flexibly respond to application scenarios of different scales and complexities, especially supporting the efficient processing of large-scale concurrent tasks.
[0101] It is not difficult to find that this embodiment is a system embodiment corresponding to the method embodiment, and this embodiment can be implemented in conjunction with the method embodiment. The relevant technical details mentioned in the method embodiment are still valid in this embodiment, and to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the method embodiment.
[0102] It should be noted that, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems raised by this application, but this does not mean that there are no other units in this embodiment.
[0103] A fourth aspect of the present application also provides an electronic device, such as Figure 8As shown, it includes: at least one processor 801; and a memory 802 communicatively connected to the at least one processor 801; wherein the memory 802 stores instructions that can be executed by the at least one processor 801, and the instructions are executed by the at least one processor 801 to enable the at least one processor 801 to execute the scheduling method of the inference model described in any of the above method embodiments.
[0104] The memory 802 and processor 801 are connected using a bus. The bus may include any number of interconnected buses and bridges, connecting various circuits of one or more processors 801 and memory 802. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and are therefore not described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor 801 is transmitted over a wireless medium via an antenna. Furthermore, the antenna receives data and transmits it to the processor 801.
[0105] The processor 801 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 802 can be used to store data used by the processor 801 when performing operations.
[0106] A fifth aspect of the present application further provides a computer-readable storage medium storing a computer program that implements the above method embodiment when executed by a processor.
[0107] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a device (which may be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0108] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and that in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.
Claims
1. A scheduling method for an inference model, characterized in that: Applied to a sub-load balancing manager, the method includes: Receiving a pre-filling result generated by a local pre-filling node; wherein the local pre-filling node is a pre-filling node deployed on the same machine as the sub-load balancing manager; Detecting whether a local decoding node supports decoding the pre-filled result; wherein the local decoding node is a decoding node deployed on the same machine as the sub-load balancing manager; In a case where the local decoding node supports decoding the pre-filled result, scheduling the local decoding node to decode the pre-filled result; In the case that the local decoding node does not support decoding processing of the pre-filled result, the pre-filled result is forwarded to other sub-load balancing managers, so that the remote decoding node is scheduled by the other sub-load balancing managers to decode the pre-filled result; wherein, the remote decoding node is a decoding node deployed on a different machine from the sub-load balancing manager.
2. The scheduling method of the inference model according to claim 1, characterized in that: The forwarding of the pre-filled result to other sub-load balancing managers includes: Sending a decoding request message to the central load balancing manager, for the central load balancing manager to select a target sub-load balancing manager from all the other sub-load balancers according to the decoding request message, and returning the address information of the target sub-load balancing manager; Receiving the address information returned by the central load balancing manager; The pre-filled result is forwarded to the location indicated by the address information.
3. The scheduling method of the inference model according to claim 1, characterized in that: The scheduling the local decoding node to perform decoding processing on the pre-filled result includes: Scheduling the idle local decoding node to perform decoding processing on the pre-filled result; or, Scheduling the local decoding node with the most idle resources to perform decoding processing on the pre-filled result; or, The local decoding node with the lowest resource utilization is scheduled to perform decoding processing on the pre-filled result.
4. The scheduling method of the inference model according to any one of claims 1 to 3, characterized in that: When supporting decoding processing of the pre-filled result, the local decoding node satisfies at least one of the following conditions: there is an idle local decoding node, there is a local decoding node whose resource utilization is less than a first threshold, there is a local decoding node whose idle resource amount is greater than a second threshold, there is a local decoding node whose resource utilization is less than the resource utilization of the local decoding nodes corresponding to other sub-load balancing managers, and there is a local decoding node whose resource amount is less than the resource amount of the local decoding nodes corresponding to other sub-load balancing managers; wherein the first threshold and / or the second threshold is: a preset value, or a parameter associated with the data volume of the pre-filled result.
5. A scheduling method for an inference model, characterized in that: Applied to a central load balancing manager, the method includes: Receiving an inference request sent by a user terminal; Forwarding the inference request to a sub-load balancing manager, so that the sub-load balancing manager that receives the inference request can, if the local decoding node supports decoding, schedule the local decoding node to decode a pre-filled result obtained by scheduling the local pre-filled node to process the inference request, or, if the local decoding node does not support decoding of the pre-filled result, forwarding the pre-filled result to another sub-load balancing manager, so that the other sub-load balancing manager can schedule a remote decoding node to decode the pre-filled result; Among them, the local pre-filling node is a pre-filling node deployed on the same machine as the sub-load balancing manager, the local decoding node is a decoding node deployed on the same machine as the sub-load balancing manager, and the remote decoding node is a decoding node deployed on a different machine from the sub-load balancing manager.
6. The scheduling method of the inference model according to claim 5, characterized in that: Forwarding the inference request to a sub-load balancing manager includes: Determining a target load balancing manager from all the sub-load balancing managers based on the status information of the local pre-filling node and the status information of the local decoding node of each sub-load balancing manager; wherein the status information is used to indicate at least one of the following node conditions: load condition, queue condition, and working status; The inference request is forwarded to the target load balancing manager.
7. The scheduling method of the inference model according to claim 6, characterized in that: Before determining the target load balancing manager from all the sub-load balancing managers based on the status information of the local pre-filling node and the status information of the local decoding node of each sub-load balancing manager, the method further includes: Receive the status information reported by the local pre-filling node and the local decoding node of each sub-load balancing manager.
8. A scheduling system, characterized in that: include: Sub-load balancing manager and central load balancing manager; Among them, different first load balancing managers are deployed on different machines, and each machine deployed by the first load balancing manager is also deployed with at least one pre-filled node and at least one decoding node. The sub-load balancing manager is used to execute the scheduling method of the inference model as described in any one of claims 1 to 4, and the central load balancing manager is used to execute the scheduling method of the inference model as described in any one of claims 5 to 7.
9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the scheduling method of the inference model as described in any one of claims 1 to 4, or execute the scheduling method of the inference model as described in any one of claims 5 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the scheduling method of the inference model according to any one of claims 1 to 4, or implements the scheduling method of the inference model according to any one of claims 5 to 7.
Citation Information
Cited By
Data processing method and device, electronic equipment and storage medium
CN120929221A