Reasoning rescheduling method, apparatus, device, medium, and computer program product

CN118798355BActive Publication Date: 2026-09-29CHINA MOBILE GROUP DESIGN INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410723443.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2026-09-29
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

[0004]本发明提供一种推理重调度方法、装置、设备、介质及计算机程序产品,用以解决现有推理请求的重调度方案存在的推理请求重调度性能低的技术问题,提升了推理重调度的整体性能

Benefits of technology

[0015]本发明提供的推理重调度方法、装置、设备、介质及计算机程序产品,在推理调度器与服务注册中心合设的情况下,基于每个推理服务器上报的负载状态信息,确定待重调度的第一推理服务器以及接收重调度的目标推理服务器;在推理调度器与推理服务器的流量网关合设的情况下,基于推理调度器记录的每个推理服务器的请求处理信息,确定待重调度的第一推理服务器以及接收重调度的目标推理服务器。最终基于第一推理服务器的请求排队时间信息以及目标推理服务器的请求数量信息,将第一推理服务器的待重调度请求发送给目标推理服务器。本申请解决了目前的技术方案只能在单个推理服务器进行推理请求重调度的问题,实现了运行同一个AI模型的多个推理服务器之间协同进行推理请求重调度,将排队时间过长的推理服务器中的推理请求重新调度到排队时间较短的推理服务器中,从而提升推理重调度的整体性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118798355B_ABST
    Figure CN118798355B_ABST
Patent Text Reader

Abstract

The application provides a reasoning rescheduling method, device and equipment and a readable storage medium. The application is applied to a reasoning scheduling system including a reasoning scheduler and multiple reasoning servers. The method comprises the following steps: in the case that the reasoning scheduler and a service registration center are combined, based on the load state information reported by each reasoning server, determining a first reasoning server to be rescheduled and a target reasoning server to receive rescheduling; in the case that the reasoning scheduler and a traffic gateway of the reasoning server are combined, based on the request processing information of each reasoning server recorded by the reasoning scheduler, determining the first reasoning server to be rescheduled and the target reasoning server to receive rescheduling; based on the request queuing time information of the first reasoning server and the request quantity information of the target reasoning server, sending the rescheduled request of the first reasoning server to the target reasoning server. The application improves the overall performance of reasoning rescheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing power scheduling technology, and in particular to a reasoning rescheduling method, apparatus, device, medium, and computer program product. Background Technology

[0002] Existing rescheduling methods for Artificial Intelligence (AI) model inference only reschedule inference requests that are queued within the same inference server. The inference server only reschedules requests with excessively long queue times within its own server, without being aware of the queuing status of other inference servers, even if those servers are also running the same AI model.

[0003] When the volume of AI model inference requests grows to a certain scale, multiple inference servers are often needed to provide inference services for a single AI model simultaneously. This multi-server inference can lead to uneven load distribution among the servers. However, existing internal rescheduling methods for individual inference servers cannot solve this load imbalance problem. This results in inference requests on high-loaded servers taking longer to process than those on low-loaded servers, impacting user experience. Existing solutions adjust the proportion of inference requests across servers using load balancers, gateways, and service registration and discovery centers, but they cannot reschedule already forwarded or assigned requests. Long-queued inference requests on high-loaded servers still cannot be processed promptly, further affecting the user experience. Summary of the Invention

[0004] This invention provides a method, apparatus, device, medium, and computer program product for inference rescheduling, which solves the technical problem of low performance in existing inference request rescheduling schemes and improves the overall performance of inference rescheduling.

[0005] This invention provides an inference rescheduling method applied to an inference scheduling system, the inference scheduling system including an inference scheduler and multiple inference servers; the method includes: In the case where the inference scheduler and the service registry are co-located, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined based on the load status information reported by each inference server. When the inference scheduler and the inference server's traffic gateway are co-located, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined based on the request processing information of each inference server recorded by the inference scheduler. Based on the request queuing time information of the first inference server and the request quantity information of the target inference server, the rescheduling request of the first inference server is sent to the target inference server.

[0006] According to the inference rescheduling method provided by the present invention, determining the first inference server to be rescheduled and the target inference server to receive the rescheduling based on the load status information reported by each inference server includes: Based on the queuing data reported by each of the inference servers, the request queuing time at the tail of the request queue of each of the inference servers is calculated. Based on the comparison between the request queuing time and the preset waiting time, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined among the multiple inference servers.

[0007] According to the inference rescheduling method provided by the present invention, determining the first inference server to be rescheduled and the target inference server to receive the rescheduling based on the request processing information of each inference server recorded by the inference scheduler includes: Based on the request volume and inference result return volume of each inference server recorded by the inference scheduler, the request queuing time at the tail of the request queue of each inference server is calculated. Based on the comparison between the request queuing time and the preset waiting time, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined among the multiple inference servers.

[0008] According to a method for inference rescheduling provided by the present invention, the step of sending the rescheduling request of the first inference server to the target inference server based on the request queuing time information of the first inference server and the request quantity information of the target inference server includes: The number of rescheduling requests for the first inference server is determined based on the number of queue requests and inference throughput of the first inference server and the preset waiting time. Determine the tail request queuing time of the target inference server; The first ranking result of all the first inference servers is determined based on the number of rescheduling requests, and the target ranking result of all the target inference servers is determined based on the tail request queuing time. Based on the first sorting result and the target sorting result, the rescheduling request of the first inference server is sent to the target inference server.

[0009] According to the inference rescheduling method provided by the present invention, determining the number of rescheduling requests for the first inference server includes: Based on the request volume and inference result return volume of each inference server recorded by the inference scheduler, the number of rescheduling requests for each inference server is calculated.

[0010] According to a method for inference rescheduling provided by the present invention, the step of sending a rescheduling request from the first inference server to the target inference server based on the first sorting result and the target sorting result includes: The first sorting result is determined based on the order of the number of rescheduling requests from largest to smallest, and the target sorting result is determined based on the order of the queuing time of the tail requests from smallest to largest. The request for rescheduling of the first inference server is sent to the target inference server; the sorting position of the first inference server in the first sorting result corresponds to the sorting position of the target inference server in the target sorting result.

[0011] The present invention also provides an inference rescheduling device, comprising the following modules: The first scheduling determination module is used to determine the first inference server to be rescheduled and the target inference server to receive rescheduling based on the load status information reported by each inference server when the inference scheduler and the service registry are co-located. The second scheduling determination module is used to determine the first inference server to be rescheduled and the target inference server to receive rescheduling based on the request processing information of each inference server recorded by the inference scheduler when the inference scheduler and the traffic gateway of the inference server are co-located. The rescheduling request sending module is used to send the rescheduling request of the first inference server to the target inference server based on the request queuing time information of the first inference server and the request quantity information of the target inference server.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the inference rescheduling method as described above.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the inference rescheduling method as described above.

[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the inference rescheduling method as described above.

[0015] The inference rescheduling method, apparatus, device, medium, and computer program product provided by this invention, when the inference scheduler and service registry are co-located, determines the first inference server to be rescheduled and the target inference server to receive the rescheduling based on the load status information reported by each inference server; when the inference scheduler and the traffic gateway of the inference server are co-located, it determines the first inference server to be rescheduled and the target inference server to receive the rescheduling based on the request processing information of each inference server recorded by the inference scheduler. Finally, based on the request queuing time information of the first inference server and the request quantity information of the target inference server, the rescheduling request of the first inference server is sent to the target inference server. This application solves the problem that the current technical solution can only perform inference request rescheduling on a single inference server, and realizes collaborative inference request rescheduling among multiple inference servers running the same AI model, rescheduling inference requests in inference servers with long queuing times to inference servers with shorter queuing times, thereby improving the overall performance of inference rescheduling. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is one of the flowcharts of the inference rescheduling method provided by the present invention.

[0018] Figure 2 This is the second flowchart of the inference rescheduling method provided by the present invention.

[0019] Figure 3 This is a schematic diagram of the inference rescheduling device provided by the present invention.

[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] The following is combined Figures 1-4 This invention describes the inference rescheduling method, apparatus, device, medium, and computer program product.

[0023] Reference Figure 1 , Figure 1 This is one of the flowcharts illustrating the inference rescheduling method in this application. The inference rescheduling method provided in this application is applied to an inference scheduling system, which includes an inference scheduler and multiple inference servers; the inference rescheduling method provided in this application may include: Step 100: In the case where the inference scheduler and the service registry are co-located, based on the load status information reported by each inference server, determine the first inference server to be rescheduled and the target inference server to receive the rescheduling. Specifically, the inference rescheduling method provided in this application is applied to an inference scheduling system, which includes an inference scheduler and multiple inference servers. The inference scheduler is responsible for selecting an inference server for each inference request and rescheduling the inference requests based on the queuing status of the inference requests on each inference server. The inference servers are responsible for executing AI model inference and cooperate with the inference scheduler to complete the rescheduling of inference requests.

[0024] This application provides two schemes for inference scheduler and inference server to work together to implement inference rescheduling. Scheme 1 is as follows: The inference scheduler and service registry are co-located, responsible for providing service registration and discovery services to inference servers. This first scheme has lower forwarding requirements for the inference scheduler and is suitable for high-traffic AI inference scenarios such as video. During inference rescheduling, the inference scheduler receives request queuing information (i.e., load status information in this embodiment) from registered inference servers and reschedules some inference requests with excessively long queue times based on the request queuing information of all registered inference servers. Specifically, it determines the first inference server requiring rescheduling and the target inference server to receive the rescheduled inference request based on the request queuing information of all registered inference servers.

[0025] Step 200: In the case where the inference scheduler and the inference server's traffic gateway are co-located, based on the request processing information of each inference server recorded by the inference scheduler, determine the first inference server to be rescheduled and the target inference server to receive the rescheduling. Specifically, the second scheme for implementing inference rescheduling in collaboration between the inference scheduler and the inference server is as follows: In Scheme 2, the inference scheduler and the traffic gateway for the inference server are co-located. This scheme makes the business address of the inference server invisible to AI inference requesters, thereby improving security and making it more suitable for AI inference scenarios targeting public users. In this scheme, the inference scheduler records the inference sending time and inference result receiving time of all inference requests to analyze the request queuing situation of each inference server. Inference requests with excessively long queue times are rescheduled. Specifically, the inference scheduler calculates and determines whether any inference server needs to be rescheduled based on the recorded request sending volume and inference result return volume of each inference server, and then determines the first inference server that needs to be rescheduled. The inference scheduler sorts all currently available inference servers that can receive scheduling requests in ascending order of their tail-end request queuing time, and selects the k inference servers with the shortest tail-end request queuing time as the target inference servers that can accept the rescheduling request.

[0026] Step 300: Based on the request queuing time information of the first inference server and the request quantity information of the target inference server, send the rescheduling request of the first inference server to the target inference server.

[0027] Specifically, the inference scheduler sorts the k first inference servers that need rescheduling in descending order of the number of requests requiring rescheduling; similarly, it sorts the target inference servers in ascending order of the queuing time of the requests at the end of their queues. The first inference servers with the same sequence number in both sorts are associated with the target inference servers. Specifically, the inference scheduler sends a rescheduling instruction to the inference server ranked n (1≤n≤k) in terms of the number of requests among the first inference servers that need rescheduling. This instruction includes the number of requests that need rescheduling on that inference server, sending these requests in batches to the inference scheduler. Simultaneously, the inference scheduler sends the same instruction to the target inference servers ranked n (1≤n≤k) in terms of queuing time, notifying them to prepare to receive rescheduling requests. This process continues until all rescheduling requests from all first inference servers have been sent to the target inference servers, completing the inference rescheduling.

[0028] In this embodiment, when the inference scheduler and service registry are co-located, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined based on the load status information reported by each inference server. When the inference scheduler and the traffic gateway of the inference server are co-located, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined based on the request processing information of each inference server recorded by the inference scheduler. Finally, based on the request queuing time information of the first inference server and the request quantity information of the target inference server, the rescheduling request of the first inference server is sent to the target inference server. This application solves the problem that the current technical solution can only perform inference request rescheduling on a single inference server, and realizes collaborative inference request rescheduling among multiple inference servers running the same AI model, rescheduling inference requests in inference servers with long queuing times to inference servers with shorter queuing times, thereby improving the overall performance of inference rescheduling.

[0029] In one embodiment, the inference rescheduling method provided in this application may further include: Step 110: Based on the queuing data reported by each inference server, calculate the request queuing time at the tail of the request queue of each inference server. Step 120: Based on the comparison between the request queuing time and the preset waiting time, determine the first inference server to be rescheduled and the target inference server to receive the rescheduling among the multiple inference servers.

[0030] Specifically, the details of Option 1 are as follows: The inference server initiates a service registration request to the inference scheduler, which registers the server and returns a successful registration response. After successful registration, the inference server sends its request queuing information to the inference scheduler at a set time frequency. This request queuing information includes parameters such as the number of requests M in the current request queue, the current number of requests received per second N, and the current inference throughput per second Q.

[0031] The inference scheduler determines whether any inference servers need to be rescheduled based on the load status information reported by each inference server. Specifically, the inference scheduler calculates the request queuing time T1 at the tail of the request queue of the inference server based on the request queuing data reported by the inference server, as shown in Formula 1.

[0032] X = (M - QTmax) × a; (2) The calculated queuing time is updated in real time based on the data reported by the inference server, thus accurately reflecting the tail queuing time of the inference server at each moment. If the queuing time exceeds the preset maximum waiting time (i.e., the preset waiting time in this embodiment), it is determined that the request queue of that inference server needs to be rescheduled. The inference scheduler traverses the request queuing data of all inference servers and performs calculations to obtain the number k of inference servers that need to be rescheduled.

[0033] For the first inference server determined to require inference request rescheduling, the inference scheduler calculates the number of requests X that need to be rescheduled. The inference scheduler calculates this based on data such as the set maximum waiting time Tmax, the number of requests in the queue M, and the inference throughput Q per second, as shown in Formula 2, where a is an adjustable coefficient.

[0034] The inference scheduler sorts all currently available inference servers that can receive scheduling requests in ascending order of the request queuing time at the tail of the queue, and selects the k inference servers with the shortest request queuing time at the tail of the queue as the target inference servers that can accept rescheduling requests.

[0035] This embodiment calculates the request queuing time at the tail of the request queue of each inference server by using the queuing data reported by each inference server, and then provides a scheme for determining the first inference server to be rescheduled and the target inference server to receive the rescheduling.

[0036] In one embodiment, the inference rescheduling method provided in this application may further include: Step 210: Based on the request sending volume and inference result return volume of each inference server recorded by the inference scheduler, calculate the request queuing time at the tail of the request queue of each inference server. Step 220: Based on the comparison between the request queuing time and the preset waiting time, determine the first inference server to be rescheduled and the target inference server to receive the rescheduling among the multiple inference servers.

[0037] Specifically, in Scheme 2 above, all inference request traffic passes through the inference scheduler and is distributed by the scheduler. Therefore, the inference scheduler can record the number of inference requests sent to each inference server and the number of inference results returned at each moment. Using the number of requests sent to each inference server and the number of inference results returned at each moment, the inference scheduler calculates the tail-end request queuing time and the number of requests requiring rescheduling for each inference server, without needing to obtain request queuing data from each inference server.

[0038] At time T, after the inference server starts running, the queuing time and number of queued requests for the tail request are shown in Formulas 3 and 4, where b is the queuing time for the tail request at time T; c is the number of requests sent at time t; d is the number of results returned at time t; and e is the number of queued requests at time T.

[0039] The inference scheduler calculates and determines whether a first inference server needs to be rescheduled based on the locally recorded request volume and inference result return volume of each inference server. Specifically, the inference scheduler uses the request volume and inference result return volume sent to each inference server at each moment to calculate the tail request queuing time of each inference server, without needing to obtain request queuing data reported by each inference server.

[0040] The calculated queuing time is continuously updated over time to accurately reflect the tail-end queuing time at each moment. If the queuing time exceeds the preset waiting time, it is determined that the request queue of that inference server needs to be rescheduled. The inference scheduler traverses all inference server queuing information data and performs calculations to obtain the number k of inference servers that need to be rescheduled.

[0041] For the first inference server, the inference scheduler calculates the number of requests that need to be rescheduled. Specifically, the inference scheduler uses the number of requests sent to each inference server at each time step and the number of inference results returned to calculate the number of requests that need to be rescheduled, without needing to have each inference server report the request queue data. As shown in Formula 5, where f is the number of results returned at time T.

[0042] The inference scheduler sorts all currently available inference servers that can receive scheduling requests in ascending order of the request queuing time at the tail of the queue, and selects the k inference servers with the shortest request queuing time at the tail of the queue as the target inference servers that can accept rescheduling requests.

[0043] This embodiment calculates the request queuing time at the tail of the request queue of each inference server by recording the request volume and inference result return volume of each inference server by the inference scheduler, and then provides another scheme for determining the first inference server to be rescheduled and the target inference server to receive the rescheduling.

[0044] Reference Figure 2 , Figure 2 This is a second flowchart illustrating the inference rescheduling method in this application. In one embodiment, the inference rescheduling method provided in this application may further include: Step 310: Determine the number of rescheduling requests for the first inference server based on the number of queue requests and inference throughput of the first inference server and the preset waiting time. Step 320: Determine the tail request queuing time of the target inference server; Step 330: Determine the first ranking result of all the first inference servers based on the number of rescheduling requests, and determine the target ranking result of all the target inference servers based on the tail request queuing time. Step 340: Based on the first sorting result and the target sorting result, send the rescheduling request of the first inference server to the target inference server.

[0045] The inference rescheduling method provided in this application embodiment may further include: Step 311: Calculate the number of rescheduling requests for each inference server based on the request volume and inference result return volume of each inference server recorded by the inference scheduler.

[0046] The inference rescheduling method provided in this application embodiment may further include: Step 341: Determine that the first sorting result is obtained based on the order of the number of rescheduling requests from largest to smallest, and determine that the target sorting result is obtained based on the order of the queuing time of the tail requests from smallest to largest; Step 342: Send the rescheduling request of the first inference server to the target inference server; the sorting position of the first inference server in the first sorting result corresponds to the sorting position of the target inference server in the target sorting result.

[0047] Specifically, in Schemes 1 and 2 above, the inference scheduler associates inference servers with the same sequence number in the two sorts (first sort result and target sort result). Specifically, the inference scheduler sends an instruction to the inference server ranked n (1≤n≤k) in the target inference server group, notifying it to receive the rescheduling request sent by the inference server ranked n in the number of requests in the first inference server group. Similarly, the inference scheduler sends a rescheduling instruction to the inference server ranked n in the number of requests in the first inference server group. This instruction includes the IP address of the inference server ranked n in the target inference server group and the number of requests that need to be rescheduled on the first inference server group, instructing it to send the requests that need to be rescheduled on the first inference server group to the inference server ranked n in the queue time in the target inference server group in a batch manner.

[0048] Each first inference server that needs to be rescheduled sends the rescheduled requests on its server to the associated target inference server in batches according to the rescheduling instructions of the inference scheduler. Each target inference server receives the rescheduling requests, puts these rescheduling requests into the rescheduling queue, and performs priority scheduling when the opportunity arises.

[0049] This embodiment improves the overall performance of inference rescheduling by using a multi-inference server collaborative inference rescheduling method.

[0050] The inference rescheduling apparatus provided by the present invention is described below. The inference rescheduling apparatus described below can be referred to in correspondence with the inference rescheduling method described above.

[0051] Please refer to Figure 3 The present invention also provides a reasoning rescheduling apparatus, comprising: The first rescheduling determination module 301 is used to determine the first inference server to be rescheduled and the target inference server to receive rescheduling based on the load status information reported by each inference server when the inference scheduler and the service registry are co-located. The second rescheduling determination module 302 is used to determine the first inference server to be rescheduled and the target inference server to receive rescheduling based on the request processing information of each inference server recorded by the inference scheduler when the inference scheduler and the inference server are co-located. The rescheduling request sending module 303 is used to send the rescheduling request of the first inference server to the target inference server based on the request queuing time information of the first inference server and the request quantity information of the target inference server.

[0052] Optionally, the first scheduling determination module includes: The first request queuing time calculation unit is used to calculate the request queuing time at the tail of the request queue of each inference server based on the queuing data reported by each inference server. The first rescheduling determination unit is used to determine, based on the comparison result of the request queuing time and the preset waiting time, the first inference server to be rescheduled among the multiple inference servers and the target inference server to receive the rescheduling.

[0053] Optionally, the second scheduling determination module includes: The second request queuing time calculation unit is used to calculate the request queuing time at the tail of the request queue of each inference server based on the request sending volume and inference result return volume of each inference server recorded by the inference scheduler. The second scheduling determination unit is used to determine the first inference server to be rescheduled and the target inference server to receive rescheduling among the multiple inference servers based on the comparison result of the request queuing time and the preset waiting time.

[0054] Optionally, the rescheduling request sending module includes: The rescheduling request quantity determination unit is used to determine the rescheduling request quantity of the first inference server based on the queue request quantity and inference throughput of the first inference server and the preset waiting time. The tail request queuing time determination unit is used to determine the tail request queuing time of the target inference server; The sorting result determination unit is used to determine the first sorting result of all the first inference servers based on the number of rescheduling requests, and to determine the target sorting result of all the target inference servers based on the tail request queuing time. The rescheduling request sending unit is used to send the rescheduling request of the first inference server to the target inference server based on the first sorting result and the target sorting result.

[0055] Optionally, the rescheduling request number determination unit includes: The rescheduling request quantity calculation unit is used to calculate the rescheduling request quantity of each inference server based on the request sending quantity and inference result return quantity of each inference server recorded by the inference scheduler.

[0056] Optionally, the rescheduling request sending unit includes: The determining unit is configured to determine that the first sorting result is obtained based on the order of the number of rescheduling requests from largest to smallest, and to determine that the target sorting result is obtained based on the order of the queuing time of the tail requests from smallest to largest; The request sending unit is used to send the rescheduling request of the first inference server to the target inference server; the sorting position of the first inference server in the first sorting result corresponds to the sorting position of the target inference server in the target sorting result.

[0057] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an inference rescheduling method, which includes: when the inference scheduler and the service registry are co-located, determining a first inference server to be rescheduled and a target inference server to receive the rescheduling based on the load status information reported by each inference server; when the inference scheduler and the traffic gateway of the inference server are co-located, determining a first inference server to be rescheduled and a target inference server to receive the rescheduling based on the request processing information of each inference server recorded by the inference scheduler; and sending the rescheduling request of the first inference server to the target inference server based on the request queuing time information of the first inference server and the request quantity information of the target inference server.

[0058] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0059] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the inference rescheduling method provided by the above methods. The method includes: when the inference scheduler and the service registry are co-located, determining a first inference server to be rescheduled and a target inference server to receive the rescheduling based on the load status information reported by each inference server; when the inference scheduler and the traffic gateway of the inference server are co-located, determining a first inference server to be rescheduled and a target inference server to receive the rescheduling based on the request processing information of each inference server recorded by the inference scheduler; and sending the rescheduling request of the first inference server to the target inference server based on the request queuing time information of the first inference server and the request quantity information of the target inference server.

[0060] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the inference rescheduling method provided by the above methods. The method includes: when the inference scheduler and service registry are co-located, determining a first inference server to be rescheduled and a target inference server to receive the rescheduling based on load status information reported by each of the inference servers; when the inference scheduler and the traffic gateway of the inference servers are co-located, determining a first inference server to be rescheduled and a target inference server to receive the rescheduling based on request processing information of each of the inference servers recorded by the inference scheduler; and sending a rescheduling request from the first inference server to the target inference server based on request queuing time information of the first inference server and request quantity information of the target inference server.

[0061] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0062] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A reasoning rescheduling method, characterized in that, The method is applied to an inference scheduling system, which includes an inference scheduler and multiple inference servers; the method includes: In the case where the inference scheduler and the service registry are co-located, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined based on the load status information reported by each inference server. When the inference scheduler and the inference server's traffic gateway are co-located, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined based on the request processing information of each inference server recorded by the inference scheduler. Based on the request queuing time information of the first inference server and the request quantity information of the target inference server, the rescheduling request of the first inference server is sent to the target inference server. Sending the request to be rescheduled from the first inference server to the target inference server based on the request queuing time information of the first inference server and the request quantity information of the target inference server includes: The number of rescheduling requests for the first inference server is determined based on the number of queue requests and inference throughput of the first inference server and the preset waiting time. Determine the tail request queuing time of the target inference server; The first ranking result of all the first inference servers is determined based on the number of rescheduling requests, and the target ranking result of all the target inference servers is determined based on the tail request queuing time. Based on the first sorting result and the target sorting result, the rescheduling request of the first inference server is sent to the target inference server.

2. The inference rescheduling method according to claim 1, characterized in that, The step of determining the first inference server to be rescheduled and the target inference server to receive the rescheduling based on the load status information reported by each of the inference servers includes: Based on the queuing data reported by each of the inference servers, the request queuing time at the tail of the request queue of each of the inference servers is calculated. Based on the comparison between the request queuing time and the preset waiting time, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined among the multiple inference servers.

3. The inference rescheduling method according to claim 2, characterized in that, The step of determining the first inference server to be rescheduled and the target inference server to receive the rescheduling based on the request processing information of each inference server recorded by the inference scheduler includes: Based on the request volume and inference result return volume of each inference server recorded by the inference scheduler, the request queuing time at the tail of the request queue of each inference server is calculated. Based on the comparison between the request queuing time and the preset waiting time, the first inference server to be rescheduled and the target inference server to receive the rescheduling are determined among the multiple inference servers.

4. The inference rescheduling method according to claim 1, characterized in that, Determining the number of rescheduling requests for the first inference server includes: Based on the request volume and inference result return volume of each inference server recorded by the inference scheduler, the number of rescheduling requests for each inference server is calculated.

5. The inference rescheduling method according to claim 1, characterized in that, Sending the rescheduling request from the first inference server to the target inference server based on the first sorting result and the target sorting result includes: The first sorting result is determined based on the order of the number of rescheduling requests from largest to smallest, and the target sorting result is determined based on the order of the queuing time of the tail requests from smallest to largest. The request for rescheduling of the first inference server is sent to the target inference server; the sorting position of the first inference server in the first sorting result corresponds to the sorting position of the target inference server in the target sorting result.

6. A reasoning rescheduling device, characterized in that, include: The first scheduling determination module is used to determine the first inference server to be rescheduled and the target inference server to receive the rescheduling, based on the load status information reported by each inference server, when the inference scheduler and the service registry are co-located. The second rescheduling determination module is used to determine the first inference server to be rescheduled and the target inference server to receive rescheduling, based on the request processing information of each inference server recorded by the inference scheduler, when the inference scheduler and the inference server are co-located. The rescheduling request sending module is used to send the rescheduling request of the first inference server to the target inference server based on the request queuing time information of the first inference server and the request quantity information of the target inference server. The module for sending requests to be rescheduled is specifically used for: The number of rescheduling requests for the first inference server is determined based on the number of queue requests and inference throughput of the first inference server and the preset waiting time. Determine the tail request queuing time of the target inference server; The first ranking result of all the first inference servers is determined based on the number of rescheduling requests, and the target ranking result of all the target inference servers is determined based on the tail request queuing time. Based on the first sorting result and the target sorting result, the rescheduling request of the first inference server is sent to the target inference server.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the inference rescheduling method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the inference rescheduling method as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the inference rescheduling method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Job scheduling method, server and server cluster

    CN116233022A