Scheduling method and device of reasoning request, electronic equipment and readable storage medium
By receiving and calculating the load level of the inference service and using the corresponding scheduling strategy to schedule inference requests, the problems of load imbalance and time extension in large-scale inference services are solved, and the processing efficiency and accuracy of inference requests are improved.
Patent Information
- Application Number
- CN202510820117.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-10
AI Technical Summary
In large-scale inference request scheduling, existing technologies have problems such as large differences in load levels between different inference services and long average delays in inference results.
By receiving the indicator metadata sent by multiple inference services, the load level of each inference service is calculated, and the corresponding scheduling strategy is adopted according to the load level to schedule inference requests to the target inference service, including strategies such as immediate processing of low load, delayed processing of high load, and appropriate waiting for balanced load.
It achieves load balancing of inference requests, reduces average latency, and improves the processing efficiency and accuracy of inference services.
Smart Images

Figure CN120768951A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technology, particularly to artificial intelligence technologies such as large models, deep learning, and cloud services. A method, device, electronic device, and readable storage medium for scheduling inference requests are provided. Background Art
[0002] As a fundamental breakthrough in the field of artificial intelligence, large models have achieved a qualitative change in cognitive capabilities through their architectural innovation and scale. The general intelligence capabilities demonstrated by large models have important practical significance and influence, and are a milestone in the development of artificial intelligence.
[0003] In the field of large-model inference link construction, different modes are usually adopted depending on the concurrent scale of inference requests: when the scale is relatively small, the push mode is usually used, and inference requests are directly pushed to the inference service using strategies such as polling or minimum number of connections; when the scale is relatively large, the pull mode is usually used, and inference requests are first placed in the queue, and the inference service then pulls the requests from the queue for inference.
[0004] Existing technologies typically use a pull model for large-scale inference, where different inference services directly pull inference requests from a queue and process them. When there are a large number of inference services, existing inference request scheduling methods suffer from significant load variations between inference services and high average latency for inference results. Summary of the Invention
[0005] According to a first aspect of the present disclosure, a method for scheduling inference requests is provided, comprising: receiving indicator metadata respectively sent by a plurality of inference services; obtaining a load level of each inference service based on the indicator metadata; and in response to receiving a processing request sent by a target inference service, adopting a scheduling strategy corresponding to the load level of the target inference service to schedule the inference request to the target inference service.
[0006] According to a second aspect of the present disclosure, a scheduling device for inference requests is provided, including: a receiving unit for receiving indicator metadata respectively sent by multiple inference services; a processing unit for obtaining the load level of each inference service based on the indicator metadata; and a scheduling unit for scheduling the inference request to the target inference service in response to receiving a processing request sent by the target inference service, adopting a scheduling strategy corresponding to the load level of the target inference service.
[0007] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0008] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method as described above.
[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0013] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0014] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0015] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0016] Figure 5 The block diagram is a block diagram of an electronic device for implementing the method for scheduling inference requests according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, and various details of the embodiments of the present disclosure are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and mechanisms are omitted in the following description.
[0018] Figure 1 Schematic diagram of the first embodiment of the present disclosure. Figure 1 As shown, the method for scheduling inference requests in this embodiment specifically includes the following steps:
[0019] S101, receiving indicator metadata respectively sent by multiple inference services;
[0020] S102. Obtaining a load level of each inference service based on the indicator metadata;
[0021] S103 : In response to receiving the processing request sent by the target reasoning service, use a scheduling strategy corresponding to the load level of the target reasoning service to schedule the reasoning request to the target reasoning service.
[0022] The scheduling method for inference requests in this embodiment, on the one hand, obtains the load level of each inference service based on the indicator metadata sent by multiple inference services respectively, which can improve the accuracy of the obtained load level; on the other hand, it determines the scheduling strategy corresponding to the target inference service in combination with the obtained load level, and then uses the scheduling strategy to schedule the inference request to the target inference service, ensuring that the scheduling process of the inference request matches the load level of the target inference service, which can improve the scheduling accuracy of the inference request, thereby making the load of multiple inference services more balanced, and effectively reducing the average delay in obtaining the inference result.
[0023] The execution subject of the method for scheduling inference requests in this embodiment is a queue server, which is used to schedule inference requests in the queue to different inference services.
[0024] In this embodiment, the inference service is a computing service that uses a pre-trained machine learning model (e.g., a large model) to process input data (the input data corresponds to the inference request) to output the corresponding inference result; wherein, the inference service can specifically refer to a machine learning model deployed on a server or a server cluster, which can receive and process data corresponding to the inference request to obtain the inference result.
[0025] The multiple inference services in this embodiment use machine learning models with the same model type when performing inference, for example, they are all Qianwen models or Wenxin models, etc.; therefore, the multiple inference services in this embodiment are used to process inference requests corresponding to the same model type.
[0026] In this embodiment, the indicator metadata received by executing S101 and sent by different inference services include three types: first time consumption data, second time consumption data and memory utilization data. Each type of indicator metadata is obtained in real time by a different inference service.
[0027] In this embodiment, the first time consumption data is the average time consumption data of the first token, the token is the semantic unit of the reasoning result, and the first token is the first semantic unit of the reasoning result; specifically, the reasoning service in this embodiment calculates the average value based on the first token time consumption of the reasoning results of all inference requests within the preset time window (the first token time consumption is the time required for the reasoning service to obtain the first token), and then uses the calculation result as the first time consumption data of the reasoning service.
[0028] In this embodiment, the second time consumption data is the average time consumption of the inter-package token, and the inter-package token refers to other tokens in the inference result except the first token, that is, other semantic units in the inference result except the first semantic unit; specifically, the inference service in this embodiment calculates the average value based on the inter-package token time consumption of the inference results of all inference requests within the preset time window (the inter-package token time consumption is the time required for the inference service to obtain other tokens), and then uses the calculation result as the second time consumption data of the inference service.
[0029] In this embodiment, the video memory utilization data is the video memory utilization acquired in real time by the inference service.
[0030] The preset time window in this embodiment can be determined based on the current time and the preset duration. For example, the preset duration before the current time is defined as the preset time window, and the preset duration can be 2s, 3s, 5s, etc.
[0031] That is to say, different inference services in this embodiment will obtain different types of indicator metadata through real-time calculation or real-time acquisition, and then report the obtained indicator metadata to the queue server, so that the queue server can determine the current load level of each inference service based on the received indicator metadata corresponding to different inference services.
[0032] In addition, this embodiment selects at least one of the first time consumption data, the second time consumption data and the video memory utilization data as the indicator metadata, so that the selected indicator metadata can more truly and accurately reflect the load situation of the inference service, thereby improving the accuracy of the obtained load level.
[0033] In this embodiment, the queue server may update the indicator metadata reported by different reasoning services according to the unique identification information of the reasoning services.
[0034] In this embodiment, after executing S101 to receive indicator metadata respectively sent by multiple inference services, executing S102 to obtain the load level of each inference service according to the received indicator metadata.
[0035] In this embodiment, the load level obtained by executing S102 is one of the first load level, the second load level, and the third load level.
[0036] In this embodiment, the first load level indicates that the load of the inference service is relatively low during the current operation, and the first load level is the low load level; the second load level indicates that the load of the inference service is relatively high during the current operation, and the second load level is the high load level; the third load level indicates that the load of the inference service is relatively balanced during the current operation, and the third load level is the balanced load level.
[0037] Specifically, when executing S102 in this embodiment to obtain the load level of each inference service based on the received indicator metadata, the implementation method that can be adopted is: obtaining the load indicator of each inference service based on the preset data type and indicator metadata; obtaining the load level of each inference service based on the load indicators of all inference services.
[0038] That is to say, on the one hand, this embodiment obtains the load indicator of each inference service based on the preset data type and indicator metadata, so that the obtained load indicator corresponds to the preset data type, thereby improving the accuracy of the load indicator. On the other hand, it combines the load indicators of all inference services to obtain the load level of each inference service, which can improve the accuracy of the obtained load level.
[0039] When executing S102, this embodiment can compare the load indicator of each inference service with the load indicator statistical results based on the load indicator statistical results of all inference services (such as the load indicator average or load indicator variance). If the load indicator is greater than or equal to the load indicator statistical results, the load level of the inference service is determined to be the second load level, otherwise it is the first load level.
[0040] In this embodiment, the preset data type can be determined based on the service goal emphasized by the reasoning service; for example, if the service goal emphasized by the reasoning service is token quality, the preset data type in this embodiment is the first time-consuming data; if the service goal of the reasoning service focuses on overall performance, the preset data type in this embodiment is the second time-consuming data; if the service goal of the reasoning service has no particular emphasis, the preset data type in this embodiment may include the first time-consuming data, the second time-consuming data and video memory utilization data, and may also include the first time-consuming data and the second time-consuming data, etc.
[0041] In the implementation of obtaining the load indicator of each inference service according to the preset data type and the indicator metadata in S102, the following implementation can be adopted: for each inference service, the target indicator metadata is selected from the indicator metadata corresponding to the inference service according to the preset data type; and the selected target indicator metadata is normalized, and the load indicator of the inference service is obtained according to the processing result.
[0042] In the normalization of the target indicator metadata in S102, the target indicator metadata can be processed into a data between 0 and 1, so that different types of target indicator metadata correspond to the same value range, thereby improving the convenience and accuracy of subsequent calculation.
[0043] If the preset data type includes only one type, the processing result after the normalization of the target indicator metadata in S102 can be used as the load indicator of the inference service; if the preset data type includes multiple types, the different processing results obtained after the normalization of different target indicator metadata in S102 are added, and the addition result is used as the load indicator of the inference service.
[0044] For example, if the preset data type is "first time-consuming data", the first time-consuming data corresponding to each inference service is normalized, and the processing result is used as the load indicator of each inference service; if the preset data type is "first time-consuming data" and "second time-consuming data", the first time-consuming data and the second time-consuming data corresponding to each inference service are normalized respectively, and the addition result between the two processing results is used as the load indicator of each inference service.
[0045] It can be understood that, in the addition of different normalized processing results in S102, the weight values corresponding to different data types can also be obtained, and then the load indicator of the inference service is obtained according to the different normalized processing results and the weight values corresponding thereto.
[0046] After obtaining the load level of each inference service in S102, the scheduling strategy corresponding to the load level of the target inference service is adopted to schedule the inference request to the target inference service in response to receiving the processing request sent by the target inference service in S103.
[0047] In the embodiment, the target inference service is any one of the plurality of inference services, and the processing request sent by the target inference service is used to pull the inference request from the queue server and process it; the inference service in the embodiment can send the processing request to the queue server when its processing capacity does not reach the upper limit.
[0048] Specifically, when executing S103, this embodiment adopts a scheduling strategy corresponding to the load level of the target reasoning service. When scheduling an inference request to the target reasoning service, the implementation method that can be adopted is: in response to determining that the load level of the target reasoning service is the first load level, without waiting, the inference request is pulled from the queue and returned to the target reasoning service.
[0049] That is to say, when this embodiment determines that the load level of the target reasoning service is the first load level, it means that the current load of the target reasoning service is low, and it will immediately pull the inference request from the queue and return it to the target reasoning service, so that the target reasoning service with low load can quickly process the inference request.
[0050] In this embodiment, when executing S103, a scheduling strategy corresponding to the load level of the target reasoning service is adopted. When scheduling an inference request to the target reasoning service, the implementation method that can be adopted is: in response to determining that the load level of the target reasoning service is the second load level, after waiting for the first period of time, the inference request is pulled from the queue and returned to the target reasoning service.
[0051] The first duration in this embodiment can be calculated using the following formula:
[0052] t1=(((x-μ) / σ)*2N)
[0053] In the above formula: t1 is the first duration; x is the load index of the target inference service; μ is the average value of the load index; σ is the variance of the load index; N is the preset time base parameter obtained through testing, and N is in milliseconds.
[0054] That is to say, when this embodiment determines that the load level of the target inference service is the second load level, it will not pull the inference request from the queue immediately, but will pull the inference request from the queue and return it to the target inference service after waiting for the first period of time, so that the target inference service under high load can process the inference request after a relatively long period of time.
[0055] In this embodiment, when executing S103, a scheduling strategy corresponding to the load level of the target reasoning service is adopted. When scheduling an inference request to the target reasoning service, the implementation method that can be adopted is: in response to determining that the load level of the target reasoning service is the third load level, after waiting for the second period of time, the inference request is pulled from the queue and returned to the target reasoning service.
[0056] The second duration in this embodiment can be calculated using the following formula:
[0057] t2=(((x-μ) / σ)*N)
[0058] In the above formula: t2 is the second duration; x is the load index of the target inference service; μ is the average value of the load index; σ is the variance of the load index; N is the preset time base parameter obtained through testing, and N is in milliseconds.
[0059] That is to say, when this embodiment determines that the load level of the target inference service is the third load level, it will not pull the inference request from the queue immediately, but will pull the inference request from the queue and return it to the target inference service after waiting for the second time period, so that the target inference service with balanced load can process the inference request after a relatively short period of time; the first time period of this embodiment is greater than the second time period.
[0060] It can be understood that when executing S103, this embodiment may also include the following contents: in response to determining that the load level of the target reasoning service is the third load level and the load index of the target reasoning service is less than the load index average, there is no need to wait, and the inference request is pulled from the queue and returned to the target reasoning service.
[0061] That is to say, when this embodiment determines that the load level of the target reasoning service is the third load level, it can further determine whether the load index of the target reasoning service is less than the load index average value. If so, it indicates that the load condition of the target reasoning service is relatively good, and there is no need to wait for the second period of time. The inference request is immediately pulled from the queue and returned to the target reasoning service for quick processing.
[0062] That is to say, this embodiment pre-sets scheduling strategies corresponding to different load levels, so that when scheduling inference requests to the target inference service, different scheduling strategies can be adopted according to different load levels, avoiding the problems in the prior art where different inference services at different load levels directly pull queue requests from the queue, resulting in load imbalance between different inference services and long average delay of inference results.
[0063] Figure 2 Schematic diagram of the second embodiment of the present disclosure. Figure 2 As shown in , this embodiment shows that when executing S102 "obtaining the load level of each reasoning service according to the load indicators of all reasoning services", the implementation methods that can be adopted are:
[0064] S201. Obtain a load index average value and a load index variance based on the load indexes of all the inference services.
[0065] S202: Obtain a load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service.
[0066] That is, the embodiment can achieve more fine-grained division of the load level and improve the accuracy of the obtained load level by determining the load level of the inference service according to the two pieces of information of the calculated load indicator average value and load indicator variance.
[0067] Specifically, when performing S202 to obtain the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service, the embodiment can be implemented in the following manner: obtaining a subtraction result between the load indicator average value and the load indicator variance; for each inference service, in response to determining that the load indicator of the inference service is less than the obtained subtraction result, taking the first load level as the load level of the inference service.
[0068] When performing S202 to obtain the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service, the embodiment can be implemented in the following manner: obtaining an addition result between the load indicator average value and the load indicator variance; for each inference service, in response to determining that the load indicator of the inference service is greater than the obtained addition result, taking the second load level as the load level of the inference service.
[0069] When performing S202 to obtain the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service, the embodiment can be implemented in the following manner: obtaining a subtraction result and an addition result between the load indicator average value and the load indicator variance; for each inference service, in response to determining that the load indicator of the inference service is less than or equal to the obtained addition result and greater than or equal to the obtained subtraction result, taking the third load level as the load level of the inference service.
[0070] That is, the embodiment can determine whether the load level of the inference service is the first load level (i.e., the low load level) according to the load indicator of the inference service and the subtraction result obtained according to the load indicator average value and the load indicator variance, can determine whether the load level of the inference service is the second load level (i.e., the high load level) according to the load indicator of the inference service and the addition result obtained according to the load indicator average value and the load indicator variance, and can determine whether the load level of the inference service is the third load level (i.e., the balanced load level) according to the load indicator of the inference service and the addition result and the subtraction result obtained according to the load indicator average value and the load indicator variance, thereby improving the accuracy of the obtained load level.
[0071] For example, when executing S202, this embodiment can first determine whether the load level of the reasoning service is the first load level based on the load indicator of the reasoning service. If not, then determine whether the load level of the reasoning service is the second load level based on the load indicator of the reasoning service. If not, determine that the load level of the reasoning service is the third load level.
[0072] For another example, when executing S202, this embodiment can also first determine whether the load level of the reasoning service is the second load level based on the load indicator of the reasoning service. If not, then determine whether the load level of the reasoning service is the third load level based on the load indicator of the reasoning service. If not, determine that the load level of the reasoning service is the first load level.
[0073] Therefore, this embodiment can adopt any determination order to obtain the load level of the inference service, and this embodiment does not limit the determination order.
[0074] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure. Figure 3 FIG shows a load view generated by this embodiment based on the load index average value, load index variance and load index of each inference service; FIG. Figure 3 As shown in the figure, the load view includes three areas: low load area, balanced load area and high load area. μ is the average value of the load indicator; σ is the variance of the load indicator. A circle in the figure represents an inference service.
[0075] Among them, the low load area includes inference services whose load indicators are less than the subtraction result between the load indicator average and the load indicator variance, and the load level of the inference services located in the low load area is the first load level; the balanced load area includes inference services whose load indicators are greater than or equal to the subtraction result between the load indicator average and the load indicator variance, and less than or equal to the addition result between the load indicator average and the load indicator variance, and the load level of the inference services located in the balanced load area is the third load level; the high load area includes inference services whose load indicators are greater than the addition result between the load indicator average and the load indicator variance, and the load level of the inference services located in the high load area is the second load level.
[0076] That is to say, this embodiment can divide different load zones based on the load indicator average value and load indicator variance obtained from the load indicators of multiple inference services, and then determine the load zone to which each inference service belongs based on the load indicator of the inference service, and the load level of each inference service can be determined based on the load zone to which the inference service belongs.
[0077] Figure 4 Schematic diagram of the fourth embodiment of the present disclosure. Figure 4As shown, the inference request scheduling device 400 of this embodiment includes:
[0078] The receiving unit 401 is used to receive indicator metadata sent by multiple inference services respectively;
[0079] The processing unit 402 is configured to obtain a load level of each inference service according to the indicator metadata;
[0080] The scheduling unit 403 is configured to, in response to receiving a processing request sent by a target reasoning service, schedule the reasoning request to the target reasoning service using a scheduling strategy corresponding to a load level of the target reasoning service.
[0081] The scheduling device for inference requests in this embodiment is located in a queue server, and the queue server is used to schedule inference requests in a queue to different inference services.
[0082] The indicator metadata received by the receiving unit 401 and sent by different inference services include three types: first time consumption data, second time consumption data, and video memory utilization data. Each type of indicator metadata is obtained in real time by a different inference service.
[0083] In this embodiment, the first time consumption data is the average time consumption data of the first token, the token is the semantic unit of the reasoning result, and the first token is the first semantic unit of the reasoning result; specifically, the reasoning service in this embodiment calculates the average value based on the first token time consumption of the reasoning results of all inference requests within the preset time window (the first token time consumption is the time required for the reasoning service to obtain the first token), and then uses the calculation result as the first time consumption data of the reasoning service.
[0084] In this embodiment, the second time consumption data is the average time consumption of the inter-package token, and the inter-package token refers to other tokens in the inference result except the first token, that is, other semantic units in the inference result except the first semantic unit; specifically, the inference service in this embodiment calculates the average value based on the inter-package token time consumption of the inference results of all inference requests within the preset time window (the inter-package token time consumption is the time required for the inference service to obtain other tokens), and then uses the calculation result as the second time consumption data of the inference service.
[0085] In this embodiment, the video memory utilization data is the video memory utilization acquired in real time by the inference service.
[0086] The preset time window in this embodiment can be determined based on the current time and the preset duration. For example, the preset duration before the current time is defined as the preset time window, and the preset duration can be 2s, 3s, 5s, etc.
[0087] That is to say, different inference services in this embodiment will obtain different types of indicator metadata through real-time calculation or real-time acquisition, and then report the obtained indicator metadata to the queue server, so that the queue server can determine the current load level of each inference service based on the received indicator metadata corresponding to different inference services.
[0088] In addition, this embodiment selects at least one of the first time consumption data, the second time consumption data and the video memory utilization data as the indicator metadata, so that the selected indicator metadata can more truly and accurately reflect the load situation of the inference service, thereby improving the accuracy of the obtained load level.
[0089] In this embodiment, the queue server may update the indicator metadata reported by different reasoning services according to the unique identification information of the reasoning services.
[0090] In this embodiment, after the receiving unit 401 receives the indicator metadata respectively sent by multiple inference services, the processing unit 402 obtains the load level of each inference service according to the received indicator metadata.
[0091] The load level obtained by the processing unit 402 is one of the first load level, the second load level, and the third load level.
[0092] In this embodiment, the first load level indicates that the load of the inference service is relatively low during the current operation, and the first load level is the low load level; the second load level indicates that the load of the inference service is relatively high during the current operation, and the second load level is the high load level; the third load level indicates that the load of the inference service is relatively balanced during the current operation, and the third load level is the balanced load level.
[0093] Specifically, when the processing unit 402 obtains the load level of each inference service based on the received indicator metadata, the implementation method that can be adopted is: obtaining the load indicator of each inference service based on the preset data type and indicator metadata; obtaining the load level of each inference service based on the load indicators of all inference services.
[0094] That is to say, on the one hand, the processing unit 402 obtains the load indicator of each inference service based on the preset data type and indicator metadata, so that the obtained load indicator corresponds to the preset data type, thereby improving the accuracy of the load indicator. On the other hand, it combines the load indicators of all inference services to obtain the load level of each inference service, thereby improving the accuracy of the obtained load level.
[0095] The processing unit 402 can compare the load indicator of each inference service with the load indicator statistical results based on the load indicator statistical results of all inference services (such as the load indicator average or the load indicator variance). If the load indicator is greater than or equal to the load indicator statistical results, the load level of the inference service is determined to be the second load level, otherwise it is the first load level.
[0096] Among them, when the processing unit 402 obtains the load indicator of each inference service based on the preset data type and indicator metadata, the implementation method that can be adopted is: for each inference service, according to the preset data type, select the target indicator metadata from the indicator metadata corresponding to the inference service; normalize the selected target indicator metadata, and obtain the load indicator of the inference service based on the processing result.
[0097] When normalizing the target indicator metadata, the processing unit 402 may process the target indicator metadata into data between [0, 1], so that different types of target indicator metadata correspond to the same numerical range, thereby improving the convenience and accuracy of subsequent calculations.
[0098] If the preset data type includes only one, the processing unit 402 can use the processing result as the load indicator of the reasoning service after normalizing the target indicator metadata; if the preset data type includes multiple, the processing unit 402 can add up the different processing results obtained after normalizing the different target indicator metadata, and use the added result as the load indicator of the reasoning service.
[0099] It is understandable that when the processing unit 402 adds different normalization processing results, it can also obtain weight values corresponding to different data types, and then obtain the load index of the inference service based on the different normalization processing results and their corresponding weight values.
[0100] When the processing unit 402 obtains the load level of each inference service based on the load indicators of all inference services, the implementation method that can be adopted is: obtaining the load indicator average value and the load indicator variance based on the load indicators of all inference services; obtaining the load level of each inference service based on the load indicator average value, the load indicator variance and the load indicator of each inference service.
[0101] That is, the processing unit 402 determines the load level of the inference service by calculating the load index average value and the load index variance, which can achieve a more refined division of the load level and improve the accuracy of the obtained load level.
[0102] Specifically, when obtaining the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service, the processing unit 402 can adopt an implementation manner that: obtaining a subtraction result between the load indicator average value and the load indicator variance; for each inference service, in response to determining that the load indicator of the inference service is less than the obtained subtraction result, taking the first load level as the load level of the inference service.
[0103] When obtaining the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service, the processing unit 402 can adopt an implementation manner that: obtaining an addition result between the load indicator average value and the load indicator variance; for each inference service, in response to determining that the load indicator of the inference service is greater than the obtained addition result, taking the second load level as the load level of the inference service.
[0104] When obtaining the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service, the processing unit 402 can adopt an implementation manner that: obtaining a subtraction result and an addition result between the load indicator average value and the load indicator variance; for each inference service, in response to determining that the load indicator of the inference service is less than or equal to the obtained addition result and greater than or equal to the obtained subtraction result, taking the third load level as the load level of the inference service.
[0105] That is, the processing unit 402 obtains the load level of each inference service according to the load indicator of each inference service, and according to the addition result and / or the subtraction result obtained from the load indicator average value and the load indicator variance, which can improve the accuracy of the obtained load level.
[0106] After the processing unit 402 obtains the load level of each inference service, the scheduling unit 403 responds to receiving a processing request sent by a target inference service, adopts a scheduling strategy corresponding to the load level of the target inference service, and schedules an inference request to the target inference service.
[0107] In the embodiment, the target inference service is any one of the plurality of inference services, and the processing request sent by the target inference service is used to pull an inference request from the queue server and process it; the inference service in the embodiment can send a processing request to the queue server when its processing capacity does not reach the upper limit.
[0108] Specifically, when the scheduling unit 403 adopts a scheduling strategy corresponding to the load level of the target reasoning service to schedule an inference request to the target reasoning service, the implementation method that can be adopted is: in response to determining that the load level of the target reasoning service is the first load level, without waiting, the inference request is pulled from the queue and returned to the target reasoning service.
[0109] That is, when the scheduling unit 403 determines that the load level of the target inference service is the first load level, it will immediately pull the inference request from the queue and return it to the target inference service, so that the target inference service with low load can quickly process the inference request.
[0110] When the scheduling unit 403 adopts a scheduling strategy corresponding to the load level of the target reasoning service to schedule an inference request to the target reasoning service, the implementation method that can be adopted is: in response to determining that the load level of the target reasoning service is the second load level, after waiting for the first period of time, the inference request is pulled from the queue and returned to the target reasoning service.
[0111] The first duration in this embodiment can be calculated using the following formula:
[0112] t1=(((x-μ) / σ)*2N)
[0113] In the above formula: t1 is the first duration; x is the load index of the target inference service; μ is the average value of the load index; σ is the variance of the load index; N is the preset time base parameter obtained through testing, and N is in milliseconds.
[0114] That is to say, when the scheduling unit 403 determines that the load level of the target inference service is the second load level, it will not pull the inference request from the queue immediately, but will pull the inference request from the queue and return it to the target inference service after waiting for the first period of time, so that the target inference service under high load can process the inference request after a relatively long period of time.
[0115] When the scheduling unit 403 adopts a scheduling strategy corresponding to the load level of the target reasoning service to schedule an inference request to the target reasoning service, the implementation method that can be adopted is: in response to determining that the load level of the target reasoning service is the third load level, after waiting for the second period of time, pull the inference request from the queue and return it to the target reasoning service.
[0116] The second duration in this embodiment can be calculated using the following formula:
[0117] t2=(((x-μ) / σ)*N)
[0118] In the above formula: t2 is the second time length; x is the load index of the target inference service; μ is the average value of the load index; σ is the variance of the load index; N is a preset time reference parameter obtained through testing, and N is in the order of milliseconds.
[0119] That is, the scheduling unit 403 will not immediately pull the inference request from the queue when determining that the load level of the target inference service is the balanced load level, but will wait for a second time length, and then pull the inference request from the queue and return it to the target inference service, so that the target inference service in the balanced load can process the inference request after a relatively short period of time.
[0120] It can be understood that the scheduling unit 403 can further include the following content: in response to determining that the load level of the target inference service is the third load level and the load index of the target inference service is less than the average value of the load index, the inference request is pulled from the queue and returned to the target inference service without waiting.
[0121] That is, the scheduling unit 403 can further determine whether the load index of the target inference service is less than the average value of the load index when determining that the load level of the target inference service is the third load level, and if so, it indicates that the load of the target inference service is relatively good, and the inference request is pulled from the queue and returned to the target inference service without waiting for a second time length, for fast processing.
[0122] In the technical solution of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0123] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0124] As Figure 5 shown, it is a block diagram of an electronic device for scheduling an inference request according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0125] As Figure 5As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0126] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0127] The computing unit 501 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the scheduling method for reasoning requests. For example, in some embodiments, the scheduling method for reasoning requests can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 508.
[0128] In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for scheduling inference requests described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the method for scheduling inference requests in any other appropriate manner (e.g., by means of firmware).
[0129] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable vehicle positioning or positioning model training device, such that when executed by the processor or controller, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can include or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0133] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0134] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service system that addresses the management difficulties and poor business scalability of traditional physical hosts and VPS services ("Virtual Private Servers," or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0135] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0136] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for scheduling an inference request, comprising: Receive indicator metadata sent by multiple inference services; Obtaining a load level of each inference service based on the indicator metadata; In response to receiving a processing request sent by a target reasoning service, the reasoning request is scheduled to the target reasoning service using a scheduling policy corresponding to a load level of the target reasoning service.
2. The method according to claim 1, wherein Obtaining the load level of each inference service according to the indicator metadata includes: Obtaining a load indicator for each inference service based on a preset data type and the indicator metadata; The load level of each reasoning service is obtained according to the load indicators of all reasoning services.
3. The method according to claim 2, wherein: Obtaining the load level of each reasoning service according to the load indicators of all reasoning services includes: Obtaining a load index average and a load index variance based on the load indexes of all the inference services; The load level of each reasoning service is obtained according to the load indicator average value, the load indicator variance, and the load indicator of each reasoning service.
4. The method according to claim 2, wherein: Obtaining the load indicator of each inference service according to the preset data type and the indicator metadata includes: For each reasoning service, selecting target indicator metadata from the indicator metadata corresponding to the reasoning service according to the preset data type; The target indicator metadata is normalized, and the load indicator of the inference service is obtained according to the processing result.
5. The method according to claim 3, wherein Obtaining the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service includes: Obtaining a subtraction result between the load indicator average value and the load indicator variance; For each reasoning service, in response to determining that the load indicator of the reasoning service is less than the subtraction result, the first load level is used as the load level of the reasoning service.
6. The method according to claim 3, wherein: Obtaining the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service includes: Obtaining a sum of the load indicator average value and the load indicator variance; For each inference service, in response to determining that the load index of the inference service is greater than the addition result, the second load level is used as the load level of the inference service.
7. The method according to claim 3, wherein: Obtaining the load level of each inference service according to the load indicator average value, the load indicator variance, and the load indicator of each inference service includes: Obtaining a subtraction result and an addition result between the load indicator average value and the load indicator variance; For each reasoning service, in response to determining that the load index of the reasoning service is less than or equal to the addition result and greater than or equal to the subtraction result, the third load level is used as the load level of the reasoning service.
8. The method according to claim 1, wherein The adopting a scheduling strategy corresponding to the load level of the target reasoning service to schedule the inference request to the target reasoning service includes: In response to determining that the load level of the target inference service is the first load level, the inference request is pulled from the queue and returned to the target inference service without waiting.
9. The method according to claim 1, wherein The adopting a scheduling strategy corresponding to the load level of the target reasoning service to schedule the inference request to the target reasoning service includes: In response to determining that the load level of the target inference service is a second load level, after waiting for a first period of time, an inference request is pulled from the queue and returned to the target inference service.
10. The method according to claim 1, wherein The adopting a scheduling strategy corresponding to the load level of the target reasoning service to schedule the inference request to the target reasoning service includes: In response to determining that the load level of the target inference service is a third load level, after waiting for a second period of time, an inference request is pulled from the queue and returned to the target inference service.
11. A scheduling device for an inference request, comprising: A receiving unit, configured to receive indicator metadata sent by multiple inference services; a processing unit, configured to obtain a load level of each inference service based on the indicator metadata; The scheduling unit is configured to, in response to receiving a processing request sent by a target reasoning service, schedule the reasoning request to the target reasoning service by adopting a scheduling strategy corresponding to a load level of the target reasoning service.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 10.