Control method for scheduling reasoning service capability and electronic device

By obtaining and dividing the quota for inference service capabilities, the message scheduling gateway/load balancing device is used to dynamically control the inference service call of AI applications, which solves the problem of resource imbalance in high concurrency of AI applications and improves system stability and user experience.

CN120455526APending Publication Date: 2025-08-08BEIJING ZTE DIGITAL NEBULA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510803014.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art cannot evenly schedule inference services when there are too many inference requests for a certain or some AI applications, resulting in uneven resource allocation and affecting user experience and business efficiency.

Method used

By obtaining the quota for inference service capabilities, dividing it into specific quotas and shared quotas, the target AI application calls inference service capabilities, and dynamically control it using the message scheduling gateway/load balancing device to ensure balanced allocation of resources.

Benefits of technology

It realizes resource balanced scheduling during high concurrency periods of AI applications, improves system stability and user experience, and improves the business efficiency of AI applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455526A_ABST
    Figure CN120455526A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a control method for scheduling reasoning service capability and an electronic device. The control method comprises the following steps: acquiring a quota of the reasoning service capability; and controlling a target artificial intelligence AI application to call the reasoning service capability according to the quota. According to the method and the device, the problem that the reasoning service cannot be scheduled in a balanced manner when the reasoning requests of a certain or some AI applications are excessive in the prior art is solved, so that the effects of improving the stability and the efficiency of the whole system and improving the service experience of a user using the AI applications are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of communications, and in particular, to a method and electronic device for controlling scheduling reasoning service capabilities. Background Art

[0002] With the increasing number of AI applications and the continuous expansion of application scenarios, the demand for inference service calls has become more diverse and highly concurrent. However, traditional inference service scheduling mechanisms often focus on load balancing between backend service instances, overlooking the issue of call balancing between front-end AI applications. In production environments, especially when faced with complex and dynamically changing business scenarios, when inference requests from some AI applications surge, existing mechanisms struggle to effectively control and allocate limited inference service capacity, resulting in uneven resource allocation. Some AI applications may face service request delays or rejections, severely impacting user experience and business efficiency. Summary of the Invention

[0003] An embodiment of the present invention provides a control method and electronic device for scheduling inference service capabilities, so as to at least solve the problem in the related art that inference services cannot be evenly scheduled when there are too many inference requests from one or some AI applications.

[0004] According to one embodiment of the present invention, a control method for scheduling inference service capabilities is provided, including: obtaining a quota of inference service capabilities; and controlling a target artificial intelligence (AI) application to call the inference service capabilities according to the quota.

[0005] According to yet another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when run.

[0006] According to another embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0007] According to yet another embodiment of the present invention, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.

[0008] The above-described embodiment of the present invention controls the target AI application's inference service capability invocation based on the obtained inference service capability quota. This resolves the problem in related technologies of being unable to evenly dispatch inference services when one or more AI applications have too many inference requests, thereby improving the stability and efficiency of the entire system and enhancing the user experience of using AI applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 This is a schematic diagram of AI application reasoning service calls in related technologies;

[0010] Figure 2 is a flow chart (1) of a method for controlling scheduling reasoning service capabilities according to an embodiment of the present invention;

[0011] Figure 3 This is a structural block diagram of a message scheduling gateway / load balancing device according to an embodiment of the present invention;

[0012] Figure 4 Flowchart (II) of the method for controlling scheduling reasoning service capabilities according to an embodiment of the present invention. DETAILED DESCRIPTION

[0013] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments.

[0014] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0015] During AI application development, inference services based on large and small models are deployed on a single GPU server or GPU server cluster. Inference service interfaces are published to upper-layer AI applications, allowing them to invoke the inference capabilities of large and small models. Whether deployed in an all-in-one appliance or in a resource pool, the inference service capabilities provided are limited based on specific application traffic models, hardware resources, and the operating environment of large and small model instances. For example:

[0016] Traffic model: Single session context 4KB, first word delay no more than 2 seconds, and output delay of all words except the first word no more than 30 milliseconds.

[0017] Hardware resources, for a single-machine inference system, are configured as follows: CPU: 2 of a certain model (2.7GHz 64C); Memory: 24GB DDR5 96GB; Hard Drive: 2 x 960GB SATA SSDs + 6 x 1.92TB NVME SSDs; GPU Card: 8 x Lxx GPU cards (48GB video memory).

[0018] Model: DeepSeek-R1-Disti l l-Qwen-32B. Deployment method: Deploy one large model instance for every two GPU cards, for a total of four instances deployed on a single machine.

[0019] Based on the above constraints, the maximum concurrent request of the DeepSeek-R1-Disti l l-Qwen-32B inference service released by this GPU server is 20QPS (Query Percent Second).

[0020] In order to ensure that the processing capabilities of multiple inference service instances on the backend are evenly called by AI applications, a message gateway or load balancing device is generally set up on the GPU server or cluster to distribute the inference call messages received from AI applications to the backend inference service instances according to a certain distribution strategy, ensuring that the four inference service instances on the backend are called evenly. Figure 1 This is a schematic diagram of AI application reasoning service calls in related technologies, such as Figure 1 shown.

[0021] However, this distribution mechanism in related technologies can only solve the problem of balanced message processing between backend inference service instances, and cannot solve the problem of balanced inference service calls by AI applications. Assume that there are four AI applications A1, A2, A3, and A4 in the application layer. Under normal circumstances, each AI application will call the inference service for inference. However, in actual engineering applications, there is an imbalance in the inference service calls of AI applications, such as Figure 1 The A4 application is a high-concurrency business during a certain period of time. One application accounts for all 20 inference service requests, resulting in no inference services available for the A1, A2, and A3 applications.

[0022] Therefore, there is an urgent need for a method to solve the problem of balanced use of inference services when there are too many inference requests for one or more AI applications.

[0023] To solve the above technical problems, this embodiment provides a control method for scheduling reasoning service capabilities. Figure 2 Flowchart (1) of the method for controlling the scheduling reasoning service capability according to an embodiment of the present invention, Figure 2 As shown, the process includes the following steps:

[0024] Step S202: Obtain the quota of the inference service capability.

[0025] For example, the same inference service can be configured with different inference service capabilities (QPS) and quotas (QPS) according to different context lengths (K) and AI application requirements.

[0026] In an exemplary embodiment, a quota includes one or more first sub-quotas and a second sub-quota.

[0027] For example, the first sub-quota can be a specific quota, and the second sub-quota can be a shared quota. A specific quota is directly authorized to a specific AI application and is used by the designated AI application. A shared quota is authorized for shared use by multiple AI applications. After the specific quota is exhausted, it is used by shared AI applications through a competitive mechanism.

[0028] In an exemplary embodiment, step S202 includes:

[0029] Obtain the user-configured reasoning service capability and / or reasoning service capability quota; or

[0030] Obtain parameters of the inference service system and configure inference service capabilities and / or inference service capability quotas based on the parameters.

[0031] Illustratively, the reasoning service capability and the quota of the reasoning service capability may be configured by a user, or may be automatically configured by obtaining parameters from the reasoning service system through an interface with a backend reasoning service system.

[0032] Exemplarily, the parameter can be historical data used by the existing AI application reasoning service capabilities in the reasoning service system, or it can be an empirical value summarized based on historical data (such as QPS ratio, specific QPS value, etc.), or it can be the business busyness of the AI application in the actual production environment, the user's business characteristics and the reasoning capabilities of the reasoning service system, etc.

[0033] In an exemplary embodiment, configuring a quota of the reasoning service capability according to parameters includes:

[0034] The reasoning service capacity is divided into one or more first sub-quotas and one second sub-quota according to the parameter, wherein the sum of the one or more first sub-quotas and the second sub-quota is less than or equal to the reasoning service capacity.

[0035] For example, if the inference service capacity of inference service A is 100QPS, the first sub-quota and the second sub-quota can be divided in a ratio of 4:1, so the first sub-quota is 80QPS and the second sub-quota is 20QPS. If there are multiple first sub-quotas (for example, 4), the first sub-quotas can be evenly distributed as 20QPS, 20QPS, 20QPS, and 20QPS, or unevenly distributed as 25QPS, 10QPS, 40QPS, and 5QPS. However, no matter how the distribution is made, the sum of one or more first sub-quotas and one second sub-quota must be less than or equal to the inference service capacity.

[0036] Step S204: Control the target artificial intelligence (AI) application's inference service capabilities based on the quota.

[0037] For example, the target AI application can be controlled to call the reasoning service capability according to the obtained quota to achieve balanced scheduling of the reasoning service capability.

[0038] In an exemplary embodiment, before step S204, the following steps are included:

[0039] receiving an inference service scheduling request of a target AI application, wherein the inference service scheduling request includes target inference service capabilities;

[0040] In response to the quota including a first sub-quota and a second sub-quota, determining whether the target reasoning service capability is greater than the first sub-quota, or in response to the quota including multiple first sub-quotas, determining whether the target reasoning service capability is greater than the first sub-quota corresponding to the target AI application.

[0041] For example, taking the quota as including a first sub-quota and a second sub-quota, the target AI application is application A, the target inference service capability of application A is 50QPS, the inference service capability is 100QPS, the first sub-quota is 80QPS, and the second sub-quota is 20QPS, then it is necessary to determine whether the target inference service capability of 50QPS is greater than the first sub-quota of 80QPS.

[0042] For example, a quota includes multiple first sub-quotas and one second sub-quota. The target AI application is application A, and its target inference service capability is 70 QPS and 100 QPS. The first sub-quota for application A is 20 QPS, while the first sub-quota for application B is 50 QPS and its second sub-quota is 30 QPS. To determine whether the target inference service capability of 50 QPS is greater than the first sub-quota of 20 QPS for application A, consider the following:

[0043] In an exemplary embodiment, step S204 includes:

[0044] In response to the target reasoning service capability being less than or equal to the first sub-quota or the first sub-quota corresponding to the target AI application, scheduling the first sub-quota or the first sub-quota corresponding to the target AI application for the target AI application according to the target reasoning service capability;

[0045] Update the current first sub-quota to the target inference service capacity;

[0046] Forwards the inference service dispatch request to the inference service instance.

[0047] For example, let's assume that the quota includes a first sub-quota and a second sub-quota, the target AI application is application A, the target inference service capability of application A is 50QPS, the inference service capability is 100QPS, the first sub-quota is 80QPS, and the second sub-quota is 20QPS. When it is determined that the target inference service capability of 50QPS is less than the first sub-quota of 80QPS, a portion of the first sub-quota, i.e., 50QPS, can be scheduled for the target AI application, and the current first sub-quota can be updated to 50QPS. This means that we can know how much of the first sub-quota has been used, and thus how many QPS of the first sub-quota are left to be used.

[0048] For example, take a quota that includes multiple first sub-quotas and one second sub-quota, the target AI application is application A, the target inference service capability of application A is 10QPS, the inference service capability is 100QPS, the first sub-quota of application A is 20QPS, the first sub-quota of application B is 50QPS, and the second sub-quota is 30QPS. When it is determined that the target inference service capability of 10QPS is less than the first sub-quota of 20QPS corresponding to the target AI application, the portion of the first sub-quota corresponding to the target AI application, i.e., 10QPS, can be scheduled for the target AI application, and the current first sub-quota can be updated to 10QPS. This is equivalent to knowing how much of the first sub-quota corresponding to the target AI application has been used, and thus knowing how many QPS of the first sub-quota corresponding to the target AI application are left to use.

[0049] In an exemplary embodiment, step S204 includes:

[0050] In response to the target reasoning service capability being greater than the first sub-quota or the first sub-quota corresponding to the target AI application, updating the current first sub-quota to the target reasoning service capability, and recording the difference between the first sub-quota or the first sub-quota corresponding to the target AI application and the current first sub-quota;

[0051] Determine whether the difference is greater than the second sub-quota;

[0052] In response to the difference being less than or equal to the second sub-quota, scheduling the second sub-quota and the first sub-quota or the first sub-quota corresponding to the target AI application for the target AI application according to the target inference service capability;

[0053] Update the current second sub-quota to the difference;

[0054] Forwards the inference service dispatch request to the inference service instance.

[0055] For example, take the quota as including a first sub-quota and a second sub-quota, the target AI application is application A, the target inference service capability of application A is 70QPS, the inference service capability is 100QPS, the first sub-quota is 60QPS, and the second sub-quota is 40QPS. When it is determined that the target inference service capacity of 70QPS is greater than the first sub-quota of 60QPS, the current first sub-quota is updated to 70QPS, so that the difference between the first sub-quota and the current first sub-quota can be recorded as 10QPS. It is further determined whether the difference is greater than the second sub-quota. When it is determined that the difference of 10QPS is less than the second sub-quota of 60QPS, the first sub-quota and the second sub-quota can be scheduled for the target AI application, which is equivalent to scheduling all of the first sub-quota of 60QPS to the target AI application, and also scheduling part of the second sub-quota, that is, 10QPS, to the target AI application. The current second sub-quota is updated to the difference of 10QPS, which is equivalent to knowing how much of the second sub-quota has been used, and thus knowing how many QPS of the second sub-quota are left to be used.

[0056] For example, a quota includes multiple first sub-quotas and one second sub-quota. The target AI application is application A. The target inference service capability of application A is 30QPS, the inference service capability is 100QPS, the first sub-quota of application A is 20QPS, and the first sub-quota of application B is 50QPS and the second sub-quota is 30QPS. When it is determined that the target inference service capability of 30QPS is greater than the first sub-quota of 20QPS corresponding to the target AI application, the current first sub-quota is updated to 30QPS, so that the difference between the first sub-quota corresponding to the target AI application and the current first sub-quota is recorded as 10QPS, and it is further determined whether the difference is greater than the second sub-quota. When it is determined that the difference of 10QPS is less than the second sub-quota of 30QPS, the first sub-quota and the second sub-quota corresponding to the target AI application can be scheduled for the target AI application, which is equivalent to scheduling all of the first sub-quota of 20QPS corresponding to the target AI application to the target AI application, and also scheduling part of the second sub-quota, that is, 10QPS, to the target AI application. The current second sub-quota is updated to the difference of 10QPS, which is equivalent to knowing how much of the second sub-quota has been used, and thus knowing how many QPS of the second sub-quota are left to be used.

[0057] In one exemplary embodiment, the method further comprises:

[0058] In response to the difference being greater than the second sub-quota, a scheduling failure response message is returned.

[0059] For example, take a quota that includes multiple first sub-quotas and one second sub-quota, the target AI application is application A, the target inference service capability of application A is 70QPS, the inference service capability is 100QPS, the first sub-quota of application A is 20QPS, the first sub-quota of application B is 50QPS, and the second sub-quota is 30QPS. First, it is determined that the target inference service capability of 70QPS is greater than the first sub-quota of 20QPS of application A, then the current first sub-quota is updated to 70QPS, and the difference between the first sub-quota and the current first sub-quota is recorded as 50QPS. It is also determined that the difference of 50QPS is greater than the second sub-quota of 30QPS. At this time, the inference service capability cannot be scheduled for application A, and a response message indicating scheduling failure is returned.

[0060] For example, the response information of scheduling failure may carry a failure code corresponding to quota exhaustion, or other failure codes agreed upon by the target AI application and the inference service system, or specific reasons for failure, such as quota exhaustion, insufficient quota, etc.

[0061] In an exemplary embodiment, the reasoning service scheduling request includes at least one of the following:

[0062] Target AI application identification information, target inference service, and target context length, wherein the target AI application identification information includes at least one of the following: target AI application name, target AI application ID, and target AI application call token.

[0063] Through the above steps, the relevant technology can solve the problem of being unable to evenly schedule inference services when there are too many inference requests for one or some AI applications, thereby improving the stability and efficiency of the entire system and improving the user experience of using AI applications.

[0064] In this embodiment, a message scheduling gateway / load balancing device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have already been described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0065] Figure 3 is a structural block diagram of a message scheduling gateway / load balancing device according to an embodiment of the present invention. Figure 3 As shown, the device includes a reasoning capability configuration / acquisition module, a quota policy generation module, and a quota policy execution module.

[0066] These three modules are used to manage on-demand quotas for inference service capabilities, ensuring that front-end AI applications can still use inference services even when business is busy.

[0067] The reasoning capability configuration / acquisition module can obtain the user-configured reasoning service capabilities and the parameters of the reasoning service system, and configure the reasoning service capabilities based on these parameters. For example, Table 1 shows a schematic diagram of the configuration of reasoning service capabilities, which shows the reasoning service capabilities of each reasoning service under different request message context length constraints.

[0068] Table 1

[0069]

[0070] The quota policy generation module can obtain the quota of the inference service capability configured by the user, or configure the quota of the inference service capability based on the parameters of the inference service system. The quota includes one or more first sub-quotas (specific quotas) and a second sub-quota (shared quota). For example, Table 2 is a configuration diagram of the quota of the inference service capability. Based on the configuration table of the inference service capability, a quota table for the AI application calling the corresponding inference service capability can be generated. Each policy contains the AI application identifier, application call token, inference service (identifier), context length, specific quota and shared quota, as shown in Table 2.

[0071] Table 2

[0072]

[0073] The inference service capability under a specified context length. For example, the DSR1-FP8-671 B service in Table 1 provides 60QPS of inference service capability when the context length is 4KB. In Table 2, the specific quota for this scenario can be configured as 48QPS and the shared quota as 12QPS. The specific allocation strategy is as follows: APP1 has a specific quota of 30QPS, APP3 has a specific quota of 12QPS, and APP1 and APP3 have a shared quota of 12QPS. When the specific quotas of individual applications are exhausted, the shared quota can be flexibly used. The shared quota can be used in a competitive preemptive manner.

[0074] The quota policy execution module can control the target artificial intelligence (AI) application's ability to call reasoning services based on quotas. For example, when receiving a reasoning service scheduling request from a target AI application, the reasoning service scheduling request is identified, and information such as the AI application identifier, the called reasoning service, and the length of the called context are identified. Based on the quota table, the reasoning service capability during the use of the AI application business is dynamically controlled. If the configuration data of the specific quota in the quota table is not exceeded, the call is made normally, and the current first sub-quota is updated. If the configuration data of the specific quota in the quota table is exceeded, a shared quota is allocated. If the shared quota can be used, the reasoning service scheduling request is called and executed normally, and the current second sub-quota is updated. If the shared quota has been exhausted or cannot be used, a response message indicating scheduling failure is returned, carrying the failure reason as quota exhaustion.

[0075] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0076] Example 1

[0077] This embodiment takes the complete process of inference service capability configuration, AI application quota configuration, and AI application access control as an example, which can be divided into two stages.

[0078] Phase 1: Configuring inference service capabilities and quotas.

[0079] Step S1: Based on specific inference resources, such as a single GPU node or cluster, based on specific traffic models, such as inference request context length, and specific model instances, configure the inference service capability configuration table.

[0080] Step S2: Based on the inference service capability configuration table, configure a specific quota table for AI applications to perform quota control on AI applications calling specific inference service capabilities.

[0081] Step S3: The quota table is synchronized to the quota policy execution module, and the inference service capability of the AI application is called and controlled based on the quota policy.

[0082] Phase 2: The user's request is processed by the target AI application and then invokes the inference service capability. Figure 4 Flowchart (II) of the method for controlling the scheduling reasoning service capability according to an embodiment of the present invention is as follows: Figure 4 shown.

[0083] Step S400: The user initiates a request to use an AI application.

[0084] Step S401: The AI application processes the user request, analyzes and assembles the inference service scheduling request, and sends it to the message scheduling gateway / load balancing device (referred to as the message gateway for short).

[0085] Step S402: The message gateway forwards the inference service scheduling request to the quota policy execution module, which executes process P1:

[0086] Identify inference service scheduling requests, identify information such as the AI application ID, the invoked inference service, and the call context length, and then control the inference request in this scenario:

[0087] If the configuration data of the specific quota in the quota table is not exceeded, the call is performed normally, and the process goes to S403A, where the specific quota is called and the current first sub-quota is updated;

[0088] If the configuration data of the specific quota of the policy table is exceeded, a shared quota is allocated. If the shared quota can be used, the inference service scheduling request is called and executed normally, the process goes to S403A, and the current second sub-quota is updated. If the shared quota is exhausted or cannot be used, a response message of scheduling failure is returned, and the process goes to S403B because the quota is exhausted.

[0089] Step S403:

[0090] S403A: The quota allocation of the inference service scheduling request is approved, and the quota policy execution module forwards the inference service scheduling request to a specific inference service instance;

[0091] S403B: The quota allocation of the inference service scheduling request is not approved, and the quota policy execution module returns a scheduling failure response message to the message gateway. The reason is: quota exhaustion.

[0092] Step S404;

[0093] S404A: The response message of successful scheduling is returned to the message gateway;

[0094] S404B: The message gateway returns a scheduling failure response to the AI application, indicating that the reason is quota exhaustion. The AI application then enters exception handling process P2: waiting for a predetermined period of time (e.g., 1 second, during which system resources may be released) before re-initiating the inference service scheduling request, and the process returns to S401; alternatively, a direct response (quota exhaustion) is returned to the user, and the process returns to S406.

[0095] Step S405: The message gateway returns a response message indicating successful scheduling to the AI application;

[0096] Step S406: The AI application returns a response message to the user, and the process ends.

[0097] Through the above steps, it can be ensured that the target AI application can evenly schedule the inference service capabilities during busy periods, improve the stability and efficiency of the entire system, and improve the user experience of using AI applications.

[0098] Example 2

[0099] This embodiment takes the automatic configuration of the reasoning service capability and the quota of the reasoning service capability by the user configuration and quota policy generation module as an example.

[0100] Table 3 is a schematic diagram of the configuration of the reasoning service capability of this implementation, as shown in Table 3.

[0101] Table 3

[0102] Inference Service Context length (K) Reasoning service capability (QPS) QWQ-32B 4 200

[0103] Assume there are three AI applications that can call the inference service: B1, B2, and B3. B1 is a heavy inference user and, during peak hours, consumes the entire inference service capacity (200 QPS). This makes it difficult for B2 and B3 to request inference services when B1 is busy, or even for them to receive no inference services at all, resulting in a poor user experience.

[0104] To solve this problem, it is necessary to configure quotas for AI applications based on the reasoning service capabilities of the reasoning service system.

[0105] Method 1: User Configuration

[0106] Based on the level of inference service usage, users can plan their computing in advance, specifying 80% of the total inference service capacity as a specific quota and 20% as a shared quota. For example, if B1 is assigned an 80QPS quota, B2 and B3 are each assigned a 40QPS quota, and 20% of the total inference service capacity is used as a shared quota, that is, 40QPS is used as a shared quota. When the specific quota is exhausted, B1, B2, and B3 can compete for the shared 40QPS quota. The first application to obtain the quota will use it until it is exhausted.

[0107] It should be noted that the allocation of the above-mentioned specific quotas and shared quotas needs to be reasonably planned according to the specific AI application scenarios, and the above-mentioned 20% shared quota is only an example.

[0108] Table 4 is a schematic diagram of the configuration of the quota of the reasoning service capability configured by the user, as shown in Table 4.

[0109] Table 4

[0110]

[0111] Method 2: Automatic configuration

[0112] The quota policy generation module can infer the parameters of the service system and configure the quota of the inference service capability based on the parameters. For example, based on the historical data of the use of the inference service capability of existing AI applications, the proportion of each inference service called by AI applications is counted, and then dynamically set and adjusted during the operation of the system.

[0113] For B1 applications, the QWQ-32B inference service has a 4K context length, and the year-on-year occupancy rate of the inference service in each period is 50%;

[0114] For B2 applications, the QWQ-32B inference service with a 4K context length had a year-on-year occupancy rate of 25% in each period.

[0115] For B3 applications, when the QWQ-32B inference service has a context length of 4K, the year-on-year occupancy rate of the inference service in each period is 25%.

[0116] For the QWQ-32B inference service, with a 4KB context length, the inference service capacity is 200 QPS, as shown in Table 3. If 80% is set as the specific quota and 20% as the shared quota, the inference service capacity quota for each application can be calculated. Table 5 shows a schematic diagram of the automatically configured inference service capacity quota.

[0117] Table 5

[0118]

[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.

[0120] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when running.

[0121] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0122] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0123] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0124] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.

[0125] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0126] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for controlling scheduling reasoning service capabilities, characterized in that: include: Get the quota for inference service capabilities; The target artificial intelligence (AI) application calls the inference service capability according to the quota control.

2. The method according to claim 1, characterized in that The quota includes one or more first sub-quotas and one second sub-quota.

3. The method according to claim 2, characterized in that Before the quota control target artificial intelligence (AI) application calls the inference service capability, the following steps are included: receiving an inference service scheduling request for the target AI application, wherein the inference service scheduling request includes a target inference service capability; In response to the quota including a first sub-quota, determining whether the target reasoning service capability is greater than the first sub-quota, or in response to the quota including multiple first sub-quotas, determining whether the target reasoning service capability is greater than the first sub-quota corresponding to the target AI application.

4. The method according to claim 3, characterized in that The ability to call the inference service according to the quota control target artificial intelligence (AI) application includes: In response to the target reasoning service capability being less than or equal to the first sub-quota or the first sub-quota corresponding to the target AI application, scheduling the first sub-quota or the first sub-quota corresponding to the target AI application for the target AI application according to the target reasoning service capability; Updating the current first sub-quota to the target inference service capability; The reasoning service scheduling request is forwarded to the reasoning service instance.

5. The method according to claim 3, characterized in that The inference service capability of the target artificial intelligence (AI) application according to the quota control includes: In response to the target reasoning service capability being greater than the first sub-quota or the first sub-quota corresponding to the target AI application, updating the current first sub-quota to the target reasoning service capability, and recording a difference between the first sub-quota or the first sub-quota corresponding to the target AI application and the current first sub-quota; Determining whether the difference is greater than the second sub-quota; In response to the difference being less than or equal to the second sub-quota, scheduling the second sub-quota and the first sub-quota or the first sub-quota corresponding to the target AI application for the target AI application according to the target inference service capability; Updating the current second sub-quota to the difference; The reasoning service scheduling request is forwarded to the reasoning service instance.

6. The method according to claim 5, characterized in that The method further comprises: In response to the difference being greater than the second sub-quota, a scheduling failure response message is returned.

7. The method according to claim 1, characterized in that The quotas for obtaining inference service capabilities include: Obtaining the user-configured reasoning service capability and / or the quota of the reasoning service capability; or Parameters of the reasoning service system are acquired, and the reasoning service capability and / or the quota of the reasoning service capability are configured according to the parameters.

8. The method according to claim 7, characterized in that Configuring the quota of the reasoning service capability according to the parameters includes: The reasoning service capability is divided into one or more first sub-quotas and one second sub-quota according to the parameter, wherein a sum of the one or more first sub-quotas and the one second sub-quota is less than or equal to the reasoning service capability.

9. The method according to claim 3, characterized in that The reasoning service scheduling request includes at least one of the following: Target AI application identification information, target inference service, and target context length, wherein the target AI application identification information includes at least one of the following: target AI application name, target AI application ID, and target AI application call token.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 9 are implemented.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.