Reasoning method, reasoning system, device, equipment and storage medium

The scheduling service actively perceives the idle resources and load conditions of the inference service, realizes the reasonable scheduling of resources, solves the problem of load imbalance in clustered deployment, improves resource utilization and system stability, and enhances user experience.

CN119294514BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411303517.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-09-23
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

The existing clustered deployment of large-model inference services makes it difficult to perform reasonable resource scheduling based on the characteristics of inference services, resulting in load imbalance and resource waste, affecting system stability and resource utilization.

Method used

The scheduling service actively perceives the idle resources and load conditions of the inference service, matches the target inference request, realizes the reasonable scheduling of resources, avoids load imbalance, and adopts clustered deployment of scheduling services and inference services to ensure system stability and high availability.

Benefits of technology

It effectively improves resource utilization, ensures system stability and response efficiency, enhances user experience, and has fault tolerance and high availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119294514B_ABST
    Figure CN119294514B_ABST
Patent Text Reader

Abstract

The present disclosure provides an inference method, an inference system, an apparatus, a device, and a storage medium, which relate to the field of data processing, especially technical fields such as artificial intelligence, big data, and large models. A specific implementation scheme is as follows: receiving a first task request through a first scheduling service among multiple scheduling services, wherein the first task request is a request sent by the first inference service among multiple inference services, and carries at least a model identifier and the current idle inference resources of the first inference service; based on the model identifier and idle inference resources carried by the first task request, determining a target inference request that matches the first task request; and sending the target inference request to the first inference service through the first scheduling service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to technical fields such as artificial intelligence, big data, and big models. Background Art

[0002] With the rapid development of artificial intelligence (AI), large models have become a research hotspot in the field of artificial intelligence (AI). They can extract knowledge from massive amounts of data and excel in a variety of tasks. To apply large models to real-world scenarios and meet the needs of a large number of users, clustered deployment of inference services for large models is often required. However, existing clustered deployment methods struggle to optimize resource scheduling based on the inherent characteristics of inference. Summary of the Invention

[0003] The present disclosure provides an inference method, an inference system, an apparatus, a device, and a storage medium.

[0004] According to one aspect of the present disclosure, there is provided a reasoning method, the method comprising:

[0005] Receiving a first task request through a first scheduling service among the multiple scheduling services, wherein the first task request is a request sent by the first inference service among the multiple inference services and carries at least a model identifier and currently idle inference resources of the first inference service;

[0006] Determining a target inference request that matches the first task request based on the model identifier and the idle inference resources carried by the first task request;

[0007] According to another aspect of the present disclosure, there is provided an inference system, comprising:

[0008] A first scheduling service is configured to receive a first task request, wherein the first task request carries at least a model identifier and currently idle inference resources of the first inference service; based on the model identifier and the currently idle inference resources carried by the first task request, determine a target inference request that matches the first task request, and send the target inference request;

[0009] The first reasoning service is used to send a first task request; further used to obtain a target reasoning request that matches the first task request; respond to the target reasoning request to perform generative reasoning to obtain at least part of the reasoning result of the target reasoning request.

[0010] According to another aspect of the present disclosure, there is provided an inference device, comprising:

[0011] an acquiring unit, configured to receive a first task request through a first scheduling service among the multiple scheduling services, wherein the first task request is a request sent by the first inference service among the multiple inference services and carries at least a model identifier and currently idle inference resources of the first inference service;

[0012] a scheduling unit, configured to determine a target inference request that matches the first task request based on a model identifier and idle inference resources carried by the first task request;

[0013] A sending unit is configured to send the target inference request to a first inference service through the first scheduling service.

[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0020] In this way, the disclosed solution can obtain the current load situation of the first inference service (such as idle inference resources) based on the first task request actively sent by the first inference service, and then enable the scheduling service to determine the target inference request that matches the load situation, and send the target inference request to the first inference service for inference. Compared with the traditional polling strategy, the disclosed solution fully combines the characteristics of the large model inference service, effectively realizes the reasonable scheduling of resources, avoids idle resources due to load imbalance, and at the same time, ensures that the inference requests are processed in a timely manner, thereby improving resource utilization and ensuring system stability.

[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0023] Figure 1 This is a diagram showing the scheduling results of an inference server cluster in an example;

[0024] Figure 2 This is a schematic flow chart of the reasoning method according to an embodiment of the present application. Figure 1 ;

[0025] Figure 3 This is a schematic flow chart of the reasoning method according to an embodiment of the present application. Figure 2 ;

[0026] Figure 4 is a schematic diagram of the structure of an inference system according to an embodiment of the present application;

[0027] Figure 5 is a schematic diagram of a scenario in a specific example of an inference system according to an embodiment of the present application;

[0028] Figure 6 is a structural diagram of an inference device according to an embodiment of the present application;

[0029] Figure 7 is a block diagram of an electronic device for implementing the inference method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0032] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0033] With the rapid development of artificial intelligence (AI), large models have become a key research direction in the field of artificial intelligence (AI). They can learn rich knowledge representations from massive amounts of data, thereby demonstrating superior performance in various tasks. However, to apply these powerful large models in real-world scenarios, they need to be service-oriented. For example, by providing services through application programming interfaces (APIs), other applications can easily access the capabilities of the large models.

[0034] In actual applications, in order to meet the needs of a large number of users, large-model inference services are usually deployed in clusters. This cluster deployment faces two challenges: (1) improving the system's resource utilization; (2) ensuring the stability of the system. Among them, improving the system's resource utilization can avoid resource waste and thus respond to more users' concurrent requests more quickly; and the stability of the system is the basic quality of a reliable commercial system. To meet the above challenges, an efficient load balancing algorithm is needed to schedule cluster resources. However, the existing cluster load balancing components are usually general-purpose and difficult to reasonably schedule resources based on the characteristics of large-model inference. This is because the existing load balancing components use general load scheduling strategies such as polling and minimum number of connections, and cannot perceive the actual load of the backend inference cluster. Therefore, it is easy to cause load imbalance in the inference cluster, resulting in wasted computing resources.

[0035] For example, suppose the backend inference server cluster has two instances, such as Figure 1 As shown in the figure, they are Inference Service 1 and Inference Service 2. The gateway responsible for load balancing uses a polling strategy to distribute traffic. The number of concurrent requests is 2. Now three requests need to be sent. The generation lengths of the models corresponding to these three requests are 128, 512, and 128 respectively. At this time, the request process is:

[0036] (1) The gateway (a unified entry point for accessing the service-oriented cluster, acting as a reverse proxy and supporting common load balancing algorithms) sends request 1 to inference service 1 and request 2 to inference service 2. Since request 1 is generated in a shorter length, inference service 1 will complete the inference earlier.

[0037] (2) After receiving the inference result of request 1, the target object sends request 3. At this time, the gateway may send request 3 to inference service 2; inference service 2 is still processing inference request 2, while inference service 1 is idle.

[0038] From the above example, we can see that if request 3 should be sent to the idle inference service 1, the polling strategy causes the gateway to send request 3 to the inference service 2, which in turn causes load imbalance in the inference server cluster.

[0039] Based on this, the disclosed solution proposes an inference method that can fully combine the characteristics of large-model inference for graphics card usage, actively perceive the load and hardware capacity of a single model inference service on the backend, and collaboratively optimize to keep the load of the inference server cluster balanced, thereby better improving the overall system resource utilization and ensuring system stability.

[0040] Specifically, Figure 2 This is a schematic flow chart of the reasoning method according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0041] Furthermore, the method includes at least part of the following content. It should be noted that the disclosed solution can be specifically applied to a scheduling server or a scheduling server cluster; further, the disclosed solution can be integrated into a scheduling service component (or load balancing component). In this case, the scheduling service component can be integrated into any device, such as a gateway, so that the device integrated with the scheduling service component has the scheduling capabilities described in the disclosed solution, thereby efficiently completing inference tasks.

[0042] Specifically, if Figure 2 As shown, including:

[0043] Step S201: receiving a first task request through a first scheduling service among a plurality of scheduling services.

[0044] Here, the first task request is a request sent by a first reasoning service among multiple reasoning services, and carries at least a model identifier and currently idle reasoning resources of the first reasoning service.

[0045] It's important to note that an inference service can be understood as a computing service that uses pre-trained machine learning models to process input data and output predictions or decision results. It's a core component of artificial intelligence (AI) and machine learning (ML) applications and can be executed on an inference server. Specifically, an inference service refers to a machine learning model deployed on a server or server cluster that receives data, processes it, and returns inference results.

[0046] Correspondingly, the scheduling service can also be specifically understood as being executed on a scheduling server or a scheduling server cluster and being responsible for scheduling.

[0047] That is, in this example, the first reasoning service proactively sends a first task request to the first scheduling service, and carries its own idle reasoning resources and model identifier, to proactively request to execute the reasoning task that it is capable of.

[0048] Furthermore, in one example, in order to meet multiple concurrency requirements, a gateway can also be set up. At this time, the first inference service can send the first task request carrying at least the model identifier and the current idle inference resources of the first inference service to the first scheduling service via the gateway.

[0049] Furthermore, in one example, the model identifier carried by the first task request can be specifically understood as a unique identifier (also referred to as a model identifier) ​​of the model used by the first inference service, for example, can be recorded as model_id.

[0050] Furthermore, in one example, the idle inference resources include but are not limited to load conditions and remaining hardware capacity.

[0051] Step S202: Based on the model identifier and idle inference resources carried by the first task request, determine a target inference request that matches the first task request.

[0052] Here, the target reasoning request carries at least the target reasoning task to be executed and the model identifier of the target model required to execute the target reasoning task (ie, the unique identifier of the target model).

[0053] That is, the first scheduling service determines a target inference request that matches the first task request based on the model identifier and idle inference resources carried by the first task request. In other words, the first scheduling service determines, for the first inference service, the target inference request that the first inference service is capable of handling based on the model identifier and idle inference resources.

[0054] Here, in one example, the target model may be specifically a large model, further, may be specifically a large language model, or may be other models, which is not limited in the present disclosure.

[0055] Step S203: Send the target reasoning request to the first reasoning service through the first scheduling service, so that the first reasoning service executes the target reasoning task carried by the target reasoning request.

[0056] In this way, the disclosed solution can obtain the current load situation of the first inference service (such as idle inference resources) based on the first task request actively sent by the first inference service, and then enable the scheduling service to determine the target inference request that matches the load situation, and send the target inference request to the first inference service for inference. Compared with the traditional polling strategy, the disclosed solution fully combines the characteristics of the large model inference service, effectively realizes the reasonable scheduling of resources, avoids idle resources due to load imbalance, and at the same time, ensures that the inference requests are processed in a timely manner, thereby improving resource utilization and ensuring system stability.

[0057] Furthermore, since the scheduling service and reasoning service in the disclosed solution can be deployed in a clustered manner, whether the scheduling service or the reasoning service fails, it will not affect the normal function of the overall system, so that the system has a certain fault tolerance and high availability, thereby further ensuring the reliability and stability of the system operation.

[0058] Furthermore, in a specific example, the target inference request can be obtained in the following manner. Specifically, the above-described determination of a target inference request matching the first task request based on the model identifier and idle inference resources carried by the first task request (e.g., step S202) specifically includes:

[0059] A target inference request that matches the model identifier in the first task request and for which the idle inference resources meet the inference requirements is selected from a preset request queue of the target database through the first scheduling service.

[0060] Here, the preset request queue stores a plurality of inference requests, and each inference request carries at least an inference task and a model identifier of a target model required for the inference task.

[0061] That is to say, after receiving the first task request, the first scheduling service will select an inference request from the preset request queue that matches the model identifier carried by the first task request and the idle inference resources, and the inference request that meets the inference requirements of the inference task as the target inference request, and then send the target inference request to the first inference service.

[0062] In this way, the disclosed solution can intelligently select inference requests that match the current actual load situation from the target database based on the current actual load situation of the inference service. In this way, reasonable allocation of resources is achieved, and resource idleness due to load imbalance is effectively avoided, thereby improving the utilization rate of computing resources.

[0063] Figure 3 This is a schematic flow chart of the reasoning method according to an embodiment of the present application. Figure 2 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 2 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0064] Furthermore, the method includes at least part of the following contents. Figure 3 As shown, including:

[0065] Step S301: receiving a first task request through a first scheduling service among a plurality of scheduling services.

[0066] Here, the first task request is a request sent by a first reasoning service among multiple reasoning services, and carries at least a model identifier and currently idle reasoning resources of the first reasoning service.

[0067] Here, for related contents such as scheduling services and reasoning services, please refer to the above description and will not be repeated here.

[0068] Step S302: Determine, through the first scheduling service, the target inference resource required for the inference task based on the characteristic information of the inference task carried in the inference request.

[0069] For example, in one example, the first scheduling service can determine the target inference resources required for each inference request in the preset request queue, thereby facilitating subsequent matching; or, in another example, the first scheduling service determines the target inference resources required for the inference request with the earliest timing one by one based on the first-in-first-out output timing of the preset request queue, and then performs matching.

[0070] For example, in one example, the characteristic information of the reasoning task carried by the reasoning request may specifically be the input information and the required output requirements of the reasoning task.

[0071] Here, the execution order of step S301 and step S302 can be interchanged. Alternatively, in one example, the first scheduling service can pre-determine the target inference resource required for each inference request, thereby laying the foundation for subsequent rapid matching of the target inference request.

[0072] Step S303: Through the first scheduling service, select a target inference request that matches the model identifier in the first task request and whose target inference resource is not greater than the idle inference resource from the multiple inference requests included in the preset request queue.

[0073] Step S304: Send the target inference request to the first inference service through the first scheduling service.

[0074] That is to say, in one example, after the first scheduling service receives the first task request, it can first determine the target inference resources required for the inference request based on the input information and output requirements of the inference task carried by the inference request in the preset request queue (for example, the Redis request queue) of the target database (for example, the Redis database); secondly, the first scheduling service, based on the model identifier carried by the first task request and the current idle inference resources of the first inference service, selects an inference request (that is, a target inference request) from the Redis request queue that matches the model identifier in the first task request and whose target inference resources are not greater than the current idle inference resources of the first inference service; finally, the first scheduling service sends the selected target inference request to the first inference service so that the inference task carried by the target inference request is executed through the first inference service.

[0075] In this way, the disclosed solution can utilize the characteristic information of the inference task carried by the inference request in the preset request queue to determine the target inference resource required for the inference task, thereby enabling the first scheduling service to select an inference request (that is, a target inference request) whose model identifier and target inference resource meet the requirements from multiple inference requests based on the model identifier in the first task request and the current idle inference resources of the first inference service, and send it to the first inference service through the first scheduling service. In this way, the first scheduling service provides the first inference service with inference tasks that meet its requirements, thereby effectively realizing the reasonable allocation of resources, avoiding resource idleness due to load imbalance, and thus improving the utilization rate of computing resources.

[0076] In a specific example of the disclosed solution, a preset request queue may be constructed in the following manner. Specifically, before determining a target inference request that matches the first task request, for example, before step S102 or step S302 described above, the method further includes:

[0077] A target inference request is obtained through a second scheduling service among the multiple scheduling services, and the target inference request is stored in the preset request queue through the second scheduling service.

[0078] Here, in one example, the second scheduling service is different from the first scheduling service.

[0079] That is, in this example, the scheduling service that stores the target inference request and the scheduling service that allocates the inference service to the target inference request may be different scheduling services.

[0080] In addition, it should be noted that, in actual applications, the second scheduling service and the first scheduling service may also be the same scheduling service.

[0081] For example, in one example, the target object distributes the target inference request to the second scheduling service among the multiple scheduling services through a gateway (for example, the gateway is provided with a distribution strategy); after receiving the target inference request, the second scheduling service writes at least the request ID (unique identifier, which can be recorded as user_id) carried by the target inference request to the Redis request queue of the Redis database for temporary storage. Alternatively, in order to facilitate subsequent matching, the second scheduling service may also write the request ID and the model identifier of the model required for the inference task carried by the target inference request into the Redis request queue. Alternatively, in another example, in order to achieve subsequent rapid matching success, the second scheduling service may also write the request ID, the model identifier of the model required for the inference task carried by the target inference request, and the target inference resources required by the target inference request into the Redis request queue.

[0082] In this way, the disclosed solution can timely store the acquired target inference request into the preset request queue via the second scheduling service. In this way, the inference request is managed with the help of the queue mechanism, laying the foundation for subsequent more efficient resource allocation and service load balancing.

[0083] In a specific example of the presently disclosed solution, after the first inference service receives the target inference request sent by the first scheduling service and executes the corresponding inference task to obtain the corresponding inference result, the method further includes: obtaining at least part of the inference result for the target inference request through the second scheduling service; further, sending at least part of the inference result for the target inference request through the second scheduling service.

[0084] Here, at least part of the inference result is at least part of the result generated after the first inference service executes the target inference request.

[0085] That is to say, in this example, the scheduling service that stores the target reasoning request, that is, the second scheduling service, will also be responsible for feeding back the reasoning results of the target reasoning request. For example, it is responsible for feeding back the reasoning results of the target reasoning request to the target object that initiates the target reasoning request, so as to ensure that the reasoning request of the target object can be responded to quickly.

[0086] In this way, the disclosed solution can utilize different scheduling services to achieve different scheduling purposes, thereby improving response efficiency while ensuring improved resource utilization and system stability, thereby effectively improving user experience.

[0087] Further, in a specific example, the above-mentioned obtaining at least part of the inference results for the target reasoning request through the second scheduling service specifically includes: querying the preset result queue through the second scheduling service whether there is at least part of the inference results corresponding to the target reasoning request; further, if it is determined that there is, obtaining at least part of the inference results corresponding to the target reasoning request from the preset result queue.

[0088] In this example, the preset result queue stores the inference results returned by at least some of the multiple inference services.

[0089] For example, in one example, after the first inference service executes the target inference request and generates at least part of the inference result of the target inference request, the first inference service sends at least part of the inference result of the target inference request to the first scheduling service through the gateway; accordingly, the first scheduling service stores at least part of the received inference result of the target inference request in a preset result queue (for example, the Redis result queue of the Redis database), so that the second scheduling service can obtain the inference result from the Redis result queue.

[0090] It is understandable that the preset result queue not only stores at least part of the inference results of the target inference request, but also can store the inference results of other inference requests. Furthermore, in actual applications, the preset result queue stores the request ID and its corresponding inference result; at this time, the second scheduling service can query the preset result queue based on the request ID of the target inference request to see whether there is an inference result corresponding to the request ID of the target inference request. If so, the second scheduling service directly obtains the inference result of the target inference request and sends the inference result of the target inference request through the gateway, for example, to the target object that initiated the target inference request.

[0091] In this way, the disclosed solution can utilize different scheduling services to achieve asynchronous processing and rapid response, thereby effectively improving the utilization of system resources and ensuring the stability of the system while improving response efficiency, thereby significantly improving the user experience.

[0092] The present disclosure provides an inference system, such as Figure 4 As shown, including:

[0093] The first scheduling service 401 is configured to receive a first task request, wherein the first task request carries at least a model identifier and currently idle inference resources of the first inference service; based on the model identifier and the currently idle inference resources carried by the first task request, determine a target inference request that matches the first task request, and send the target inference request;

[0094] The first reasoning service 402 is used to send a first task request; further used to obtain a target reasoning request that matches the first task request; respond to the target reasoning request to perform generative reasoning to obtain at least part of the reasoning result of the target reasoning request.

[0095] In this way, the disclosed solution can obtain the current load situation of the first inference service (such as idle inference resources) based on the first task request actively sent by the first inference service, and then enable the scheduling service to determine the target inference request that matches the load situation, and send the target inference request to the first inference service for inference. Compared with the traditional polling strategy, the disclosed solution fully combines the characteristics of the large model inference service, effectively realizes the reasonable scheduling of resources, avoids idle resources due to load imbalance, and at the same time, ensures that the inference requests are processed in a timely manner, thereby improving resource utilization and ensuring system stability.

[0096] Furthermore, since the scheduling service and reasoning service in the disclosed solution can be deployed in a clustered manner, whether the scheduling service or the reasoning service fails, it will not affect the normal function of the overall system, so that the system has a certain fault tolerance and high availability, thereby further ensuring the reliability and stability of the system operation.

[0097] In a specific example of the disclosed solution, the reasoning system further includes:

[0098] A second scheduling service is used to obtain a target reasoning request and store the target reasoning request in a preset request queue;

[0099] A database is used to store preset request queues.

[0100] In this way, the disclosed solution can timely store the acquired target inference request into the preset request queue via the second scheduling service. In this way, the inference request is managed with the help of the queue mechanism, laying the foundation for subsequent more efficient resource allocation and service load balancing.

[0101] In a specific example of the disclosed solution, the second scheduling service is further configured to obtain at least an inference result corresponding to the target inference request and to send at least a portion of the inference result for the target inference request. This allows different scheduling services to be used to achieve different scheduling objectives, thereby improving response efficiency while ensuring increased resource utilization and system stability, thereby effectively enhancing the user experience.

[0102] It should be noted that for the relevant content of the first scheduling service, the second scheduling service, and the first reasoning service, please refer to the above description and will not be repeated here.

[0103] The following detailed description of the disclosed solution is based on specific examples. This solution proposes a load balancing system for large-model service-based deployment clusters. This system leverages the characteristics of large-model inference services (e.g., graphics card usage) to achieve load optimization through the collaborative work of a load balancing component (also known as the scheduling service described above) and an inference server cluster. The core concept is to enable the load balancing component to proactively perceive the load status of the backend inference service, such as the current remaining hardware resources (e.g., graphics memory size) and the remaining available slots for inference. This allows for end-to-end collaborative optimization, improving overall system resource utilization and ensuring stability.

[0104] Furthermore, in one example, the load balancing component distributes user requests (i.e., corresponding to the aforementioned inference requests) to matching inference services for processing by determining the relationship between the resources required for the request and the currently available inference resources of the inference service. Compared to traditional polling strategies, the disclosed solution is aware of the load on the backend inference service. Furthermore, the inference service proactively "pulls" user requests from the load balancing component, rather than pushing user requests without any awareness. This ensures that user requests are always executed on the appropriate inference service, effectively avoiding idle resources caused by load imbalance.

[0105] Before introducing the implementation steps of the disclosed solution, some terms used in this example are explained:

[0106] (1) Gateway: Traffic entrance; specifically, a unified entrance for accessing the service-based cluster, acting as a reverse proxy.

[0107] (2) Scheduling server cluster: used for load balancing processing, which can also be understood as the load balancing component of this example, mainly including multiple scheduling services, for example, which can be recorded as scheduling service 1, scheduling service 2, ..., scheduling service n.

[0108] (3) Remote dictionary server (Redis): used to store user requests and inference results, mainly used for data transfer.

[0109] (4) Large model inference server cluster: contains multiple inference services (for example, can be recorded as inference service 1, inference service 2, ..., inference service n), mainly used to perform specific inference tasks.

[0110] It should be noted that before executing the core steps of this disclosure, the following preparatory work needs to be done:

[0111] Specifically, if Figure 5 As shown, the core steps of the disclosed solution include:

[0112] Step S500: Preliminary work: Set a unique identifier (which can be recorded as model_id) for the model in the inference server cluster so that when the inference service subsequently pulls user requests, it only pulls requests with the unique identifier model_id.

[0113] Step S501: The scheduling service responds to the user request.

[0114] Step S5011: the target object sends the target user request to the gateway, so that the gateway distributes the target user request to the scheduling service of the scheduling server cluster, for example, to the scheduling service 2 (corresponding to the second scheduling service described above).

[0115] Step S5012: the scheduling service 2 writes the received target user request into the Redis request queue (corresponding to the above preset request queue) of the Redis database for temporary storage.

[0116] Here, the target user request may specifically refer to requesting the target model to perform a specified reasoning task. In this case, the target user request may specifically carry a unique identifier of the requested target model and the required reasoning task.

[0117] For example, in one example, the scheduling service 2 may write the request id (eg, user_id) of the target user request and the unique identifier model_id of the target model required for the target user request into the Redis request queue for temporary storage.

[0118] Step S502: The scheduling service responds to the inference request of the inference service.

[0119] Step S5021: The inference service 1 of the inference server cluster (corresponding to the first inference service described above) sends a task pull request (used to obtain a user request, corresponding to the first task request described above) to the gateway, so as to send the task pull request to the scheduling service of the scheduling server cluster through the gateway, for example, to the scheduling service 1 (corresponding to the first scheduling service described above).

[0120] Here, the task pull request may specifically carry a unique identifier of the model used by the inference service (also referred to as a model identifier, which may be recorded as model_id) and the current idle inference resources of the inference service 1.

[0121] Step S5022: Scheduling service 1 determines the user request that matches the task pull request from the Redis request queue based on the model identifier and idle inference resources carried by the task pull request, such as determining the target user request that matches the task pull request, and pops the target user request out of the queue, and then sends it to inference service 1 through the gateway.

[0122] For example, in one example, the scheduling service 1 can determine the target inference resources required for the user requests contained in the Redis request queue, and select from the user requests contained in the Redis request queue the target user request that matches the model identifier in the task pull request and whose required target inference resources are no greater than the idle inference resources.

[0123] For example, in one example, the target reasoning resource required for the user request is obtained based on feature information (such as input information and required output requirements) of the reasoning task carried by the user request.

[0124] Step S5023: After receiving the target user request, the inference service 1 performs generative reasoning to obtain at least part of the inference results of the target user request, and sends at least part of the inference results of the target user request to the scheduling service 1 through the gateway, so that the scheduling service 1 writes at least part of the inference results of the target user request into the Redis result queue for temporary storage.

[0125] Step S5024: Scheduling service 2 queries the Redis result queue to see whether there is at least part of the inference result corresponding to the target user request. Further, if it is determined that there is, at least part of the inference result corresponding to the target user request is obtained from the Redis result queue (for example, at least part of the inference result corresponding to the target user request is popped out of the queue), and returned to the target object in streaming or whole sentence format.

[0126] In this way, the disclosed solution can fully combine the characteristics of large-scale model reasoning for the use of graphics cards, actively perceive the load and hardware capacity of the back-end single model reasoning service, thereby achieving end-to-end collaborative optimization and better improving the overall system resource utilization and system stability.

[0127] In summary, the system architecture of the reasoning system used in the disclosed solution has the following advantages:

[0128] First, the inference service actively pulls user requests from the scheduling service. Moreover, when pulling user requests, it will actively report the current load status of the inference service (such as idle inference resources, etc.). At this time, the scheduling service, as a load balancing component, can perceive the load status of the backend inference service and use this as a basis to select user requests that match the inference service, thereby achieving real and effective load balancing for the inference server cluster.

[0129] Second, the scheduling service can also adopt a clustered deployment solution. When one or more scheduling services fail, the inference service can make a judgment based on the communication delay between the current inference service and the scheduling service, and then reconnect to the normal scheduling service;

[0130] Third, inference services can also adopt a clustered deployment solution. When one or more inference services fail and cannot pull user requests, it will not affect the normal function of the entire system, thereby ensuring the reliability and stability of the system.

[0131] The present disclosure provides an inference device, such as Figure 6 As shown, including:

[0132] An acquiring unit 601 is configured to receive a first task request through a first scheduling service among the multiple scheduling services, wherein the first task request is a request sent by the first inference service among the multiple inference services and carries at least a model identifier and currently idle inference resources of the first inference service;

[0133] The scheduling unit 602 is configured to determine a target inference request that matches the first task request based on the model identifier and the idle inference resources carried by the first task request;

[0134] The sending unit 603 is configured to send the target inference request to the first inference service through the first scheduling service.

[0135] In a specific example of the disclosed solution, the scheduling unit is specifically configured to:

[0136] Selecting, through a first scheduling service, from a preset request queue of a target database, a target inference request that matches the model identifier in the first task request and for which the idle inference resources meet the inference requirements;

[0137] The preset request queue stores a plurality of inference requests, and each inference request carries at least an inference task and a model identifier of a model required to be used in the inference task.

[0138] In a specific example of the disclosed solution, the scheduling unit is specifically configured to:

[0139] Determine the target inference resource required for the inference task based on the feature information of the inference task carried in the inference request;

[0140] A target inference request that matches the model identifier in the first task request and whose target inference resource is not greater than the idle inference resource is selected from the multiple inference requests included in the preset request queue.

[0141] In a specific example of the disclosed solution, the acquiring unit is further configured to acquire the target inference request through a second scheduling service among the multiple scheduling services; wherein the second scheduling service is different from the first scheduling service;

[0142] The scheduling unit is further configured to store the target inference request in the preset request queue through the second scheduling service.

[0143] In a specific example of the disclosed solution, the acquiring unit is further configured to acquire, through the second scheduling service, at least part of the inference results for the target inference request; wherein at least part of the inference results are generated after the first inference service executes the target inference request;

[0144] The sending unit is further configured to send at least part of the inference result for the target inference request through the second scheduling service.

[0145] In a specific example of the disclosed solution, the acquisition unit is specifically configured to:

[0146] querying, through the second scheduling service, a preset result queue whether there are at least some inference results corresponding to the target inference request; wherein the preset result queue stores inference results returned by at least some of the multiple inference services;

[0147] If it is determined that the target reasoning request exists, at least part of the reasoning results corresponding to the target reasoning request is obtained from the preset result queue.

[0148] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0149] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0150] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0151] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0152] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. Computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0153] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0154] The computing unit 701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the inference method. For example, in some embodiments, the inference method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the inference method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the inference method by any other suitable means (e.g., via firmware).

[0155] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0156] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0157] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0158] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0159] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0160] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0161] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0162] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A reasoning method, comprising: Receiving a first task request through a first scheduling service among the multiple scheduling services, wherein the first task request is a request sent by the first inference service among the multiple inference services and carries at least a model identifier and currently idle inference resources of the first inference service; Determining a target inference request that matches the first task request based on the model identifier and the idle inference resources carried by the first task request; sending the target inference request to the first inference service via the first scheduling service; The step of determining a target inference request that matches the first task request based on the model identifier and the idle inference resources carried by the first task request includes: Determining, through the first scheduling service, target inference resources required for the inference task based on characteristic information of the inference task carried by the inference request contained in a preset request queue of the target database; the preset request queue stores multiple inference requests, each of which carries at least the inference task and a model identifier of a model required to be used by the inference task; A target inference request that matches the model identifier in the first task request and whose target inference resource is not greater than the idle inference resource is selected from the multiple inference requests included in the preset request queue.

2. The method according to claim 1, further comprising: Obtaining a target inference request through a second scheduling service among the multiple scheduling services; wherein the second scheduling service is different from the first scheduling service; The target inference request is stored in the preset request queue through the second scheduling service.

3. The method according to claim 2, further comprising: Obtaining, through the second scheduling service, at least part of the inference results for the target inference request; wherein at least part of the inference results are generated after the first inference service executes the target inference request; At least part of the reasoning result for the target reasoning request is sent through the second scheduling service, so that the target object corresponding to the target reasoning request obtains at least part of the reasoning result.

4. The method according to claim 3, wherein: The obtaining, through the second scheduling service, at least part of the inference result for the target inference request includes: querying, through the second scheduling service, a preset result queue whether there are at least some inference results corresponding to the target inference request; wherein the preset result queue stores inference results returned by at least some of the multiple inference services; If it is determined that the target reasoning request exists, at least part of the reasoning results corresponding to the target reasoning request is obtained from the preset result queue.

5. An inference system comprising: A first scheduling service is configured to receive a first task request, wherein the first task request carries at least a model identifier and currently idle inference resources of the first inference service; based on the model identifier and the currently idle inference resources carried by the first task request, determine a target inference request that matches the first task request, and send the target inference request; A first reasoning service is configured to send a first task request; further configured to obtain a target reasoning request that matches the first task request; and respond to the target reasoning request to perform generative reasoning and obtain at least a partial reasoning result of the target reasoning request; The first scheduling service is specifically used for: Determining target inference resources required for the inference task based on characteristic information of the inference task carried by the inference request contained in a preset request queue of the target database; the preset request queue stores multiple inference requests, each of which carries at least the inference task and a model identifier of a model required for the inference task; A target inference request that matches the model identifier in the first task request and whose target inference resource is not greater than the idle inference resource is selected from the multiple inference requests included in the preset request queue.

6. The reasoning system according to claim 5, further comprising: A second scheduling service is used to obtain a target reasoning request and store the target reasoning request in a preset request queue; A database is used to store preset request queues.

7. The inference system according to claim 6, wherein: The second scheduling service is further used to obtain at least part of the reasoning results corresponding to the target reasoning request, and send at least part of the reasoning results for the target reasoning request.

8. An inference device, comprising: an acquiring unit, configured to receive a first task request through a first scheduling service among the multiple scheduling services, wherein the first task request is a request sent by the first inference service among the multiple inference services and carries at least a model identifier and currently idle inference resources of the first inference service; a scheduling unit, configured to determine a target inference request that matches the first task request based on a model identifier and idle inference resources carried by the first task request; a sending unit, configured to send the target inference request to a first inference service via the first scheduling service; The scheduling unit is specifically configured to: Determining, through the first scheduling service, target inference resources required for the inference task based on characteristic information of the inference task carried by the inference request contained in a preset request queue of the target database; the preset request queue stores multiple inference requests, each of which carries at least the inference task and a model identifier of a model required to be used by the inference task; A target inference request that matches the model identifier in the first task request and whose target inference resource is not greater than the idle inference resource is selected from the multiple inference requests included in the preset request queue.

9. The device according to claim 8, wherein The acquiring unit is further configured to acquire the target inference request through a second scheduling service among the multiple scheduling services; wherein the second scheduling service is different from the first scheduling service; The scheduling unit is further configured to store the target inference request in the preset request queue through the second scheduling service.

10. The device according to claim 9, wherein The acquiring unit is further configured to acquire, through the second scheduling service, at least part of the inference results for the target inference request; wherein at least part of the inference results are generated after the first inference service executes the target inference request; The sending unit is further configured to send at least part of the inference result for the target inference request through the second scheduling service, so that the target object corresponding to the target inference request obtains at least part of the inference result.

11. The device according to claim 10, wherein The acquisition unit is specifically configured to: querying, through the second scheduling service, a preset result queue whether there are at least some inference results corresponding to the target inference request; wherein the preset result queue stores inference results returned by at least some of the multiple inference services; If it is determined that the target reasoning request exists, at least part of the reasoning results corresponding to the target reasoning request is obtained from the preset result queue.

12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Large model reasoning system and method and cloud server

    CN117119049A

  • Model reasoning task processing method and device, computer equipment and medium

    CN117196036A