Load balancing method and device for reasoning service, electronic equipment and storage medium
By creating a replica set of proxy services in the inference service node and performing load balancing operations, the resource bottlenecks and service congestion problems when nodes are under great pressure in large-scale deployment environments are solved, and inference efficiency and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202510851598.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
In a large-scale deployment environment, how to distribute user requests to the appropriate inference nodes based on the current operating status of the system, reduce resource bottlenecks and service congestion, especially when the node cannot provide services, the inference efficiency is poor.
By obtaining the load status information of the inference service node, creating a replica set of proxy services, and performing load balancing operations between proxy services and replicas, distributing session requests based on request allocation policies, releasing unnecessary replicas to replace proxy services, realizing elastic recycling and resource utilization.
It reduces service interruptions and resource bottlenecks when nodes are under great pressure, improves overall throughput capability and request processing efficiency, and realizes flexible recycling and efficient utilization of resources.
Smart Images

Figure CN120358236A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, and a storage medium for load balancing of inference services. Background Art
[0002] With the development of science and technology, artificial intelligence technology has been continuously evolving, and inference services have become one of the core forces driving the intelligent transformation of society. For example, the forwarding path of a user request can be determined based on resource usage information. However, in a large-scale deployment environment, when a node fails to provide services, how to distribute user requests to appropriate inference nodes according to the current operating state of the system to reduce resource bottlenecks and service congestion has become the focus of research. Summary of the Invention
[0003] The present disclosure provides a method, an apparatus, an electronic device, and a storage medium for load balancing of inference services. Its main purpose is to solve the problem of poor inference efficiency when the node pressure of inference service nodes is relatively large.
[0004] According to a first aspect of the present disclosure, there is provided a method for load balancing of inference services, including: During the process of an inference service node set processing a session request, obtaining load status information of any inference service node in the inference service node set; When the load status information indicates that the node pressure of the any inference service node is greater than a pressure threshold, creating a proxy service corresponding to the any inference service node; Creating a replica set of the proxy service according to the node pressure; According to a request distribution policy, distributing the session request corresponding to the any inference service node to each replica in the replica set for processing, and performing a load balancing operation between the proxy service and each replica; When the pressure information corresponding to the replica set meets the pressure requirement, determining a first replica in the replica set, releasing replicas other than the first replica in the replica set, and replacing the proxy service with the first replica.
[0005] According to a second aspect of the present disclosure, there is provided a load balancing apparatus for inference services, including: An information acquisition set for obtaining load status information of any inference service node in the inference service node set during the process of the inference service node set processing a session request; A service creation unit for creating a proxy service corresponding to the any inference service node when the load status information indicates that the node pressure of the any inference service node is greater than a pressure threshold; A set creation unit for creating a set of replicas of the proxy service according to the node pressure; A request distribution unit for distributing the session requests corresponding to any one of the inference service nodes to each replica in the set of replicas for processing according to a request allocation policy, and performing a load balancing operation between the proxy service and each replica; A proxy release unit for determining a first replica in the set of replicas, releasing the replicas other than the first replica in the set of replicas, and replacing the proxy service with the first replica when the pressure information corresponding to the set of replicas meets the pressure requirements.
[0006] According to a third aspect of the present disclosure, there is provided an electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the foregoing first aspect.
[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.
[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method described in the foregoing first aspect when executed by a processor.
[0009] Through the present disclosure, during the process of processing a session request by an inference service node set, the load status information of any inference service node in the inference service node set is obtained; when the load status information indicates that the node pressure of the any inference service node is greater than a pressure threshold, a proxy service corresponding to the any inference service node is created; according to the node pressure, a replica set of the proxy service is created; according to a request distribution policy, the session request corresponding to the any inference service node is distributed to each replica in the replica set for processing, and a load balancing operation is performed between the proxy service and each replica; when the pressure information corresponding to the replica set meets the pressure requirement, a first replica is determined in the replica set, replicas other than the first replica in the replica set are released, and the first replica is used to replace the proxy service. Therefore, when the node pressure of a certain node is relatively large, load balancing can be performed through the proxy service, the situation that the service is interrupted due to the inability of the node to process the session request can be reduced, the resource bottleneck and service congestion can be reduced, load balancing operations can be performed among multiple replicas, the overall throughput capacity can be improved, the request processing efficiency can be improved, and when the pressure is reduced, degradation processing can be performed, only one replica is retained, resources that are not needed can be released, elastic recovery and resource utilization efficiency can be realized, and thus the inference efficiency when the node pressure is relatively large can be improved while the resource utilization efficiency is improved.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them: Figure 1 is a schematic flowchart of a load balancing method for an inference service provided by an embodiment of the present disclosure; Figure 2 is a schematic flowchart of another load balancing method for an inference service provided by an embodiment of the present disclosure; Figure 3 is an example schematic diagram of a load balancing method provided by an embodiment of the present disclosure; Figure 4 is an example schematic diagram of a key-value cache synchronization method provided by an embodiment of the present disclosure; Figure 5 is an example schematic diagram of a proxy service removal method provided by an embodiment of the present disclosure; Figure 6 is a schematic structural diagram of a load balancing device for an inference service provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0013] According to some embodiments, a pre-trained model can, for example, possess powerful language understanding and generation capabilities. It has not only played an important role in multiple fields such as education, healthcare, law, finance, and scientific research, but also profoundly changed people's production and lifestyle. Among them, in the educational scenario, the pre-trained model can serve as a personalized intelligent tutor to help students plan their learning according to their own levels; in the healthcare scenario, it can assist doctors in medical record analysis and preliminary diagnosis to improve the diagnosis and treatment efficiency; in various professional tasks, the application of the pre-trained model has significantly reduced the threshold for obtaining complex knowledge, enabling more people to participate equally in the knowledge society, and greatly promoting information fairness and knowledge inclusiveness. However, with the continuous expansion of the service scale of the pre-trained model and the increasing complexity of application scenarios, the technical infrastructure supporting it also faces huge challenges. Pre-trained model inference is not only computationally intensive but also has extremely high requirements for real-time performance and availability. Therefore, load balancing, as the core scheduling mechanism connecting user requests and inference nodes, has become increasingly important. It not only concerns the performance of the service but also determines the key factor for the stable operation and flexible expansion of the entire system.
[0014] In some embodiments, load balancing is not only the key scheduler for the efficient and stable operation of the pre-trained model service but also an indispensable cornerstone in its evolution towards large-scale, real-time, and intelligent. An efficient load balancing mechanism, for example, needs to comprehensively consider multi-dimensional information such as the usage rate of the Graphics Processing Unit (GPU), the usage rate of the Central Processing Unit (CPU), memory occupancy, response latency, Key-Value Cache (kv cache) hit rate, and network conditions of model replicas, so that each request can be processed along the optimal path. Among them, in the pre-trained model service with multi-node and multi-replica deployment, the load balancer also undertakes functions such as fault tolerance, hot migration, and traffic control. When a node fails or is overloaded, it can promptly switch the traffic to healthy nodes to ensure that the service does not interrupt and the user experience does not decline. Therefore, how to perform load balancing has become the focus of attention.
[0015] The load balancing method, device, electronic device, and storage medium for an inference service according to embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0016] Figure 1 It is a flowchart showing a load balancing method for an inference service provided by embodiments of the present disclosure.
[0017] As Figure 1 shown, the method includes the following steps: Step 101, during the process of the inference service node set processing a session request, obtain the load status information of any inference service node in the inference service node set; According to some embodiments, the execution subject of the embodiments of the present disclosure may be an electronic device, for example. The electronic device may be a server, for example. The electronic device does not specifically refer to a certain fixed device. For example, when the device identifier corresponding to the electronic device changes, the electronic device may also change accordingly.
[0018] In some embodiments, the inference service node set may be a collective formed by at least one inference service node, for example. The inference service node may be used to process session requests, for example. The inference service node set does not specifically refer to a certain fixed set. For example, when the number of services corresponding to the inference service node set changes, the inference service node set may also change accordingly. For example, when a certain inference service node in the inference service node set changes, the inference service node set may also change accordingly. The inference service node set may also be referred to as a backend service node set, for example.
[0019] According to some embodiments, the session request may refer to at least one session request, and the session requests processed by the inference service node set may be a session request set, for example. The session request does not specifically refer to a certain fixed request. For example, when the number of sessions corresponding to the session request changes, the session request may also change accordingly.
[0020] In some embodiments, the load status information is used to indicate the load information borne by any service node. The load status information does not specifically refer to a certain fixed information. For example, when the service node changes, the load status information may also change accordingly. For example, when the node identifier of any service node changes, the load status information may also change accordingly.
[0021] In some embodiments, the load status information of any inference service node in the inference service node set may be obtained during the process of the inference service node set processing a session request.
[0022] Step 102, when the load status information indicates that the node pressure of any inference service node is greater than the pressure threshold, create a proxy service corresponding to any inference service node; In some embodiments, the node pressure can be, for example, the pressure of any service node. This node pressure does not specifically refer to a fixed pressure. For example, when the number of session request processes corresponding to any service node changes, the node pressure can also change accordingly. For example, when the resources required for any service node to process a session request change, the node pressure can also change accordingly.
[0023] In some embodiments, the pressure threshold can be, for example, a threshold used to determine whether to create a proxy service for any inference service node. This pressure threshold does not specifically refer to a fixed threshold. For example, when a modification instruction for the pressure threshold is received, the pressure threshold can also change accordingly. For example, when the business requirement information changes, the pressure threshold can also change accordingly.
[0024] According to some embodiments, the proxy service can be, for example, a proxy service corresponding to any inference service node. This proxy service does not directly participate in business calculations and is only responsible for forwarding requests. Among them, different inference service nodes can correspond to different proxy services.
[0025] According to some embodiments, when the load status information indicates that the node pressure of any inference service node is greater than the pressure threshold, a proxy service corresponding to the inference service node is created. That is, a proxy service corresponding to the inference service node with a node pressure greater than the pressure threshold can be created.
[0026] Step 103: Create a replica set of the proxy service according to the node pressure; In some embodiments, the replica set can be, for example, a collective formed by at least one replica. The replica set can include, for example, multiple replicas corresponding to the proxy service. This replica set does not specifically refer to a fixed set. For example, when the node pressure changes, the number of replicas corresponding to the replica set can also change accordingly.
[0027] In some embodiments, a replica set of the proxy service can be created according to the node pressure.
[0028] Step 104: Distribute the session requests corresponding to any inference service node to each replica in the replica set for processing, and perform a load balancing operation between the proxy service and each replica; According to some embodiments, the request allocation policy can be, for example, a policy of allocating the session requests corresponding to any inference service node to the replica set. This request allocation policy does not specifically refer to a certain fixed policy. For example, when the traffic of any inference service node changes, the request allocation policy can also change accordingly. For example, when the replica set changes, the request allocation policy can also change accordingly. Among them, the session requests corresponding to any inference service node can be, for example, at least one.
[0029] Among some embodiments, performing a load balancing operation between the proxy service and each replica can be used to achieve balanced traffic distribution and reduce the situation where the request processing efficiency is low due to unbalanced traffic between replicas.
[0030] Among some embodiments, according to the request allocation policy, the session requests corresponding to any inference service node are distributed to each replica in the replica set, and a load balancing operation is performed between the proxy service and each replica.
[0031] Step 105, when the pressure information corresponding to the replica set meets the pressure requirement, determine the first replica in the replica set, release the replicas in the replica set except the first replica, and use the first replica to replace the proxy service.
[0032] After some embodiments, the pressure requirement can be, for example, a requirement for determining whether to perform a downgrade process on the replica set. This pressure requirement does not specifically refer to a certain fixed requirement. For example, the pressure requirement can be that the pressure information is less than a certain pressure threshold. For example, when the pressure threshold changes, the pressure requirement can also change accordingly.
[0033] Among some embodiments, the first replica can be, for example, the replica left in the replica set, which can replace the proxy service and be the replica of any inference service node. The "first" in the first replica is used to distinguish it from the other replicas and does not specifically refer to a certain fixed replica. Among them, other replicas in the replica set can be released, for example.
[0034] According to some embodiments, using the first replica to replace the proxy service can be, for example, the proxy service going offline, and the first replica becoming the primary replica node to achieve single-point service.
[0035] Among some embodiments, when the pressure information corresponding to the replica set meets the pressure requirement, perform a downgrade process on the replica set, retain the first replica, and use the first replica to replace the proxy service.
[0036] Through the present disclosure, during the process of processing session requests by an inference service node set, the load status information of any inference service node in the inference service node set is obtained; in the case where the load status information indicates that the node pressure of any inference service node is greater than a pressure threshold, a proxy service corresponding to any inference service node is created; according to the node pressure, a replica set of the proxy service is created; according to a request distribution policy, the session requests corresponding to any inference service node are distributed to each replica in the replica set for processing, and a load balancing operation is performed between the proxy service and each replica; in the case where the pressure information corresponding to the replica set meets the pressure requirements, a first replica is determined in the replica set, replicas other than the first replica in the replica set are released, and the first replica is used to replace the proxy service. Therefore, when the node pressure of a certain node is relatively large, load balancing can be performed through the proxy service, the situation where the service is interrupted due to the inability of the node to process session requests can be reduced, the resource bottleneck and service congestion can be reduced, load balancing operations can be performed among multiple replicas, the overall throughput capacity can be improved, the request processing efficiency can be improved, and when the pressure decreases, degradation processing can be performed, only one replica is retained, resources that are not needed can be released, elastic recovery and resource utilization efficiency can be achieved, and thus the inference efficiency when the node pressure is relatively large can be improved while the utilization efficiency of resources is improved.
[0037] It should be noted that there may be multiple steps in the embodiments of the present disclosure. For the convenience of description, these steps are numbered, but these labels are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.
[0038] Further, in a possible implementation manner of this embodiment, Figure 2 is a schematic flowchart of another load balancing method for an inference service provided by an embodiment of the present disclosure. As Figure 2 shown, this method includes the following steps: Step 201, during the process of processing session requests by an inference service node set, obtain the load status information of any inference service node in the inference service node set; Among them, related descriptions can be as described above, and will not be elaborated here.
[0039] In some embodiments, the embodiments of the present disclosure can be used, for example, in a pre-training model inference service scenario with large-scale and dynamic load balancing. The pre-training model can be, for example, a model trained well on a large dataset through unsupervised or self-supervised learning. The pre-training model does not specifically refer to a certain fixed model. When the specific type of the pre-training model changes, the pre-training model can also change accordingly.
[0040] Step 202, when the node pressure of any inference service node is indicated to be greater than the pressure threshold by the load status information, create a proxy service corresponding to any inference service node; Among them, the relevant descriptions can be as described above, and will not be elaborated here.
[0041] According to some embodiments, the open sea method further includes: When starting the set of inference service nodes and each inference service node in the set of inference service nodes passes the consistent hashing load balancing check, route the same session request in the session request to the same inference service node. Therefore, the utilization rate of the kv cache can be improved, the inference speed can be increased, and the inference efficiency can be improved.
[0042] According to some embodiments, obtaining the load status information of any inference service node in the set of inference service nodes includes: Use the monitoring services corresponding to the inference service nodes in the set of inference service nodes to collect data from each inference service node, and obtain the index information corresponding to each inference service node, where the index information includes at least one of the central processing unit usage rate, memory information, graphics processing unit usage rate, and active session count; Perform aggregation processing and analysis processing on the index information to obtain the load status information corresponding to each inference service node. Therefore, the load status information can be determined according to the index information, the accuracy of obtaining the load status information can be improved, the accuracy of creating the proxy service can be improved, the situation where the inference speed is slow due to excessive node pressure can be reduced, and the inference efficiency can be improved. Among them, the usage rate can also be referred to as the utilization rate, and the embodiments of the present disclosure do not limit this.
[0043] In some embodiments, for example, monitoring services can be deployed on each inference service node, where there is no limitation on the monitoring service. There is no limitation on the data collected by the monitoring service. For example, at least one of the central processing unit (CPU) usage rate, memory information, graphics processing unit (GPU) usage rate, and active session count can be collected. The embodiments of the present disclosure do not limit this. The collected index information can be matched with the alarm rules, for example.
[0044] In some embodiments, for example, the data collected by the monitoring service can be obtained once every preset duration. The preset duration can be 5 seconds, for example. The specific example of obtaining the data can be the central monitoring module in the electronic device.
[0045] Among them, the index information can be subjected to aggregation processing and analysis processing to obtain the load status information corresponding to each inference service node. Among them, for example, the index information can also be visually processed and displayed.
[0046] Among them, according to the load status information, it is determined whether the node pressure of any inference service node is greater than the pressure threshold. Specifically, for example, each piece of information in the load status information is greater than the corresponding threshold. Among them, the judgment rule or the alarm rule can be determined based on the rule setting instruction, and the embodiments of the present disclosure do not limit this. Among them, the alarm information can be one of the metric information or multiple. For example, when the CPU usage rate is greater than 80%, the GPU usage rate is greater than 70%, and the number of active sessions is greater than 100, it is determined that the node pressure of any inference service node is greater than the pressure threshold.
[0047] According to some embodiments, the method further includes: Obtain the business requirement information corresponding to the session request; According to the business requirement information, obtain the pressure threshold and the request allocation strategy. Therefore, the matching degree between the pressure threshold and the business requirement information can be improved, the matching degree between the request allocation strategy and the business requirement information can be improved, the accuracy of obtaining the pressure threshold can be improved, and the accuracy of creating the proxy service can be improved.
[0048] According to some embodiments, the method further includes: Obtain the business requirement information corresponding to the session request; According to the business requirement information, obtain the pressure threshold. Therefore, the matching degree between the pressure threshold and the business requirement information can be improved, the accuracy of obtaining the pressure threshold can be improved, and the accuracy of creating the proxy service can be improved.
[0049] According to some embodiments, for example, when the resource occupancy rate of any inference service node exceeds the threshold or the number of active sessions exceeds the threshold, a proxy service corresponding to any inference service node can be created. That is, the expansion process can be triggered.
[0050] Among them, for example, when it is determined that an alarm message is received, the overloaded node can be determined, and a proxy service corresponding to the overloaded node is created, and the overloaded node is the node whose node pressure is greater than the pressure threshold.
[0051] Among them, for example, a proxy service can be deployed in front of any inference service node through an automated script or a container orchestration service. The Proxy proxy service can run as an independent container or process and has the ability of service registration and discovery. After the Proxy service is started, it can register itself with the service discovery server and mark the original node as its replica.
[0052] Step 203, create a replica set of the proxy service according to the node pressure; Among them, the relevant description can be as described above, and will not be repeated here.
[0053] According to some embodiments, after creating a replica set of the proxy service, for example, the replica set can be checked at preset intervals to determine that the running status of each replica meets the status requirements, improving the availability of the service.
[0054] In some embodiments, after creating the proxy service, the load balancer or Domain Name System (DNS) configuration can be updated to switch all session requests directed to any inference service node to the Proxy proxy service. Among them, for example, "hot switching" or "lossless migration" mechanisms can be adopted to reduce the situation of session interruption and improve the continuity of session request processing. Among them, the type of the load balancer is not limited.
[0055] According to some embodiments, obtaining a replica set of the proxy service according to node pressure includes: Obtaining the number of replicas corresponding to the proxy service according to node pressure and load information; According to the number of replicas, creating a replica set of the proxy service and adding any inference service node as a replica to the replica set.
[0056] Among them, the load information can include at least one of request rate, average response time, and replica utilization rate. Therefore, the number of replicas can be determined according to the load information and node pressure, reducing the waste of resources caused by excessive creation of replicas. Specifically, the replica set can be created by a container orchestration server. Each replica in the replica set is registered under the Proxy service and confirmed to be available after passing the status check. For example, the GPU resources and CPU resources corresponding to each replica can be determined.
[0057] According to some embodiments, the method further includes: Obtaining the business requirement information corresponding to the session request; Obtaining a request allocation policy according to the business requirement information. Among them, the request allocation policy can be executed before step 203. Therefore, the policy can be dynamically switched, which can adapt to different business load scenarios, improve the accuracy of obtaining the request allocation policy, improve the matching degree between the request allocation policy and the traffic of any inference service node, and improve the request processing efficiency.
[0058] According to some embodiments, the external service port of any inference service node is taken over by the proxy service, and any inference service node is used as one of the replicas of the proxy service.
[0059] Among them, the request allocation policy can include various types, for example, it can include round robin, least connections, least load, etc., which can achieve uniform traffic sharing.
[0060] Step 204: According to the request distribution policy, distribute the session requests corresponding to any inference service node to each replica in the replica set for processing, and perform load balancing operations between the proxy service and each replica; Among them, the relevant descriptions can be as described above and will not be elaborated here.
[0061] According to some embodiments, distributing the session requests of any inference service node to each replica in the replica set according to the request distribution policy includes: Switch the session requests of any inference service node to the proxy service; Control the proxy service to distribute the session requests to each replica in the replica set according to at least one of the request distribution policy and the session identifier.
[0062] For example, Figure 3 is an example schematic diagram of a load balancing method provided by an embodiment of the present disclosure. As Figure 3 shown, the Proxy proxy service can uniformly receive and forward all session requests, and distribute them to the replica set according to the session identifier id or the load balancing policy, which can improve session stickiness and reduce the situation where the same session is processed alternately by different replicas. Among them, the inference service nodes can include, for example, the pre-trained model service A, the pre-trained model service B, and the pre-trained model service C. Among them, the node pressure of the pre-trained model service B is greater than the pressure threshold, and a proxy service corresponding to the pre-trained model service B can be created. Traffic distribution can be performed by means of round-robin or the least connection number load balancing method. Among them, replicas B1, B2, and B3 are the replica set of the pre-trained model service B. Key-value cache synchronization can be performed between replicas B1, B2, and B3. Among them, Client is the client and User is the user.
[0063] According to some embodiments, the method further includes: Obtain at least one uncompleted session request during the process of switching the traffic of any inference service node to the proxy service; In the case of confirming that the traffic of any inference service node is switched to the proxy service, distribute at least one session request to each replica. Therefore, the uncompleted requests in the cache or queue can be continuously processed after the switch is completed, making the user requests unaware and improving the user experience.
[0064] According to some embodiments, each replica can be detected every preset duration, and when the detection fails, it can be removed and recreated to improve the high availability of each replica.
[0065] Step 205, when each replica has processed the target session request and there is a change in the key-value cache, synchronize the key-value cache among the replicas through a serialization and broadcasting mechanism. According to some embodiments, the key-value cache can be used, for example, to refer to the Key and Value cached during the autoregressive inference process, which can reduce repeated calculations. Among them, the change in the key-value cache can be caused, for example, by adding new tokens or expanding the execution context. Among them, the embodiments of the present disclosure do not limit the change method of the key-value cache.
[0066] In some embodiments, serialization can be, for example, performing a serialization operation on the changed key-value cache. Broadcasting can be, for example, an operation of broadcasting the changed key-value cache among the replicas. Among them, the method of broadcasting is not limited. This broadcasting method can include, for example, a full-volume broadcasting method and an incremental broadcasting method.
[0067] In some embodiments, when each replica has processed the target session request and there is a change in the key-value cache, synchronize the key-value cache among the replicas through a serialization and broadcasting mechanism.
[0068] According to some embodiments, when each replica has processed the target session request and there is a change in the key-value cache, synchronizing the key-value cache among the replicas through a serialization and broadcasting mechanism includes: When each replica has processed the target session request and there is a change in the key-value cache, upload the serialized key-value cache to the centralized storage system; Send synchronization notification information to a subset of replicas through a network broadcasting method, where the synchronization notification information includes information related to the change in the key-value cache; When it is determined that each replica has received the synchronization notification information, according to the session identifier and the key-value cache version number in the synchronization notification information, pull the key-value cache closest to the current time from the centralized storage system and perform deserialization to obtain the deserialized key-value cache; When the deserialized key-value cache passes the verification, store the deserialized key-value cache and delete the stored key-value cache. Therefore, kv cache synchronization between replicas can be achieved based on a high-speed network, improving the synchronization efficiency of the key-value cache and the request processing efficiency.
[0069] According to some embodiments, the kv cache can be, for example, tensor or dictionary (dict) structure data, which can be serialized into binary data. Among them, it can be serialized into a memory object or directly written to a disk file, which can improve the convenience of subsequent storage and transmission of the kv cache. The dictionary structure is a data structure that stores data in the form of key-value pairs.
[0070] In some embodiments, the centralized storage system may be, for example, at least one of a Remote Dictionary Server (Redis), a distributed key-value (KV), and an object storage.
[0071] According to some embodiments, uploading the serialized key-value cache to the centralized storage system includes: Performing a serialization operation on the changed key-value cache using the serialization interface of the inference framework to obtain the key-value corresponding to the serialized key-value cache; Uploading the serialized key-value cache to the centralized storage system according to the key-value. This can improve the accuracy of storing the serialized key-value cache, improve the subsequent retrieval efficiency, and improve the convenience of subsequent use.
[0072] According to some embodiments, sending a synchronization notification message to all replicas via network broadcast includes: Sending a synchronization notification message to all replicas via a message queue or a distributed bus, where the synchronization notification message includes at least one of a session identifier, a key-value cache version number, a change timestamp information, and a central storage location information, and the central storage location information is used to indicate the storage location of the serialized key-value cache in the centralized storage system.
[0073] In some embodiments, for example, each replica can be controlled to subscribe to the kv cache change topic (topic), that is, the "publish-subscribe" mode can be adopted to reduce the synchronization duration when the kv cache changes. Each replica process continuously listens to the message queue / event bus and waits for the kv cache change event. Once a kv cache change notification related to the local session is received, it immediately enters the synchronization process.
[0074] According to some embodiments, when pulling the latest kv cache binary data from the central storage, for example, idempotent pulling and resume from breakpoint can be supported, which can reduce the probability of repeated or partial pulling of the kv cache.
[0075] In some embodiments, pulling the key-value cache closest to the current time from the centralized storage system and performing deserialization to obtain the deserialized key-value cache. For example, the key-value cache closest to the current time can be pulled from the centralized storage system and deserialized into a local Tensor or dict structure using the framework interface to obtain the deserialized key-value cache.
[0076] According to some embodiments, the deserialized key-value cache verification can be, for example, verifying metadata information such as the session_id, number of tokens, version number, etc. of the kv cache, which can improve the accuracy of data acquisition, so that the acquired data is the most recent data closest to the current time and has not been concurrently overwritten.
[0077] Among some embodiments, if there is a higher version of the kv cache locally, discard the data of this synchronization. Or in the case of data anomalies or verification failures, record the log and wait for the next synchronization.
[0078] According to some embodiments, when the deserialized key-value cache verification passes, the deserialized key-value cache can be stored and the stored key-value cache can be deleted, so that the context is consistent when processing this session next time. For example, atomic replacement can also be performed on the deserialized key-value cache to reduce the situation of data inconsistency caused by concurrent reading and writing and improve the accuracy of data storage.
[0079] Among some embodiments, when storing the deserialized key-value cache, for example, only the changed part can be synchronized, which can reduce bandwidth and latency, improve the accuracy of data storage and at the same time improve the data storage efficiency. Among them, the changed part to be synchronized can be, for example, a newly added token. In the case of failure to synchronize the changed part, it can be retried again, or degraded to a full synchronization. Among them, the failure to synchronize the changed part can include, for example, pull timeout or deserialization failure. For example, the synchronization status and failure logs can also be recorded for subsequent monitoring and alerting.
[0080] According to some embodiments, the synchronization notification information is sent to the replica subset by means of network broadcast, including: Obtain the network broadcast method corresponding to the business requirement information, where the network broadcast method includes one of the full broadcast method and the incremental broadcast method; Send the synchronization notification information to the replica subset corresponding to the network broadcast method. Improve the accuracy of sending the synchronization notification information and reduce the waste of resources.
[0081] Among them, the full broadcast method can be, for example, notifying all replicas, and the incremental broadcast method can be, for example, notifying the relevant replicas in the replica set, that is, the department replicas. Among them, the relevant replicas can be, for example, the replicas corresponding to the same session. Message retry and deduplication can be supported in the network broadcast method, which can improve availability and consistency.
[0082] According to some embodiments, when the network broadcast method is the incremental broadcast method, the method further includes: At each preset time interval, perform a full verification on the key-value caches of each replica set to obtain the key-value cache verification result, where the key-value cache verification result is used to indicate the consistency between the key-value caches of each replica set. Therefore, full verification can be performed regularly to prevent data drift caused by long-term incremental synchronization and improve the consistency of key-value cache storage.
[0083] In some embodiments, for example, multi-replica redundancy processing and primary-backup switching processing can be performed, which can reduce the impact on the overall kv cache consistency in case of a single-point failure.
[0084] In some embodiments, Figure 4 is an example schematic diagram of a key-value cache synchronization method provided by an embodiment of the present disclosure. As Figure 4 shown, for example, it can include that replica A finishes processing a session request, and it can be determined whether the key-value cache has changed. In the case of no change, it can wait for the next request. In the case of a change, the key-value cache can be serialized, stored in the central storage system, and a change notice can be broadcast. When other replicas listen to the notice, other replicas can pull the latest key-value cache, that is, the key-value cache closest to the current time, deserialize the key-value cache, and perform a consistency check. When the consistency check passes, the local key-value cache can be replaced. During this process, for example, incremental synchronization and / or failure retry operations can be performed. In the case where the consistency check fails, the key-value cache can be discarded.
[0085] Step 206, when the pressure information corresponding to the replica set meets the pressure requirement and lasts for a preset time interval, select the first replica in the replica set as the primary replica; Among them, the related descriptions can be as described above and will not be elaborated here.
[0086] According to some embodiments, selecting the first replica in the replica set as the primary replica includes: Select the first replica in the replica set as the primary replica by using at least one of the selection methods of priority weight and polling.
[0087] According to some embodiments, when distributing the traffic of any inference service node to each replica in the replica set, for example, through the monitoring system integrated in the Proxy server or the docked monitoring system, continuously collect indicators such as the CPU memory, GPU usage rate, active session number, and request rate of each replica in all replicas. Among them, for example, the monitoring can be performed through the monitoring scheme in some embodiments, or through a pre-set lightweight monitoring module.
[0088] In some embodiments, the pressure requirement can be determined according to the service requirement information. The pressure requirement can include, for example, that the CPU memory occupancy rate is less than 30% and the number of sessions is less than 10. The preset duration can be, for example, 5 minutes, which reduces the situation of frequent scaling down caused by occasional fluctuations and improves the accuracy of scaling up and down control.
[0089] According to some embodiments, Figure 5 is an example schematic diagram of a proxy service removal method provided by an embodiment of the present disclosure. As Figure 5 shown, when the pressure information corresponding to the replica set meets the pressure requirement and lasts for the preset duration, the scaling-down process can be triggered, and the trigger event and related monitoring data can be recorded for subsequent traceability and parameter optimization. Among them, when it is determined that the pressure corresponding to the replica set is less than the pressure threshold, the proxy service can be removed.
[0090] In some embodiments, for example, based on strategies such as the inspection information, current load information, and historical stability of each replica, a replica that meets the running requirements and has a lower load can be elected as the primary service node. Among them, the election can be carried out, for example, by means of priority weights or polling.
[0091] Step 207, synchronize the key-value caches of all session requests to the first replica to obtain a second replica; The relevant descriptions can be, for example, as described above and will not be elaborated here.
[0092] In some embodiments, specifically, the kv cache of all sessions can be merged / synchronized to the primary replica through a central storage or a message queue, and the kv cache of the sessions of the primary replica can be detected by means of incremental synchronization detection or conflict detection, so as to improve the consistency of data acquisition of the second replica, that is, the primary replica.
[0093] In some embodiments, data detection can be performed on the second replica to determine whether the kv cache of the second replica includes the key-value caches of all session requests, reduce the probability of data loss, and improve the integrity of the provided service.
[0094] Step 208, when the second replica passes the verification, use the second replica as the primary service node, and release at least one third replica in the replica set according to the replica release order, where at least one third replica is the remaining replicas in the replica set except the first replica; The relevant descriptions can be, for example, as described above and will not be elaborated here.
[0095] According to some embodiments, releasing at least one third replica in the replica set includes: Release at least one third replica in ascending order of the load information of at least one third replica in the replica set; Release at least one third replica in descending order of the offline priority of non-primary replicas. Therefore, the resources occupied by at least one third replica can be released, improving resource utilization.
[0096] According to some embodiments, before releasing at least one third replica, for example, it is also possible to confirm whether the requests corresponding to at least one third replica are completed, and transfer the uncompleted requests to other nodes for processing. After releasing at least one third replica, the available resource pool of the resource management system can be monitored.
[0097] Step 209, adjust the routing of the proxy service and remove the proxy service.
[0098] Among them, the related descriptions can be as described above and will not be elaborated here.
[0099] In some embodiments, when it is determined that all at least one third replicas are released and the running state of the second replica meets the state requirements, the removal of the proxy service can be triggered.
[0100] According to some embodiments, adjusting the routing of the proxy service includes at least one of the following: Control the proxy service to retain the routing of the first replica and remove the forwarding configuration of at least one third replica; Point the ingress load balancing configuration to the first replica.
[0101] According to some embodiments, for example, the ingress load balancing configuration can be switched from Proxy to directly point to the primary replica to achieve seamless traffic migration. Among them, during this process, for example, hot swapping can be supported, which can reduce the situation of session interruption.
[0102] In some embodiments, after releasing the proxy service, the resources occupied by the proxy service can be released and the registration in the discovery system can be updated, which can improve the convenience for external requests to access the primary replica.
[0103] In some embodiments, the electronic device can continuously monitor the pressure of each inference service node, and re-trigger the scaling process when the pressure is greater than the pressure threshold, enabling elastic scaling closed-loop control and improving the utilization efficiency of resources.
[0104] In some embodiments, creating a proxy service corresponding to any inference service node can decouple the business traffic from the nodes with relatively high pressure, and can improve the flexibility of scaling and traffic migration. Secondly, kv cache synchronization can be performed, which can improve the integrity of the context corresponding to the primary replica and improve data consistency. Then, when the pressure of the replica set is less than the pressure threshold, degradation processing can be performed, which can reduce resource occupancy, improve resource utilization efficiency, and can achieve an organic combination of high availability, high performance, dynamic scalability, and context consistency, thereby improving the inference efficiency.
[0105] According to an embodiment of the present disclosure, the present disclosure also provides a load balancing device for an inference service.
[0106] Exemplarily, Figure 6 FIG. is a schematic structural diagram of a load balancing device for an inference service provided by an embodiment of the present disclosure. The load balancing device 600 for the inference service includes: an information acquisition set 601, a service creation unit 602, a set creation unit 603, a request distribution unit 604, and a proxy release unit 605; wherein, The information acquisition set 601 is configured to obtain the load status information of any inference service node in the inference service node set during the process of the inference service node set processing a session request; The service creation unit 602 is configured to create a proxy service corresponding to any inference service node when the load status information indicates that the node pressure of any inference service node is greater than the pressure threshold; The set creation unit 603 is configured to create a replica set of the proxy service according to the node pressure; The request distribution unit 604 is configured to distribute the session request corresponding to any inference service node to each replica in the replica set for processing according to a request distribution policy, and perform a load balancing operation between the proxy service and each replica; The proxy release unit 605 is configured to determine a first replica in the replica set, release replicas other than the first replica in the replica set, and replace the proxy service with the first replica when the pressure information corresponding to the replica set meets the pressure requirement.
[0107] Further, when the information acquisition set 601 is configured to obtain the load status information of any inference service node in the inference service node set, it is specifically configured to: Collect data of each inference service node by using the monitoring service corresponding to each inference service node in the inference service node set, and obtain the index information corresponding to each inference service node, where the index information includes at least one of the central processing unit usage rate, memory information, graphics processing unit usage rate, and active session number; Aggregate and analyze the metric information to obtain the load status information corresponding to each inference service node.
[0108] Further, the information acquisition set 601 is also used for: Obtain the business requirement information corresponding to the session request; According to the business requirement information, obtain the pressure threshold and the request distribution policy.
[0109] Further, when the set creation unit 603 is used to create a replica set of the proxy service according to the node pressure, it is specifically used for: Obtain the number of replicas corresponding to the proxy service according to the node pressure and the load information; According to the number of replicas, create a replica set of the proxy service, and add any inference service node as a replica to the replica set.
[0110] Further, when the request distribution unit 604 is used to distribute the session requests of any inference service node to each replica in the replica set, it is specifically used for: Switch the session request of any inference service node to the proxy service; Control the proxy service to distribute the session request to each replica in the replica set according to at least one of the request distribution policy and the session identifier.
[0111] Further, the request distribution unit 604, the method is also specifically used for: Obtain at least one unfinished session request during the process of switching the session request of any inference service node to the proxy service; When it is confirmed that the session request of any inference service node is switched to the proxy service, distribute at least one session request to each replica.
[0112] Further, when the proxy release unit 605 is used to determine the first replica in the replica set, release the replicas other than the first replica in the replica set, and replace the proxy service with the first replica when the pressure information corresponding to the replica set meets the pressure requirement, it is specifically used for: When the pressure information corresponding to the replica set meets the pressure requirement and lasts for a preset duration, select the first replica in the replica set as the primary replica; Synchronize the key-value caches of all session requests to the first replica to obtain a second replica; When the second replica passes the verification, use the second replica as the primary service node, and release at least one third replica in the replica set according to the replica release order, where at least one third replica is the remaining replicas in the replica set other than the first replica; Adjust the routing of the proxy service and remove the proxy service.
[0113] Further, when the proxy release unit 605 is used to select the first replica in the replica set as the primary replica, it is specifically used for: Select the first replica in the replica set as the primary replica by using at least one of the selection methods of priority weight and polling.
[0114] Further, when the proxy release unit 605 is used to adjust the routing of the proxy service, it is specifically used for at least one of the following: Control the proxy service to retain the routing of the first replica and remove the forwarding configuration of at least one third replica; Point the ingress load balancing configuration to the first replica.
[0115] Further, when the proxy release unit 605 is used to release at least one third replica in the replica set, it is specifically used for at least one of the following: Release at least one third replica in ascending order of the load information of at least one third replica in the replica set; Release at least one third replica in descending order of the offline priority of non-primary replicas.
[0116] Further, the proxy release unit 605 is further used for: When each replica has processed the target session request and there is a change in the key-value cache, synchronize the key-value caches among the replicas through the serialization and broadcast mechanism.
[0117] Further, when the proxy release unit 605 is used to synchronize the key-value caches among the replicas through the serialization and broadcast mechanism when each replica has processed the target session request and there is a change in the key-value cache, it is specifically used for: When each replica has processed the target session request and there is a change in the key-value cache, upload the serialized key-value cache to the centralized storage system; Send a synchronization notification message to the replica subset through network broadcast, where the synchronization notification message includes relevant information about the change in the key-value cache; When it is determined that each replica has received the synchronization notification message, pull the key-value cache closest to the current time from the centralized storage system according to the session identifier and the key-value cache version number in the synchronization notification message and perform deserialization to obtain the deserialized key-value cache; When the deserialized key-value cache passes the verification, store the deserialized key-value cache and delete the stored key-value cache.
[0118] Further, when the proxy release unit 605 is used to upload the serialized key-value cache to the centralized storage system, it is specifically used for: Perform serialization operations on the changed key-value cache using the serialization interface of the inference framework to obtain the key values corresponding to the serialized key-value cache. Upload the serialized key-value cache to the centralized storage system according to the key values.
[0119] Further, the proxy release unit 605, when used to send synchronization notification information to all replicas through network broadcasting, is specifically used for: Send synchronization notification information to all replicas through a message queue or a distributed bus, where the synchronization notification information includes at least one of a session identifier, a key-value cache version number, a change timestamp information, and a central storage location information, and the central storage location information is used to indicate the storage location of the serialized key-value cache in the centralized storage system.
[0120] Further, the proxy release unit 605, when used to send synchronization notification information to a subset of replicas through network broadcasting, is specifically used for: Obtain the network broadcasting method corresponding to the service requirement information, where the network broadcasting method includes one of a full-scale broadcasting method and an incremental broadcasting method; Send synchronization notification information to the subset of replicas corresponding to the network broadcasting method.
[0121] Further, when the network broadcasting method is an incremental broadcasting method, the proxy release unit 605, the method is further specifically used for: At preset time intervals, perform a full-scale verification on the key-value caches of each replica set to obtain the key-value cache verification result, where the key-value cache verification result is used to indicate the consistency between the key-value caches of each replica set.
[0122] It should be noted that the descriptions of the features in the embodiments corresponding to the load balancing device of the inference service can refer to the relevant descriptions of the embodiments corresponding to the load balancing method of the inference service, and will not be elaborated here one by one.
[0123] An embodiment of the present disclosure also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the load balancing method of the inference service.
[0124] An embodiment of the present disclosure also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the load balancing method of the inference service when running.
[0125] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs, such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), external hard drives, magnetic disks, or optical discs.
[0126] Embodiments of the present disclosure also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the load balancing method for inference services.
[0127] Embodiments of the present disclosure also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the load balancing method for inference services.
[0128] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0129] The above has introduced in detail a load balancing method for inference services provided by the present disclosure. Specific examples are used herein to elaborate on the principles and implementation manners of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present disclosure, several improvements and modifications can be made to the present disclosure, and these improvements and modifications also fall within the protection scope of the claims of the present disclosure.
Claims
1. A load balancing method for inference services, characterized in that, Including: During the process of processing a session request by a set of inference service nodes, obtaining the load status information of any one of the inference service nodes in the set of inference service nodes; When the load status information indicates that the node pressure of the any one inference service node is greater than a pressure threshold, creating a proxy service corresponding to the any one inference service node; Creating a replica set of the proxy service according to the node pressure; According to a request distribution policy, distributing the session request corresponding to the any one inference service node to each replica in the replica set for processing, and performing a load balancing operation between the proxy service and each replica; When the pressure information corresponding to the replica set meets the pressure requirement, determining a first replica in the replica set, releasing replicas in the replica set except the first replica, and replacing the proxy service with the first replica.
2. The method according to claim 1, wherein The obtaining the load status information of any one inference service node in the set of inference service nodes includes: Using the monitoring services corresponding to the respective inference service nodes in the set of inference service nodes to collect data from the respective inference service nodes, and obtaining the metric information corresponding to the respective inference service nodes, where the metric information includes at least one of the central processing unit usage rate, memory information, graphics processing unit usage rate, and number of active sessions; Performing an aggregation process and an analysis process on the metric information to obtain the load status information corresponding to the respective inference service nodes.
3. The method according to claim 1, wherein The method further includes: Obtaining the service requirement information corresponding to the session request; According to the service requirement information, obtaining the pressure threshold and the request distribution policy.
4. The method according to claim 1, characterized in that, The creating a replica set of the proxy service according to the node pressure includes: Obtaining the number of replicas corresponding to the proxy service according to the node pressure and the load information; According to the number of replicas, creating a replica set of the proxy service, and adding the any one inference service node as a replica to the replica set.
5. The method according to claim 1, characterized in that, The distributing the session request of the any one inference service node to each replica in the replica set according to the request distribution policy includes: Switching the session request of the any one inference service node to the proxy service; Controlling the proxy service to distribute the session request to each replica in the replica set according to at least one of the request distribution policy and the session identifier.
6. The method according to claim 1, wherein The method further includes: Obtaining at least one uncompleted session request during the process of switching the session request of the any one inference service node to the proxy service; When it is confirmed that the session request of the any one inference service node is switched to the proxy service, distributing the at least one session request to each replica.
7. The method according to claim 1, characterized in that The when the pressure information corresponding to the replica set meets the pressure requirement, determining a first replica in the replica set, releasing replicas in the replica set except the first replica, and replacing the proxy service with the first replica includes: When the pressure information corresponding to the replica set meets the pressure requirement and lasts for a preset duration, selecting the first replica in the replica set as the primary replica; Synchronize the key - value caches of all session requests to the first replica to obtain a second replica; When the second replica passes the verification, use the second replica as the primary service node, and release at least one third replica in the replica set according to the replica release order, where the at least one third replica is the remaining replicas in the replica set except the first replica; Adjust the routing of the proxy service and remove the proxy service.
8. The method according to claim 7, wherein The selecting the first replica in the replica set as the primary replica includes: Select the first replica in the replica set as the primary replica by using at least one of priority weight and round - robin selection methods.
9. The method according to claim 7, wherein The adjusting the routing of the proxy service includes at least one of the following: Control the proxy service to retain the routing of the first replica and remove the forwarding configuration of the at least one third replica; Point the ingress load - balancing configuration to the first replica.
10. The method according to claim 7, characterized in that, The releasing at least one third replica in the replica set includes at least one of the following: Release the at least one third replica in ascending order of the load information of the at least one third replica in the replica set; Release the at least one third replica in descending order of the non - primary replica offline priority.
11. The method according to claim 1, characterized in that, The method further includes: After each replica processes the target session request and there is a change in the key - value cache, synchronize the key - value caches between the replicas through a serialization and broadcast mechanism.
12. The method according to claim 11, wherein The synchronizing the key - value caches between the replicas through a serialization and broadcast mechanism after each replica processes the target session request and there is a change in the key - value cache includes: After each replica processes the target session request and there is a change in the key - value cache, upload the serialized key - value cache to a centralized storage system; Send a synchronization notification message to a subset of replicas through network broadcasting, where the synchronization notification message includes information related to the change in the key - value cache; When it is determined that each replica has received the synchronization notification message, pull the key - value cache closest to the current time from the centralized storage system according to the session identifier and the key - value cache version number in the synchronization notification message and perform deserialization to obtain the deserialized key - value cache; When the deserialized key - value cache passes the verification, store the deserialized key - value cache and delete the stored key - value cache.
13. The method according to claim 12, characterized in that, The uploading the serialized key - value cache to a centralized storage system includes: Use the serialization interface of the inference framework to perform serialization operations on the changed key - value cache to obtain the key - value corresponding to the serialized key - value cache; Upload the serialized key - value cache to the centralized storage system according to the key - value.
14. The method according to claim 12, wherein The sending the synchronization notification message to all replicas through network broadcasting includes: Send synchronization notification information to all replicas through a message queue or a distributed bus, where the synchronization notification information includes at least one of a session identifier, a key-value cache version number, a change timestamp information, and a central storage location information, and the central storage location information is used to indicate the storage location of the serialized key-value cache in the centralized storage system.
15. The method according to claim 12, characterized in that, The sending of the synchronization notification information to a subset of replicas through a network broadcast manner includes: Obtain the network broadcast manner corresponding to the service requirement information, where the network broadcast manner includes one of a full-scale broadcast manner and an incremental broadcast manner; Send the synchronization notification information to the subset of replicas corresponding to the network broadcast manner.
16. The method according to claim 15, wherein Wherein, When the network broadcast manner is an incremental broadcast manner, the method further includes: At preset time intervals, perform a full-scale verification on the key-value caches of the respective replica sets, and obtain a key-value cache verification result, where the key-value cache verification result is used to indicate the consistency between the key-value caches of the respective replica sets.
17. A load balancing device for an inference service, characterized in that, Including: An information acquisition set for acquiring the load status information of any one of the inference service nodes in the inference service node set during the process of the inference service node set processing a session request; A service creation unit for creating a proxy service corresponding to the any one of the inference service nodes when the load status information indicates that the node pressure of the any one of the inference service nodes is greater than a pressure threshold; A set creation unit for creating a replica set of the proxy service according to the node pressure; A request distribution unit for distributing the session request corresponding to the any one of the inference service nodes to each replica in the replica set for processing according to a request distribution policy, and performing a load balancing operation between the proxy service and the respective replicas; A proxy release unit for determining a first replica in the replica set, releasing the replicas other than the first replica in the replica set, and replacing the proxy service with the first replica when the pressure information corresponding to the replica set meets the pressure requirement.
18. An electronic device, characterized in that It includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 16.
19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 16.
20. A computer program product, characterized in that, Including a computer program, which realizes the method according to any one of claims 1 to 16 when executed by a processor.
Citation Information
Patent Citations
Model cache scheduling method and system under edge AI reasoning scene
CN118113442A
Distributed key value storage architecture and storage method and system based on architecture
CN118733594A
Inference service management method, equipment, medium and computer program product
CN119415273A
Serverless large model reasoning service system, method, equipment and medium
CN119440739A
Inference service monitoring method and device, computer equipment and storage medium
CN119759694A
Cited By
Model reasoning scheduling system
CN121078050A
Dynamic reasoning load balancing service method for doctor-patient dialogue model
CN121148746A
Inference service copy pool management method, electronic equipment and storage medium
CN122287909A