Service matching method, electronic equipment, program product and storage medium
By receiving inference requests in the container orchestration platform and matching the target container group according to the request length and hardware resource amount, the problem of poor load balancing in the model inference service is solved, and better load balancing and service quality are achieved.
Patent Information
- Application Number
- CN202510408497.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the container orchestration platform provides model inference services, the load balancing effect is poor, resulting in load imbalance in the container group and affecting service quality.
By receiving inference requests, determining the request length, and matching the target container group for inference requests based on the request length and the amount of hardware resources of the container group, ensuring that the hardware resources can match the requested resource consumption, thereby achieving better load balancing.
By dynamically matching container groups, the load balancing effect of model inference services is improved according to the request length and hardware resource amount, the load imbalance of container groups is avoided, and the service quality is improved.
Smart Images

Figure CN119917293A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of service load balancing, and in particular to a service matching method, electronic device, program product and storage medium. Background Art
[0002] With the continuous development of artificial intelligence technology, pre-trained language models have been widely used. To support the application deployment of pre-trained language models, the container orchestration platform can be used to deploy pre-trained language models in multiple container groups (Pods), and these container groups jointly provide model inference services to the outside world. However, the load balancing effect provided by the container orchestration platform for the container group that provides model inference services is not good. When dealing with a large number of inference requests, it is easy to cause load imbalance in the container group, affecting the service quality. Summary of the invention
[0003] The purpose of this application is to provide a service matching method, device, program product and storage medium, which can match a suitable target container group for an inference request based on the request length and the amount of hardware resources corresponding to the container group providing the model inference service, and can consider the performance loss of the model inference service due to the request length, thereby achieving a better load balancing effect.
[0004] In order to solve the above technical problems, the present application provides a service matching method, including: receiving an inference request sent to a model inference service; Determine the request length of the inference request, and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; the model inference service is provided by at least two container groups; The inference request is sent to the target container group, so that the target container group provides model inference service according to the inference request.
[0005] Optionally, matching a target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service includes: Determine whether the request length is greater than the preset length; If the request length is greater than the preset length, determining the predicted resource consumption corresponding to the request length; According to the remaining amount of hardware resources corresponding to the container group, the target container group is matched among the container groups that meet the predicted resource consumption.
[0006] Optionally, matching the target container group in the container group that meets the predicted resource consumption includes: The container group that meets the predicted resource consumption is used as a candidate container group; According to the remaining amount of hardware resources of each candidate container group, a selection weight is set for each candidate container group; the selection weight is positively correlated with the remaining amount of hardware resources; A target container group is matched for an inference request from among candidate container groups based on the selection weight.
[0007] Optionally, it also includes: According to the conversation identification information in the reasoning request, find the historical reasoning request that belongs to the same conversation as the reasoning request; Determine the predicted resource consumption corresponding to the request length, including: The predicted resource consumption is determined according to the request length of the inference request and the request length of historical inference requests.
[0008] Optionally, after determining whether the request length is greater than a preset length, the method further includes: If the request length is not greater than the preset length, the container group corresponding to the client to which the inference request belongs is used as the target container group, or the container group with the minimum number of request processing is used as the target container group based on the current number of request processing of the container group.
[0009] Optionally, determine the request length of the inference request, including: Read the request length from the content-length field of the inference request; Alternatively, a message length of a request message including the inference request is determined, and the message length is used as the request length.
[0010] Optionally, the model inference service includes a service component, and the service component includes an external service component, an internal service component, and a request buffer; Receive an inference request sent to the model inference service, including: Control external service components to receive reasoning requests; Based on the overall performance value of the container group, determine whether the container group is in an overall slow running state; If the container group is in an overall slow running state, the external service component is controlled to send the inference request to the request buffer, and the internal service component is controlled to obtain the inference request from the request buffer; Sending the inference request to the target container group includes: The inference request in the external service component or the inference request in the internal service component is sent to the target container group.
[0011] Optionally, the model inference service also includes a service mesh; Determine the request length of the inference request, and match the target container group for the inference request based on the request length and the amount of hardware resources corresponding to the container group that provides the model inference service, including: If the container group is not in an overall slow running state, the control service grid performs the steps of determining the request length of the inference request in the external service component, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; If the container group is in an overall slow running state, the control service grid executes the steps of determining the request length of the inference request in the internal service component, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0012] Optionally, it also includes: If the container group is not in an overall slow running state, add the domain name of the external service component to the load balancing configuration of the service grid, and control the service grid to perform the steps of determining the request length of the inference request in the external service component according to the load balancing configuration, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; If the container group is in an overall slow running state, the domain name of the internal service component is added to the load balancing configuration of the service grid, and the service grid is controlled to determine the request length of the inference request in the external service component according to the load balancing configuration, and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0013] Optionally, before receiving the inference request sent to the model inference service, the method further includes: Receive a model inference service creation request, and determine custom resources required for the model inference service according to the model inference service creation request; The control resource controller resolves custom resources into container-native resources to create container groups, external service components, and internal service components corresponding to the model inference service, and marks version information for the container group; The control resource controller creates a service mesh for the model inference service and adds load balancing configuration to the service mesh.
[0014] Optionally, it also includes: Receive a model inference service update request, and determine the latest custom resources required by the model inference service according to the model inference service update request; The resource controller is controlled to convert the latest custom resources into container-native resources, obtain the new version of the container group, external service components, and internal service components of the model inference service, and mark the new version information for the new version of the container group; Control the resource controller to create a new version of the service mesh for the model inference service and add a load balancing configuration for the new version of the service mesh; Find and delete the old version of the container group, external service components, and internal service components of the Model Inference Service.
[0015] Optionally, add load balancing configuration to the service mesh, including: The control resource controller reads the global load balancing configuration from the preset configuration file and adds the global load balancing configuration to the service mesh.
[0016] Optionally, it also includes: The control strategy controller receives a load balancing configuration update request, updates a preset configuration file using the load balancing configuration update request, and updates the load balancing configuration in the service grid of each model inference service using the preset configuration file.
[0017] The present application also provides an electronic device, comprising: Memory for storing computer programs; The processor is used to implement the above service matching method when executing a computer program.
[0018] The present application also provides a computer program product, including a computer program or instructions, which implements the above-mentioned service matching method when the computer program or instructions are executed by a processor.
[0019] The present application also provides a non-volatile computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the above-mentioned service matching method is implemented.
[0020] The beneficial effect of the present application is that when receiving an inference request sent to a model inference service, the present application can determine the request length of the inference request, and can match the target container group for the inference request based on the request length and the amount of hardware resources corresponding to the container group providing the model inference service, and the model inference service is provided by at least two container groups. This is because the length of the inference request is closely related to the consumption of hardware resources by the model inference service. When the text input by the user through the inference request is long, the model inference service needs to consume a large amount of hardware resources to parse the user's input text, and it also tends to generate a longer output text, which also requires a large amount of hardware resources to store the output text. Therefore, the present application can match the appropriate target container group for the inference request based on the request length and the amount of hardware resources corresponding to the container group providing the model inference service, so as to achieve a better load balancing effect. The present application also provides an electronic device, a computer program product, and a computer-readable storage medium, which have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A flow chart of a service matching method provided in an embodiment of the present application; Figure 2 A schematic diagram of an inference service provided in an embodiment of the present application; Figure 3 A schematic diagram of another reasoning service provided by an embodiment of the present application; Figure 4 A schematic diagram of an inference service version update provided in an embodiment of the present application; Figure 5 A schematic diagram of updating a global load balancing configuration provided in an embodiment of the present application; Figure 6 A structural block diagram of a service matching device provided in an embodiment of the present application; Figure 7 A structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0024] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0025] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0026] With the continuous development of artificial intelligence technology, pre-trained language models have been widely used. Among them, pre-trained language models refer to neural network models that are pre-trained with a large amount of text data and used for natural language processing. They can be used for language dialogue, document summarization, article expansion, etc. To support the application deployment of pre-trained language models, the container orchestration platform can be used to deploy pre-trained language models in multiple container groups (Pods), and these container groups jointly provide model inference services to the outside. Among them, the container orchestration platform can create a container group (Pod) in the device node (Node), create a container (Container) in the container group, and deploy the model inference service in the container. At this time, the device node can provide the underlying hardware computing power for the container group, such as providing the container group with a processor, memory, graphics card, etc. For the same model inference service, the container orchestration platform can create container groups in different device nodes to support the operation of the model inference service through different device nodes; and the hardware configuration of different device nodes can be different.
[0027] In the related technology, although the container orchestration platform can provide a load balancing mechanism for multiple container groups, the load balancing effect achieved for the model inference service is not good, which can easily lead to load imbalance in the container group and affect the service quality. This is because the common load balancing mechanism mainly considers balancing the number of requests connected to each container group, without considering balancing the length of the requests connected to each container group; at the same time, unlike common network services, the performance of the model inference service is not only affected by the number of access requests, but also by the length of the access requests. When the request input by the user is long, the model inference service needs to consume a lot of hardware resources for input parsing and output generation, and it takes a lot of processing time; when the same container group processes a large number of long inference requests at the same time, its hardware resources will be consumed quickly, resulting in load imbalance.
[0028] In view of this, in order to solve the technical problem of how to improve the load balancing effect of container groups in model reasoning service scenarios, the present application can provide a service matching method, which can match a suitable target container group for the reasoning request based on the request length and the amount of hardware resources corresponding to the container group providing the model reasoning service, and can consider the performance loss of the model reasoning service caused by the request length, thereby achieving a better load balancing effect.
[0029] The service matching method provided by this application is introduced below. It should be noted that this method is executed by a container orchestration platform. This embodiment does not limit the hardware equipment for deploying the container orchestration platform, and can be set according to actual application requirements, for example, it can be deployed in a personal computer or a server. In particular, the container orchestration platform can be deployed in a server cluster, and multiple server devices can be used as device nodes for deployable container groups to improve overall processing performance.
[0030] For easier understanding, please refer to Figure 1 , Figure 1 A flow chart of a service matching method provided in an embodiment of the present application, the method may include: S100: Receive an inference request sent to a model inference service.
[0031] In this embodiment, the inference request refers to a request input by a user requesting the model inference service to perform model inference. The request may include the user's input data, which may be in the form of text, file, etc. It should be noted that the request length of the inference request is affected by the text length or file size, that is, the longer the text input by the user, or the larger the file size input by the user, the longer the request length of the inference request.
[0032] S200. Determine the request length of the inference request, and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; the model inference service is provided by at least two container groups.
[0033] In this embodiment, the model reasoning service is provided by at least two container groups, which can be deployed on the same device node or on different device nodes, and different device nodes can have different hardware configurations. Considering that the hardware resource consumption and processing time of the model reasoning service when processing reasoning requests of different lengths are different, especially when processing longer reasoning requests, it needs to consume a large amount of hardware resources to parse user input, and tends to generate longer output texts, which requires a large amount of hardware resources to generate and store the output texts, and takes a long time to process. Therefore, in order to achieve a better load balancing effect, this embodiment can determine the request length of the reasoning request, and match the target container group for the reasoning request in these multiple container groups according to the request length and the hardware resource amount corresponding to the container group providing the model reasoning service, so as to ensure that the hardware resource amount of the target container group can match the hardware resource consumption of the reasoning request, and ensure that the target container group can quickly process the reasoning request.
[0034] It should be noted that the hardware resource amount here can be either the native hardware resource amount of the container group or the remaining hardware resource amount of the container group. In addition, this embodiment does not limit the specific hardware resource amount, and can be memory amount, video memory amount, processor occupancy rate, etc.
[0035] Specifically, considering that the model reasoning service usually consumes a large amount of hardware resources and generates a long processing time when responding to a longer reasoning request, but consumes less hardware resources and can be processed quickly when responding to a shorter user input, this embodiment can distinguish between longer reasoning requests and shorter reasoning requests. First, this embodiment can distinguish between longer reasoning requests and shorter reasoning requests by determining whether the request length of the reasoning request is greater than the preset length. If the request length is greater than the preset length, this embodiment needs to determine the predicted resource consumption corresponding to the request length. Among them, the predicted resource consumption can be determined based on the historical resource consumption consumed by the model reasoning service in the past to process the reasoning request of the request length, such as taking the most frequently occurring historical resource consumption as the predicted resource consumption. Subsequently, this embodiment can determine the remaining hardware resources corresponding to each container group, and determine whether these remaining hardware resources can meet the predicted resource consumption, so as to select the target container group from these container groups that meet the requirements.
[0036] Based on this, according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service, the target container group is matched for the inference request, which may include: Step 11: Determine whether the request length is greater than a preset length.
[0037] Step 12: If the request length is greater than the preset length, determine the predicted resource consumption corresponding to the request length.
[0038] Step 13: According to the remaining amount of hardware resources corresponding to the container group, match the target container group among the container groups that meet the predicted resource consumption.
[0039] It is worth pointing out that, considering that a large amount of context information will be generated during the dialogue between the user and the inference model, and the inference model needs to refer to this context information when processing the current inference request, longer context information will also affect the performance of the model inference service. Therefore, when determining the predicted resource consumption, the length of the context information can also be considered. To this end, the present embodiment can search for historical inference requests that belong to the same dialogue as the current inference request, and estimate the length of the context information based on the historical inference requests. That is, when determining the predicted resource consumption, in addition to considering the request length of the current inference request, the present embodiment can also consider the request length of the historical inference request that belongs to the same dialogue as the current inference request. In this way, when determining the predicted resource consumption, the present embodiment can be closer to the actual working mode of the inference model, and thus can achieve a better load balancing effect.
[0040] Based on this, the method may further include: Step 21: According to the dialog identification information in the inference request, search for historical inference requests that belong to the same dialog as the inference request.
[0041] Determining the predicted resource consumption corresponding to the request length may include: Step 31: Determine the predicted resource consumption according to the request length of the inference request and the request length of the historical inference requests.
[0042] Furthermore, considering that there may be multiple container groups that meet the predicted resource consumption, and the hardware resource surpluses corresponding to these multiple container groups may be different, and thus the corresponding performance may also be different. On the basis of ensuring allocation fairness, in order to allocate as many container groups as possible with more hardware resource surpluses, this embodiment can use the container groups that meet the predicted resource consumption as candidate container groups, and set a selection weight for each candidate container group according to the hardware resource surpluses of each candidate container group. Among them, the selection weight is positively correlated with the hardware resource surplus, that is, the more the hardware resource surpluses, the greater the selection weight; the probability of the candidate container group being selected as the target container group is also positively correlated with the selection weight, that is, the greater the selection weight, the easier it is for the candidate container group to be selected as the target container group. In this way, this embodiment can increase the probability of container groups with larger hardware resource surpluses being selected, thereby achieving a better load balancing effect.
[0043] Based on this, matching the target container group in the container group that meets the predicted resource consumption may include: Step 41: Container groups that meet the predicted resource consumption are used as candidate container groups.
[0044] Step 42: According to the remaining amount of hardware resources of each candidate container group, a selection weight is set for each candidate container group; the selection weight is positively correlated with the remaining amount of hardware resources.
[0045] Step 43: Match the target container group for the inference request in the candidate container group according to the selection weight.
[0046] It is understandable that the above selection weights can be dynamically adjusted according to the real-time changes in the remaining amount of hardware resources. For example, a possible selection weight configuration is as follows: - match: - headers: content-length: gte: 1000 #When the text length is greater than or equal to 1000 characters, set the following weight.
[0047] route: - destination: pod: llm-high-resource-pods weight: 70 #The weight of the container group with more remaining resources is 70.
[0048] - route: - destination: pod: llm-normal-pods weight: 30 #The weight of the container group with more remaining resources is 30.
[0049] Furthermore, if the request length is less than the preset length, considering that shorter inference requests generally consume less hardware resources for the model inference service, in order to improve scheduling efficiency, scheduling can be performed according to the client to which the inference request belongs, or according to the number of requests currently processed by each container group. When scheduling according to the client to which the inference request belongs, considering that a large amount of context information will be generated during the conversation between the user and the inference model, and the inference model needs to refer to this context information when processing the current inference request, if the inference request of the same user is diverted to the same container group, it can be ensured that all of this context information is stored in the graphics card memory of the same container group, which can accelerate the inference process. Therefore, the container group corresponding to the client to which the inference request belongs can be used as the target container group. When scheduling according to the current number of requests processed by the container group, in order to balance the number of requests processed by each container group, the container group with the smallest number of requests processed can be used as the target container group.
[0050] Based on this, after determining whether the request length is greater than the preset length, the following may also be included: Step 51: If the request length is not greater than the preset length, the container group corresponding to the client to which the inference request belongs is used as the target container group, or the container group with the minimum request processing number is used as the target container group according to the current request processing number of the container group.
[0051] Furthermore, this embodiment does not limit how to determine the request length of the inference request, and can be set according to actual application requirements. For example, the request length can be read from the content length field (content-length) of the inference request, or the message length of the request message containing the inference request can be determined and used as the request length.
[0052] Based on this, the request length of the inference request is determined, which may include: Step 61: Read the request length from the content length field of the inference request; or, determine the message length of the request message containing the inference request, and use the message length as the request length.
[0053] S300: Send the inference request to the target container group, so that the target container group provides model inference service according to the inference request.
[0054] In this embodiment, after the matching of the target container group is completed, the inference request can be sent to the target container group, so that the target container group provides the model inference service according to the inference request.
[0055] Based on the above embodiment, when receiving an inference request sent to a model inference service, the present application can determine the request length of the inference request, and match the target container group for the inference request based on the request length and the amount of hardware resources corresponding to the container group providing the model inference service, wherein the model inference service is provided by at least two container groups. This is because the length of the inference request is closely related to the consumption of hardware resources by the model inference service. When the text input by the user through the inference request is long, the model inference service needs to consume a large amount of hardware resources to parse the user's input text, and it also tends to generate a longer output text, which in turn requires a large amount of hardware resources to store the output text. Therefore, the present application can match the appropriate target container group for the inference request based on the request length and the amount of hardware resources corresponding to the container group providing the model inference service, thereby achieving a better load balancing effect.
[0056] Based on the above embodiment, in order to achieve a better load balancing effect and avoid the model inference service from crashing due to overload under high concurrency, this embodiment can also improve the entry of the inference request. Based on this, the model inference service can include a service component, and the service component includes an external service component, an internal service component and a request buffer. Receiving an inference request sent to the model inference service may include: S101. Control the external service component to receive an inference request.
[0057] In this embodiment, the external service component is a unified entry for reasoning requests, which can provide an external domain name for accessing the model reasoning service. By accessing this external domain name, the client can send the reasoning request to the external service component.
[0058] S102: Determine whether the container group is in an overall slow running state according to the overall performance value corresponding to the container group.
[0059] S103: If the container group is in an overall slow running state, the external service component is controlled to send the inference request to the request buffer, and the internal service component is controlled to obtain the inference request from the request buffer.
[0060] Send the inference request to the target container group, which can include: S301: Send the inference request in the external service component or the inference request in the internal service component to the target container group.
[0061] In steps S102 and S103, in order to prevent the model inference service from crashing due to overload under high concurrency, a request buffer and an internal service component can be further introduced. The request buffer is used to temporarily store excessive inference requests, and it may include a waiting queue for storing inference requests. The internal service component is a substitute for the external service component, which stores the communication information of each container group and is used to directly connect to the container group. The internal service component can dynamically take out new inference requests from the request buffer according to the processing status of the container group, and send them to the container group for processing. It can be seen that by introducing the request buffer and the internal service component, excessive inference requests can be temporarily stored, thereby alleviating the high concurrency in the container group.
[0062] Furthermore, since the introduction of the request buffer and the internal service component prolongs the transmission path of the inference request, it is easy to affect the processing efficiency of the container group when the concurrency pressure is not large. Therefore, this embodiment can judge whether the container group is in an overall slow operation state according to the overall performance value corresponding to the container group. Among them, the overall slow operation state means that all container groups providing model inference services have slow operation conditions with slow processing speed and long response time, that is, there is a high concurrency situation. Therefore, in steps S102, S103 and S301, if it is determined that the container group is in an overall slow operation state, the external service component is controlled to send the inference request to the request buffer, and the internal service component is controlled to obtain the inference request from the request buffer, and then the inference request in the internal service component is sent to the container group. If it is determined that the container group is not in an overall slow operation state, the inference request can be directly sent from the external service component to the container group.
[0063] It should be noted that this embodiment does not limit the specific overall performance value, for example, it may be the overall video memory usage of each container group, the overall request response time, etc. If it is determined that the overall video memory usage of the container group is greater than a preset threshold, or the overall request response time is greater than a preset duration, it can be determined that the container group is in an overall slow running state.
[0064] Further, in order to improve the execution effect of load balancing, this embodiment can introduce a service grid in the model reasoning service, which is specifically used to perform load balancing operations (i.e., step S200). Among them, the service grid can realize intelligent traffic routing, load balancing, fault recovery and fuse mechanism, thereby significantly improving the stability and efficiency of the service. For example, through the traffic segmentation function of the service grid, requests can be proportionally distributed to different versions of reasoning service instances, supporting grayscale release and A / B testing; through the fuse and retry mechanism, the problem instance can be automatically isolated and the request can be retried when the backend service fails, avoiding service avalanche. In addition, the service grid also provides a wealth of observability tools that can monitor traffic, latency and error rate in real time, helping the operation and maintenance team to quickly locate and solve problems. It should be pointed out that since this embodiment specifically adjusts the entry of the reasoning request, so that the reasoning request needs to pass through the external service component to reach the container group, or needs to pass through the external service component, the request buffer, and the internal service component to reach the container group, it is necessary to dynamically adjust the position where the service grid performs load balancing. If the container group is not in an overall slow running state, you need to control the service grid to perform load balancing in the external service components; if the container group is in an overall slow running state, you need to control the service grid to perform load balancing in the internal service components.
[0065] Based on this, the model inference service also includes a service grid; determining the request length of the inference request, and matching the target container group for the inference request based on the request length and the amount of hardware resources corresponding to the container group providing the model inference service, which may include: S201. If the container group is not in an overall slow running state, the control service grid executes the steps of determining the request length of the inference request in the external service component, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0066] Specifically, the domain name of the external service component can be added to the load balancing configuration of the service grid, so that the service grid can automatically perform load balancing in the external service component according to the load balancing configuration.
[0067] Based on this, step S201 may include: Step 71: If the container group is not in an overall slow running state, add the domain name of the external service component to the load balancing configuration of the service grid, and control the service grid to determine the request length of the inference request in the external service component according to the load balancing configuration, and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0068] For easier understanding, please refer to Figure 2 , Figure 2A schematic diagram of an inference service provided by an embodiment of the present application. When the container group is not in an overall slow running state, the inference request can be directly sent from the external service component to the container group. At this time, the communication information of each container group can be copied from the internal service component to the external service component, so that the external service component points the service endpoint to the container group. When the load balancing configuration points to the external service component, the service grid performs load balancing processing in the external service component.
[0069] S202. If the container group is in an overall slow running state, the control service grid executes the steps of determining the request length of the inference request in the internal service component, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0070] Specifically, the domain name of the internal service component can be added to the load balancing configuration of the service grid, so that the service grid can automatically perform load balancing in the internal service component according to the load balancing configuration.
[0071] Based on this, step S202 may include: Step 81: If the container group is in an overall slow running state, add the domain name of the internal service component to the load balancing configuration of the service grid, and control the service grid to determine the request length of the inference request in the external service component according to the load balancing configuration, and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0072] It is worth pointing out that for longer reasoning requests, this embodiment sends the request to the final target container group in two scenarios: (1) load balancing from the external service component to the target container group; (2) forwarding from the request buffer to the internal service component, and then load balancing to the target container group through the internal service component. Therefore, when facing the special scenario of longer reasoning requests, it is necessary to configure the text capture logic of the service grid for both the internal service component and the external service component. In this way, both scenarios can achieve load balancing of longer reasoning requests.
[0073] For easier understanding, please refer to Figure 3 , Figure 3 A schematic diagram of another inference service provided by an embodiment of the present application, which shows the connection relationship and data transmission relationship between the external service component, the request buffer, the internal service component, and the container group. When the container group is in an overall slow running state, the service endpoint of the external service component points to the request buffer, and the request buffer forwards the request to the internal service component. When the load balancing configuration points to the internal service component, the service grid performs load balancing processing in the internal service component.
[0074] Based on the above embodiments, the creation and update methods of container groups, external service components, internal service components, and service grids are described in detail below. In one possible scenario, before receiving the inference request sent to the model inference service, the following may also be included: Step 91: Receive a model inference service creation request, and determine custom resources required for the model inference service according to the model inference service creation request.
[0075] Step 92: The control resource controller resolves the custom resources into container native resources to create a container group, external service components, and internal service components corresponding to the model inference service, and marks version information for the container group.
[0076] In this embodiment, in order to facilitate the creation of model reasoning services, the software resources and hardware resources required for the reasoning service can be abstracted into custom resources, and a resource controller can be set, and the resource controller can be used to convert the custom resources into container native resources, such as container groups and service components. In this way, this embodiment can easily create container groups and service components based on the required custom resources, thereby improving the convenience of creating model reasoning services.
[0077] Furthermore, to facilitate version management, after the container group is created, version information can be marked for the container group. The version information is used to indicate the version of the container group (model inference service).
[0078] Step 93: Control the resource controller to create a service grid for the model inference service and add a load balancing configuration to the service grid.
[0079] In this step, after completing the creation of the container group and service components, you can further control the resource controller to create a service grid for the model reasoning service and create a load balancing configuration for the service grid to provide a load balancing function for the model reasoning service.
[0080] In one possible case, the method may further include: Step 1001: receiving a model reasoning service update request, and determining the latest custom resource required by the model reasoning service according to the model reasoning service update request; Step 1002: Control the resource controller to convert the latest custom resource into a container native resource, obtain a new version of the container group, external service components, and internal service components of the model reasoning service, and mark the new version information for the new version of the container group; Step 1003: Control the resource controller to create a new version of the service grid for the model inference service, and add a load balancing configuration for the new version of the service grid; Step 1004: Find and delete the old version of the container group, external service components, and internal service components of the model reasoning service.
[0081] In this embodiment, when updating the resources of the model reasoning service, the resource controller will be controlled to recreate the container group and service components according to the latest custom resource requirements, and create a service grid for the recreated container group and service components. It is necessary to find the old version of the container group, external service component, and internal service component corresponding to the model reasoning service and delete the old version. For easy understanding, please refer to Figure 4 , Figure 4 A schematic diagram of an inference service version update provided in an embodiment of the present application. It can be seen that after creating a new version of the model inference service, the old version of the model inference service can be deleted. By setting the above-mentioned inference service update method, this embodiment can achieve a reliable upgrade of the inference service.
[0082] It can be seen that this embodiment realizes the full life cycle automated management of the reasoning service from creation, update to deletion through customized resource logic and supporting resource controller, which greatly simplifies the operation and maintenance process and reduces labor costs.
[0083] Furthermore, in order to improve the management efficiency of load balancing configuration, this embodiment sets a configuration file in the container orchestration platform and uses the configuration file to save the global load balancing configuration. When creating a service grid, the resource controller can be controlled to read the global load balancing configuration from the configuration file and set it for the service grid.
[0084] Based on this, add load balancing configuration to the service grid, which can include: Step 1101: Control the resource controller to read the global load balancing configuration from the preset configuration file and add the global load balancing configuration to the service grid.
[0085] Specifically, the global load balancing configuration can be saved in a configmap file.
[0086] In addition, when the global load balancing configuration is updated, in order to promptly apply the update to the service grid of each model reasoning service, this embodiment may also set a policy controller for monitoring and receiving load balancing configuration update requests. When receiving the load balancing configuration update request, the policy controller may use the load balancing configuration update request to update the preset configuration file, and use the preset configuration file to update the load balancing configuration in the service grid of each model reasoning service.
[0087] Based on this, the method may further include: Step 1201: The control policy controller receives a load balancing configuration update request, uses the load balancing configuration update request to update a preset configuration file, and uses the preset configuration file to update the load balancing configuration in the service grid of each model inference service.
[0088] For easier understanding, please refer to Figure 5 , Figure 5 A schematic diagram of updating a global load balancing configuration provided in an embodiment of the present application. When the policy controller detects that the global load balancing configuration has been updated, it can globally update the load balancing configuration of each inference service, such as synchronizing the latest load balancing configuration to each model inference service in the form of broadcast.
[0089] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0090] The embodiment of the present application also provides a service matching device. Figure 6 , Figure 6 This is a structural block diagram of a service matching device provided in an embodiment of the present application, and the device may include: A receiving module 601 is used to receive an inference request sent to a model inference service; The load balancing module 602 is used to determine the request length of the inference request and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; the model inference service is provided by at least two container groups; The request sending module 603 is used to send the reasoning request to the target container group, so that the target container group provides the model reasoning service according to the reasoning request.
[0091] Optionally, the load balancing module 602 may include: A judgment submodule is used to judge whether the request length is greater than a preset length; The consumption determination submodule is used to determine the predicted resource consumption corresponding to the request length if the request length is greater than the preset length; The first matching submodule is used to match the target container group among the container groups that meet the predicted resource consumption according to the remaining amount of hardware resources corresponding to the container groups.
[0092] Optionally, the first matching submodule may include: A candidate group setting unit, used to set a container group that meets the predicted resource consumption as a candidate container group; A weight setting unit, used to set a selection weight for each candidate container group according to the remaining amount of hardware resources of each candidate container group; the selection weight is positively correlated with the remaining amount of hardware resources; A matching setting unit is used to match a target container group for an inference request from a candidate container group according to a selection weight.
[0093] Optionally, the load balancing module 602 may further include: A query submodule, used to search for historical reasoning requests belonging to the same conversation as the reasoning request according to the conversation identification information in the reasoning request; The consumption determination submodule can be used to: The predicted resource consumption is determined according to the request length of the inference request and the request length of historical inference requests.
[0094] Optionally, the load balancing module 602 may further include: The first matching submodule is used to, if the request length is not greater than a preset length, use the container group corresponding to the client as the target container group according to the client to which the inference request belongs, or use the container group with the minimum request processing number as the target container group according to the current request processing number of the container group.
[0095] Optionally, the load balancing module 602 may include: The request length determination submodule is used to read the request length from the content length field of the inference request, or to determine the message length of the request message containing the inference request and use the message length as the request length.
[0096] Optionally, the model inference service includes a service component, and the service component includes an external service component, an internal service component, and a request buffer; The receiving module 601 may include: A first control submodule, used to control the external service component to receive the inference request; The performance detection submodule is used to determine whether the container group is in an overall slow running state according to the overall performance value corresponding to the container group; A second control submodule, configured to control the external service component to send the inference request to the request buffer if the container group is in an overall slow running state, and control the internal service component to obtain the inference request from the request buffer; The request sending module 603 can be used to: Send the inference request in the external service component or the inference request in the internal service component to the target container group.
[0097] Optionally, the model inference service also includes a service mesh; The load balancing module 602 may include: A second control submodule is used for controlling the service grid to perform the steps of determining the request length of the inference request in the external service component if the container group is not in an overall slow running state, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; The third control submodule is used to control the service grid to execute the steps of determining the request length of the inference request in the internal service component if the container group is in an overall slow running state, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0098] Optionally, the second control submodule can be used to add the domain name of the external service component to the load balancing configuration of the service grid if the container group is not in an overall slow running state, and control the service grid to perform the steps of determining the request length of the inference request in the external service component according to the load balancing configuration, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; The third control submodule can be used to add the domain name of the internal service component to the load balancing configuration of the service grid if the container group is in an overall slow running state, and control the service grid to execute the request length of the inference request in the external service component according to the load balancing configuration, and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
[0099] Optionally, the device may further include: A creation request receiving module is used to receive a model inference service creation request and determine the custom resources required for the model inference service according to the model inference service creation request; The container group creation module is used to control the resource controller to parse custom resources into container native resources to create the container group, external service components, and internal service components corresponding to the model inference service, and mark the version information for the container group; The service grid creation module is used to control the resource controller to create a service grid for the model inference service and add load balancing configuration to the service grid.
[0100] Optionally, the device may further include: An update request receiving module, used to receive a model reasoning service update request, and determine the latest custom resources required by the model reasoning service according to the model reasoning service update request; The container group update module is used to control the resource controller to convert the latest custom resources into container native resources, obtain the new version of the container group, external services and internal services of the model inference service, and mark the new version information for the new version of the container group; The service mesh update module is used to control the resource controller to create a new version of the service mesh for the model inference service and add load balancing configuration for the new version of the service mesh; The deletion module is used to find and delete the old version of the container group, external service components, and internal service components of the model inference service.
[0101] Optionally, the service mesh creation module may include: The load balancing configuration reading submodule is used to control the resource controller to read the global load balancing configuration from the preset configuration file and add the global load balancing configuration to the service grid.
[0102] Optionally, the device may further include: The load balancing configuration update module is used to control the policy controller to receive the load balancing configuration update request, use the load balancing configuration update request to update the preset configuration file, and use the preset configuration file to update the load balancing configuration in the service grid of each model inference service.
[0103] For the description of the features in the embodiment corresponding to the service matching device, reference may be made to the relevant description of the embodiment corresponding to the service matching method, which will not be described in detail here.
[0104] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above service matching method embodiments.
[0105] Please refer to Figure 7 , Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present application. The embodiment of the present application provides an electronic device 10, including a processor 11 and a memory 12; wherein the memory 12 is used to store a computer program; the processor 11 is used to execute the service matching method provided in the aforementioned embodiment when executing the computer program.
[0106] For the specific process of the above service matching method, reference may be made to the corresponding contents provided in the above embodiments, which will not be described in detail here.
[0107] Furthermore, the memory 12, as a carrier for storing resources, may be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the storage method may be temporary storage or permanent storage.
[0108] In addition, the electronic device 10 also includes a power supply 13, a communication interface 14, an input / output interface 15 and a communication bus 16; wherein the power supply 13 is used to provide working voltage for each hardware device on the electronic device 10; the communication interface 14 can create a data transmission channel between the electronic device 10 and an external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input / output interface 15 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0109] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned service matching method embodiments when running.
[0110] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0111] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above service matching method embodiments are implemented.
[0112] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned service matching method embodiments are implemented.
[0113] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0114] The above is a detailed introduction to a service matching method, device, electronic device, program product and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A service matching method, characterized in that: include: receiving an inference request sent to a model inference service; Determine the request length of the inference request, and match the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; the model inference service is provided by at least two container groups; The inference request is sent to the target container group, so that the target container group provides the model inference service according to the inference request.
2. The service matching method according to claim 1, characterized in that: Matching a target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service includes: Determine whether the request length is greater than a preset length; If the request length is greater than the preset length, determining the predicted resource consumption corresponding to the request length; According to the remaining amount of hardware resources corresponding to the container group, the target container group is matched in the container group that meets the predicted resource consumption.
3. The service matching method according to claim 2, characterized in that: Matching the target container group in the container group that meets the predicted resource consumption includes: taking a container group that meets the predicted resource consumption as a candidate container group; According to the remaining amount of hardware resources of each candidate container group, a selection weight is set for each candidate container group; the selection weight is positively correlated with the remaining amount of hardware resources; A target container group is matched for the inference request in the candidate container group according to the selection weight.
4. The service matching method according to claim 2, characterized in that: After determining whether the request length is greater than a preset length, the method further includes: If the request length is not greater than the preset length, the container group corresponding to the client to which the inference request belongs is used as the target container group, or the container group with the minimum request processing number is used as the target container group according to the current request processing number of the container group.
5. The service matching method according to claim 1, characterized in that: Determining a request length of the inference request includes: Reading the request length from the content length field of the inference request; Or, determine the message length of the request message containing the inference request, and use the message length as the request length.
6. The service matching method according to any one of claims 1 to 5, characterized in that: The model reasoning service includes a service component, and the service component includes an external service component, an internal service component and a request buffer; Receive an inference request sent to the model inference service, including: Controlling the external service component to receive the inference request; Judging, according to the overall performance value corresponding to the container group, whether the container group is in an overall slow running state; If the container group is in the overall slow running state, controlling the external service component to send the inference request to the request buffer, and controlling the internal service component to obtain the inference request from the request buffer; Sending the inference request to the target container group includes: The inference request in the external service component or the inference request in the internal service component is sent to the target container group.
7. The service matching method according to claim 6, characterized in that: The model reasoning service also includes a service grid; Determining a request length of the inference request, and matching a target container group for the inference request according to the request length and an amount of hardware resources corresponding to a container group providing the model inference service, including: If the container group is not in the overall slow running state, controlling the service grid to execute the steps of determining the request length of the inference request in the external service component, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; If the container group is in the overall slow running state, the service grid is controlled to execute the steps of determining the request length of the inference request in the internal service component, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
8. The service matching method according to claim 7, characterized in that: Also includes: If the container group is not in the overall slow running state, the domain name of the external service component is added to the load balancing configuration of the service grid, and the service grid is controlled to perform the steps of determining the request length of the inference request in the external service component according to the load balancing configuration, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service; If the container group is in the overall slow running state, the domain name of the internal service component is added to the load balancing configuration of the service grid, and the service grid is controlled to execute the steps of determining the request length of the inference request in the external service component according to the load balancing configuration, and matching the target container group for the inference request according to the request length and the amount of hardware resources corresponding to the container group providing the model inference service.
9. The service matching method according to claim 8, characterized in that: Before receiving an inference request sent to the model inference service, it also includes: Receive a model reasoning service creation request, and determine custom resources required for the model reasoning service according to the model reasoning service creation request; Controlling the resource controller to parse the custom resource into a container native resource to create a container group, an external service component, and an internal service component corresponding to the model reasoning service, and marking version information for the container group; Control the resource controller to create a service grid for the model reasoning service, and add a load balancing configuration to the service grid.
10. The service matching method according to claim 9, characterized in that: Also includes: Receiving a model reasoning service update request, and determining the latest custom resource required by the model reasoning service according to the model reasoning service update request; Control the resource controller to convert the latest custom resource into a container native resource, obtain a new version of the container group, external service components, and internal service components of the model reasoning service, and mark the new version information for the new version of the container group; Control the resource controller to create a new version of the service grid for the model reasoning service, and add a load balancing configuration for the new version of the service grid; Find and delete the old version of the container group, external service components, and internal service components of the model inference service.
11. The service matching method according to claim 9, characterized in that: Add load balancing configuration to the service grid, including: Control the resource controller to read the global load balancing configuration from a preset configuration file, and add the global load balancing configuration to the service grid.
12. The service matching method according to claim 11, characterized in that: Also includes: The control strategy controller receives a load balancing configuration update request, uses the load balancing configuration update request to update the preset configuration file, and uses the preset configuration file to update the load balancing configuration in the service grid of each model inference service.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the service matching method according to any one of claims 1 to 12 when executing the computer program.
14. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the service matching method according to any one of claims 1 to 12 is implemented.
15. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the service matching method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Service resource control system and service resource control method
CN102904942A
Access control method and apparatus
CN106649471A
Distributed high-concurrency power algorithm analysis system and method
CN116755903A
Task processing method, device and equipment running on cloud computing platform
CN119537040A
Model virtualization deployment method and device, storage medium and computer equipment
CN119597394A