Model servitization method and system and storage medium
By deploying grid agents and multiple model service instances in the service mesh, dynamic loading and routing of model services is achieved, and the problem of poor routing efficiency under dynamic loading of multiple models is solved, and the routing hit rate and efficiency are improved.
Patent Information
- Application Number
- CN202311860160.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In the scenario of dynamic loading of multiple models, due to frequent changes in the loading position of the model, the routing failure problem is caused. It takes multiple re-routing to find the required model, resulting in poor routing efficiency.
Deploy a grid proxy and multiple model service instances for model services in a service mesh, and a single model service instance is used to dynamically load the model. The grid agent receives the model call request and routes to the target service instance, which parses the model description information and continues to route to the model service instance that can provide the target model that conforms to the model description information.
In this way, even if the grid proxy does not correctly hit the model service instance, it no longer needs to be rerouted, which can effectively improve routing hit rate and efficiency.
Smart Images

Figure CN120238579A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to a model service method, system, and storage medium. Background Art
[0002] Model service means that after the model is deployed, services are provided externally through ways such as API interfaces for other services or clients to call.
[0003] Currently, the types of machine learning frameworks are already very diverse. For example, PyTorch, TensorFlow, Keras, etc. Therefore, existing model service solutions usually already support multiple machine learning frameworks. The model formats trained under different machine learning frameworks will be different. Considering the different usage requirements of users for model formats, some model service solutions further propose the idea of multi-model dynamic loading to control the deployment of models in different formats according to usage requirements.
[0004] However, under this idea of multi-model dynamic loading, since the loading location of the model changes frequently, routing failure problems often occur, and it is necessary to re-route many times before it is possible to find the required model, resulting in poor routing efficiency. Summary of the Invention
[0005] Multiple aspects of this application provide a model service method, system, and storage medium to improve the routing efficiency in the model service scenario.
[0006] An embodiment of this application provides a model service method. In a service mesh, a mesh proxy and multiple model service instances are deployed for a model service. A single model service instance is used to dynamically load a model. The method is applicable to a target service instance and includes:
[0007] Receiving a model call request forwarded by the mesh proxy;
[0008] The target service instance parses model description information from the model call request;
[0009] The target service instance continues to route the model call request to route the model call request to a model service instance that can provide a target model that meets the model description information;
[0010] Among them, the mesh proxy routes the received model call request to the target service instance.
[0011] The embodiment of the present application further provides a distributed microservice system, including a first node and multiple second nodes. A grid proxy deployed for the model service is included on the first node, and multiple model service instances deployed for the model service in the service grid are distributed on the multiple second nodes. A single model service instance is used to dynamically load a model;
[0012] The grid proxy is used to route the received model call request to a target service instance;
[0013] The target service instance is used to parse model description information from the model call request; and continue to route the model call request to route the model call request to a model service instance that can provide a target model that conforms to the model description information.
[0014] The embodiment of the present application further provides a service grid system, including a grid proxy and multiple model service instances deployed for the model service. A single model service instance is used to dynamically load a model. The grid proxy and the multiple model service instances cooperate with each other to execute the foregoing model service method.
[0015] The embodiment of the present application further provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the foregoing model service method.
[0016] In the embodiment of the present application, a model service architecture for realizing single service and multiple models based on a service grid is proposed. A grid proxy is deployed for the model service in the service grid, and multiple model service instances are deployed to provide multi-model dynamic loading support for the model service; on this basis, the grid proxy can route the received model call request to one of the target service instances, and further propose that the target service instance parses the model description information from the model call request to accurately perceive which model is required by the model call request. Moreover, the target service instance can continue to route the model call request, so as to route the model call request to a model service instance that can provide the required model. Accordingly, even if the grid proxy provided in the service gateway does not correctly hit the model service instance, it is no longer necessary to re-route, but the model call request can be routed to the correct model service instance based on the continue routing ability designed for the model service instance in this embodiment, thereby effectively improving the routing hit rate and improving the routing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 Schematic diagram of the structure of a service mesh system provided by an exemplary embodiment of the present application;
[0019] Figure 2 Schematic diagram of an optional deployment solution of a service mesh system provided by an exemplary embodiment of the present application;
[0020] Figure 3 Logical schematic diagram of an exemplary routing mechanism of a mesh proxy provided by an exemplary embodiment of the present application;
[0021] Figure 4 Schematic flow diagram of a model serviceification method provided by another exemplary embodiment of the present application;
[0022] Figure 5 Schematic diagram of the structure of a distributed microservice system provided by yet another exemplary embodiment of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0024] Before starting to describe the technical solutions provided by the embodiments of the present application, several technical concepts that may be involved in the present application are briefly explained as follows.
[0025] Model service:
[0026] Model service refers to a software service based on a machine learning model. It provides prediction or inference functions for the model by encapsulating the machine learning model into a callable API or service.
[0027] Model loading:
[0028] Model loading refers to the process of loading a machine learning model from a storage medium (such as a disk, a database, or cloud storage) into memory or a computing device for use in inference or other machine learning tasks.
[0029] Service mesh:
[0030] Often used to describe the microservice network that constitutes an application, as an infrastructure layer for handling inter-service communication. It is responsible for constructing the complex service topology of modern cloud-native applications to reliably deliver requests.
[0031] Mesh proxy:
[0032] A service mesh generally consists of a control plane and a data plane. Specifically, the control plane is a set of services running in a dedicated namespace. These services perform some control and management functions, including aggregating telemetry data, providing user-facing APIs, and providing control data to data plane proxies. Together, they drive the behavior of the data plane. The data plane, on the other hand, consists of a series of mesh proxies running beside each service instance.
[0033] The following describes in detail the technical solutions provided by each embodiment of the present application with reference to the accompanying drawings.
[0034] Figure 1 FIG. is a schematic structural diagram of a service mesh system provided by an exemplary embodiment of the present application. The service mesh system can be used to support the realization of model serviceization. Hereinafter, it is simply referred to as "service mesh".
[0035] Refer to Figure 1 , in this embodiment, a model serviceization architecture for implementing single-service multi-model based on a service mesh is proposed. Among them, single-service multi-model can be understood as supporting the provision of multiple models in a model service. For this purpose, in this embodiment, a mesh proxy and multiple model service instances are deployed for the model service in the service mesh. Among them, the multiple model service instances deployed for the model service share a mesh proxy, and the mesh proxy is used to route the in / out traffic of the model service.
[0036] In this embodiment, the mesh proxy can provide a unified way to manage and control the communication between the model service and other services, which can enhance the observability, resilience, and security of the communication between the model service and other services. The mesh proxy is responsible for intercepting, proxying, and processing requests and responses between the model service and other services. The mesh proxy is transparent to the model service and does not require any changes or code modifications to the model service. In this embodiment, the mesh proxy can handle model invocation requests from different protocols or formats and convert them into the formats required by the model service. It can convert a model invocation request from one protocol (such as HTTP or gRPC) to a protocol supported by the model service so that the model service can correctly process the request. In addition, decompression of request data or response data is also supported to control the size of the transmitted data. The mesh proxy can also collect and pass metrics, logs, and trace data related to the traffic of the model service for observation and monitoring. This helps to monitor the performance, availability, and security of the model service in real time and supports troubleshooting and performance optimization. Other functions of the mesh proxy in this embodiment are not limited and not exhausted. In addition, the implementation form of the mesh proxy is not limited in this embodiment, and the specific implementation and functions of the mesh proxy may vary depending on the tools and frameworks of the service mesh. For example, different implementation forms such as Envoy, Linkerd, and Istio can be used.
[0037] To support this system structure of a single service with multiple service instances, an optional deployment solution is provided in this embodiment. Figure 2 Schematic diagram of an optional deployment solution for a service mesh system provided in an exemplary embodiment of the present application. Refer to Figure 2 , in this optional deployment solution, multiple model service instances are mapped to a unified service access interface, which can be described as Figure 2 the model virtual service 104 shown in. The model virtual service 104 can serve as the access entry for all model service instances. The mesh proxy 113 deployed following the model virtual service 104 can route all model invocation requests to the appropriate model service instance. In practical applications, by using a fixed request service name to initiate a model invocation request to the model virtual service 104, the access and traffic to multiple models can be uniformly managed and controlled.
[0038] In this way, all models will only expose one service access interface (for example, Kubernetes Service). By mapping all models to the same service access interface, a simple and consistent way to access multiple models can be achieved. The client or other services can use the request service name corresponding to this service access interface to initiate requests, and some parameters can be carried in the model invocation request to specify the required model, which can effectively simplify the integration and use of the client.
[0039] Reference Figure 2 , in this alternative deployment solution, multiple runtime environments can be deployed for the model virtual service 104 in the service mesh, such as Figure 2 the runtime environments 105a and 105b shown in. The runtime environment can be understood as the runtime environment used by the model service instance. In the service mesh, a Deployment is a resource object used to define and manage the deployment and update of an application. In this alternative deployment solution, in a single runtime environment, multiple model service instances can be deployed and managed through the Deployment function. That is, multiple model service instances can use the same runtime environment. In this way, by using Deployment, multiple model service instances can be conveniently managed and scaled, and elastic upgrades and rollbacks can be achieved.
[0040] In addition, in this alternative deployment solution, static property tags can also be configured for multiple runtime environments respectively, which are used to describe the static properties of the runtime environment. Among them, the static properties can include but are not limited to the supported model formats and the supported compatible protocols, etc. It should be understood that the static properties can include various properties that do not change dynamically and various properties with a very low change frequency. No more examples of static properties are given here, and they can be configured as needed. Reference Figure 2 , in this alternative deployment solution, the static data tags configured for the runtime environment can be maintained in the runtime label data, such as the runtime label data 107a and 107b shown in the figure. The static property tags can help the grid proxy perform tasks such as routing, load balancing, and management. Usually, different static property tags can be configured for multiple runtime environments so that the multiple runtime environments are respectively used to support different inference requirements.
[0041] In this way, in this alternative deployment solution, multiple runtime environments 105 can be provided for the model service in the service mesh, and one or more model service instances can be deployed and managed in a single runtime environment 105. It should be understood that the above deployment solution is only exemplary. In this embodiment, other deployment solutions can also be adopted to implement the system structure of a single service with multiple service instances, and it is not limited to this. No more examples of deployment solutions are given here.
[0042] Based on this, in this embodiment, the grid proxy can receive the model call request and route the model call request to the target service instance. That is, the grid proxy provided for the model service in the service mesh can route the received model call request.
[0043] In this embodiment, the routing policy in the grid proxy is not limited. Based on the model service-oriented concept provided in this embodiment, there is no requirement for the routing hit rate of the grid proxy in this embodiment. In this embodiment, the model service instance hit by the grid proxy from multiple model service instances is described as the target model service instance. The routing mechanism in the grid proxy will not be elaborated here, nor will it be specifically limited.
[0044] In this embodiment, even if the target service instance hit by the grid proxy cannot provide the model required by the model call request, this routing will not exit. Instead, a function of continuing routing is designed on each model service instance to continue routing through the target service instance hit by the grid proxy.
[0045] Refer to Figure 1 , in this embodiment, for the target service instance, after receiving the model call request, the model description information can be parsed from the model call request. Among them, the model description information is used to point to one or more models through attribute description. The model description information may include but is not limited to model number ID, model name, model version, and model format, etc. Of course, these are only exemplary, and the model description information may also include other attribute information with identity pointing, and no more examples are given here. Through the parsed model description information, the target service instance can perceive which specific model or models are required by the model call request. For the convenience of writing, in this embodiment, the model that conforms to the model description information is described as the target model.
[0046] Continue to refer to Figure 1 , after parsing the model description information from the model call request, the target service instance can continue to route the model call request based on the model description information to route the model call request to the model service instance that can provide the target model. In this way, this routing correctly hits the model service instance. The correct hit here can be understood as that the hit model service instance can provide the model required by the model call request, so that the response to the model call request can be successfully completed.
[0047] Refer to Figure 1 It can be seen that in this embodiment, a routing mechanism combining the grid proxy and the model service instance is adopted to process and manage the traffic routing related to model services. As a basic routing component, if the grid proxy does not hit the correct model service instance, the target service instance hit by the grid proxy can be used as a continuing routing component, and according to the model description information in the model call request, the model call request can be routed to the model service instance that can provide the model required by the model call request. Therefore, through the routing mechanism proposed in this embodiment, the one-time hit rate of routing can be improved.
[0048] To support the continued routing ability in a model service instance, an exemplary deployment scheme within the service instance is proposed in this embodiment. Refer to Figure 2 In this exemplary deployment scheme: A model-specific proxy and a model server can be deployed in a single model service instance, such as Figure 2 the model-specific proxy 108a, the model-specific proxy 108b, and the model servers 106a and 106b shown in
[0049] Among them, the model-specific proxy can be used to perform the aforementioned continued routing operation. It should be understood that the model-specific proxy 108 is a proxy specifically proposed in this embodiment for a machine learning model, and is used to process and manage service traffic related to the model. As a continued routing component, the model-specific proxy can route according to the model description information in the model call request to ensure that the model call request can be correctly sent to the target model. It can be understood that the aforementioned grid proxy in this embodiment is responsible for routing at the service instance level, while the model-specific proxy here is responsible for routing at the model level, so that the model call request can correctly reach the required model and balance the load to improve system performance and scalability.
[0050] Refer to Figure 2 In the model service instance, the model server can be used to dynamically load multiple models. Among them, the model server can be understood as a service software that can load and manage multiple models simultaneously, and is used to provide a server environment for the models used in the service instance. For example, it can load and manage models in tensorflow format and pytorch format at the same time. In addition, the model server has the ability to dynamically load models, and can dynamically add, update, or delete models as needed during runtime without shutting down or restarting the server. The model-specific proxy is deployed beside the model server and is responsible for continuing to route the model call requests received by the service instance. Finally, the model call requests will be forwarded to the model server it follows through a certain model-specific proxy, so as to ensure that the model call requests are correctly routed to the required models and balance the load to improve system performance and scalability.
[0051] It should be understood that in addition to the components mentioned in the aforementioned exemplary deployment scheme in the service mesh of this embodiment, other components for supporting model serviceification can also be deployed. Continuing to refer to Figure 2 In the service mesh, the following can also be deployed:
[0052] Model service resource object 101
[0053] The model service resource object 101 is represented and managed through the Kubernetes Custom Resource, which is a mechanism to extend the native Kubernetes resources and can define and manage custom resource types in Kubernetes. Specifically, it defines a new resource type through the Custom ResourceDefinition (CRD) description and can create, update, and delete model services in a declarative manner. In this embodiment, the structure and behavior of the model service resources can be defined through the model service resource object 101, which can at least include the following aspects:
[0054] Model service description: This CRD defines the basic information of the model service, such as model name, version, labels, image, container configuration, etc. These information describe the characteristics and attributes of the model service.
[0055] Resource configuration: The required resource configuration for the model service can be specified, such as CPU, memory, storage, etc. This helps allocate appropriate resources for the model service to meet its performance and requirements.
[0056] Extended attributes: Other custom attributes can also be included to meet specific model usage requirements. These attributes can be defined and extended according to the actual situation.
[0057] Model service resource controller 102
[0058] In this embodiment, the model service resource controller 102 can be used as a controller component to monitor and manage the status and lifecycle of the model service resource object 101.
[0059] The model service resource controller 102 can be used to interact with the control plane of the service mesh and monitor operations such as the creation, update, and deletion of the model service resource object 101. It can ensure that the status of the model service resources is consistent with the desired status and take corresponding actions as needed. Specifically, in this embodiment, the functions of the model service resource controller 102 can include:
[0060] Creation and update: When receiving a create or update request for the CRD, the model service resource controller ensures that the relevant resources of the model service (such as Pods, Deployments, etc.) are created or updated according to the defined requirements.
[0061] Lifecycle management: The model service resource controller is responsible for managing the entire lifecycle of the model service resources, including start, stop, scale in and out, etc. It will create or delete replicas of the model service according to needs to meet the load and resource requirements.
[0062] Monitoring and regulation: The model service resource controller monitors the status, health, and performance metrics of model service resources. Based on this monitoring data, it can automatically perform operations such as load balancing, fault recovery, and horizontal scaling to ensure the availability and performance of the model service.
[0063] Error handling and rollback: When errors or faults occur in model service resources, the model service resource controller takes corresponding measures for error handling and rollback to ensure the stable operation of the model service.
[0064] Mesh controller 112
[0065] In a service mesh, mesh proxies are lightweight network proxies deployed between services, which are responsible for handling communication and traffic control between services. The mesh controller 112 is responsible for managing and configuring the mesh proxies to ensure that they work according to predefined policies and rules.
[0066] Model repository 111
[0067] In this embodiment, the model repository 111 refers to a data warehouse or system for storing and managing machine learning models. It is a place for centralized management and organization of machine learning models, providing functions such as model version control, metadata management, and access control. The model repository provides the function of storing models and can save model files, code, and related resources.
[0068] Model loader, such as Figure 2 the model loaders 110a and 110b shown in
[0069] In this embodiment, a model loader 110 can also be deployed in a model service instance, which is a component or module responsible for loading a machine learning model into memory and preparing it for prediction or inference. It is a bridge for loading the model from the storage medium into memory, providing the necessary resources and environment for the model service for model inference.
[0070] Through the mutual cooperation between the various components deployed in the service mesh, the concept of model serviceization in this embodiment can be supported. For example, the model service resource object 101, the model service resource controller 102, and the model service elastic scaler 103 can cooperate with each other to support the deployment of the model service and the elastic deployment of model service instances, etc. The model loader 110 and the model repository 111 cooperate with each other to support the model loading process in the model service instance. It should be understood that other components may also be included in the service mesh, which will be mentioned in other embodiments later, and the description of more components will not be given here for the time being.
[0071] In summary, in this embodiment, a model service architecture for implementing single-service multi-model based on a service mesh is proposed. A mesh proxy is deployed for the model service in the service mesh, and multiple model service instances are deployed to provide multi-model dynamic loading support for the model service. On this basis, the mesh proxy can route the received model call request to one of the target service instances, and further propose that the target service instance parses the model description information from the model call request to accurately perceive which model is required by the model call request. Moreover, the target service instance can perform further routing on the model call request, so as to route the model call request to the model service instance that can provide the required model. Accordingly, even if the mesh proxy provided in the service gateway does not correctly hit the model service instance, there is no need to re-route, but the model call request can be routed to the correct model service instance based on the further routing ability designed for the model service instance in this embodiment, thereby effectively improving the routing hit rate and improving the routing efficiency.
[0072] In the above or following embodiments, several preferred routing mechanisms are further proposed for the mesh proxy.
[0073] In a preferred routing mechanism: the mesh proxy can parse the static requirements for the model service instance from the model call request; discover candidate service instances that meet the static requirements from multiple model service instances; select a target service instance from the candidate service instances to route the model call request to the target service instance.
[0074] Among them, the static requirements can be understood as the requirements for the template service instance in terms of static attributes. That is to say, the static requirements can be used to point to a class of model service instances by describing the static attributes. In other words, the static requirements can characterize which type of model service instance the inference request needs. For this reason, in this preferred routing mechanism, static attributes can be defined for multiple model service instances respectively. The static attributes here are the same concept as the static attributes mentioned in the runtime environment deployment part in the foregoing embodiment. Therefore, the static attributes here can also include but are not limited to the model formats that can be supported and the compatible protocols that can be supported, etc.
[0075] In the foregoing embodiment, based on Figure 2Based on the provided optional deployment solutions, here, a single model service instance can inherit the static attribute tags configured in its runtime environment to define static attributes for the model service instance. For example, if the model formats supported by a runtime environment include tensorflow and keras, then each model service instance using this runtime environment also supports these two model formats, tensorflow and keras. On this basis, in practical applications, the required static attribute tags are carried in the specified fields of the model call request. In this way, the grid proxy can parse the static attribute tags from the model call request and discover candidate service instances with this static attribute tag from multiple model service instances.
[0076] To further improve the routing efficiency of the grid proxy, a solution for discovering candidate service instances is also proposed in this preferred routing mechanism: model service instances with the same static attribute tag can be divided into the same routing subset. In this way, the routing subset will be associated with the static attribute tag. On this basis, the grid proxy can look up the routing subset corresponding to the static attribute tag in the static requirements, and the model service instances included in the found routing subset are used as candidate service instances.
[0077] To support the grid proxy to perform routing based on the routing subset, refer to Figure 2 , in this solution, a dynamic subset routing policy generator 114 can be deployed in the service mesh. The dynamic subset routing policy generator 114 can be understood as a component that generates a dynamic routing policy 115 for the grid proxy according to the static attributes of the model service instances. It can collect the static attributes of each model service instance from the static attribute description data 107, and can determine the routing and traffic distribution of the model call request based on the static attributes of the model service instances to implement a flexible dynamic routing policy. These dynamic routing policies can include, but are not limited to, "routing the model call request to a specific model service instance", "distributing traffic proportionally to different model service instances", or performing traffic control according to other rules. The dynamic subset routing policy generator 114 is also used to apply the generated dynamic routing policy to the routing and traffic control of the model service (that is, configure it into the grid proxy 113 of the model service), so that through the model virtual service 104, the model call request is routed to the corresponding model service instance according to the dynamic routing policy, and the traffic is distributed and controlled according to the dynamic routing policy.
[0078] Among them, in the policy content of "routing the model call request to a specific model service instance", static attribute tags associated with each routing subset and model service instances included in each routing subset can be maintained. It should be understood that in the case where the static attribute tags of the model service instance change, the solution can support automatically re-partitioning the model service instance into a suitable routing subset.
[0079] Based on this, in this solution, the grid proxy can first route the model call request to a suitable routing subset based on the dynamic routing policy configured by the dynamic routing policy generator, and then further route the model call request to a certain model service instance in the hit routing subset.
[0080] Figure 3 It is a logical schematic diagram of an exemplary routing mechanism provided by an exemplary embodiment of this application. Refer to Figure 3 , after receiving the model call request, the grid proxy can parse out the static requirements from the specified fields of the model call request. Among them, the specified field can be the headers field in the model call request. For example, the static requirement carried in the model call request is "model format - tensorflow". Based on this, the grid proxy can find a suitable routing subset for the model call request. Refer to Figure 3 , for the routing subset on the left found by the grid proxy for the model call request, the static attribute tags associated with this routing subset include tags that support the tensorflow model format and tags that support the keras model format. That is, this routing subset includes the model format - tensorflow required by the model call request. Refer to Figure 3 , the routing subset hit by the grid proxy for the model call request includes two model service instances: model service instance a and model service instance b. These two model service instances use the same runtime environment 105a, and both of these model service instances support the model format - tensorflow. Therefore, they meet the static requirements of the model call request 1, and these two model service instances are used as candidate service instances corresponding to the model call request.
[0081] It should be understood that in this preferred routing mechanism, the model call request carries static requirements, and the dynamic routing policy configured in the grid proxy also includes policy content related to the static requirements. On this basis, the grid proxy can discover candidate service instances based on the static requirements according to this preferred routing mechanism.
[0082] For example, in this preferred routing mechanism, an exemplary policy content corresponding to the static requirements in the dynamic routing policy can be configured as:
[0083]
[0084]
[0085] In this example, if the grid agent parses the static requirements from the model call request as the model format tensorflow and the model format tensorflow, it can find the corresponding subset of routes according to the above policy content, and use the model service instances in the hit route subset as candidate service instances.
[0086] In addition, in this preferred routing mechanism, after the grid agent hits the candidate service instances, it can further select the target service instances from the candidate service instances according to the policy content related to load balancing or other policy content in the dynamic routing policy. The process of selecting the target service instances from the candidate service instances is not limited here. In some cases, it can even be randomly selected, and no more details will be provided here.
[0087] On this basis, in this preferred routing mechanism, an exemplary continued routing process in the target service instances may generally include:
[0088] 1. When receiving a model call request, the target service instance can query the model loading status registry 109 to find out whether there is a model service instance that has loaded the target model.
[0089] Some techniques can be used to implement this lookup process, such as:
[0090] A. Use a distributed cache to store the model loading status registry 109 and query it when needed.
[0091] B. Use a service discovery mechanism (such as Kubernetes service name, label selector) to dynamically obtain the network location of the model service instance holding the target model.
[0092] C. Use a distributed coordination service (such as ZooKeeper, etcd) to maintain the status and location information of each model service instance.
[0093] 2. If a model service instance that has loaded the target model is found, the target service instance can forward the model call request to one of the model service instances it has found.
[0094] In this way, the target service instance can dynamically find the model service instance that can provide the target model, so as to route the model call request to the correct location to achieve an efficient routing process.
[0095] Accordingly, in this preferred routing mechanism, static requirements can be parsed from the model call request and used as the routing basis. By means of a routing subset or the like, candidate service instances that meet the static requirements are filtered out. Finally, the grid proxy can route the model call request to a target service instance further selected from the candidate service instances. This enables the model call request to be routed to a model service instance within a reasonable range. Since the model service instances within this reasonable range all meet the static requirements of the model call request, the probability that the target service instance itself can provide the model required by the model call request can be effectively increased, and the probability of needing to cross model service instances during the subsequent routing process can be reduced, thereby improving the overall efficiency of routing.
[0096] In another preferred routing mechanism: The grid proxy can parse the service instance entry address from the model call request and route the model call request to the model service instance corresponding to the service instance entry address.
[0097] In this preferred routing mechanism, the service instance entry address is carried in the model call request. For example, the IP address of a certain model service instance, etc. In this case, the grid proxy can directly route the model call request to the model service instance corresponding to the service instance entry address.
[0098] In practical applications, in this case, the policy content corresponding to the service instance entry address will be configured in the grid proxy. This policy content can also be Figure 2 generated by the dynamic subset routing policy generator 114 in and configured into the grid proxy.
[0099] For example, in this preferred routing mechanism, an exemplary policy content corresponding to the service instance entry address in the dynamic routing policy can be configured as:
[0100]
[0101] In this example, if the grid proxy parses the service instance entry address from the model call request, it can execute the following command according to the above policy content: $curl –H “ip: XXX” http: / / modelmesh-serving-service / modelId, and directly route the model call request to the target service instance with the corresponding IP address (XXX in this command refers to the parsed IP address).
[0102] Accordingly, in this preferred routing mechanism, the grid agent can directly route the model call request to the corresponding model service instance according to the service instance entry address parsed from the model call request. In this routing mechanism, the routing efficiency of the grid agent layer can be effectively improved. Therefore, although the hit accuracy rate under this routing mechanism is lower, from the overall routing efficiency perspective, there is still a positive improvement.
[0103] In summary, in this embodiment, several preferred routing mechanisms are proposed for the grid agent. It should be understood that these routing mechanisms are only exemplary, and this embodiment is not limited thereto. Additionally, in practical applications, as described above, the dynamic routing policy in the grid agent can include multiple policy contents to support multiple routing mechanisms. For the grid agent, it can select the appropriate policy content according to the parameter type carried in the model call request. For example, if the model call request carries static requirements, the first preferred routing mechanism described above can be selected; if the model call request carries the service instance entry address, the second preferred routing mechanism described above can be selected. The specific selection rules can be configured as needed in the grid agent and are not limited here. In this way, in this embodiment, the grid agent can implement routing the model call request to the target service instance.
[0104] In the above or following embodiments, several preferred continue-routing mechanisms are also provided for the target service instance.
[0105] In a preferred continue-routing mechanism: A model loading status registry can be deployed in the service mesh. Refer to Figure 2 , the model loading status registry 109 can be used to store and manage the model loading status of each model service instance. The model loading status can describe the model description information corresponding to the models already loaded on the model service instance, the number of models already loaded, and the remaining loading space, etc. Of course, it can also include other metadata and configuration information related to model loading, and no more examples are given here. Additionally, a database such as etcd is usually used to implement this model loading status registry to provide auxiliary functions such as storage and detection during the continue-routing process of each model service instance.
[0106] Based on this, after receiving the model call request forwarded by the grid agent, the target service instance can query this model loading status registry. After the query, the following several situations may occur:
[0107] In one case, the target service instance queries from this model loading status registry that it has loaded the target model required by the model call request. In this case, the target service instance can respond to this model call request through itself, that is, respond to this model call request through the target model already loaded on itself. Refer to Figure 3, if the model call request requires Model A, and the model call request is routed by the grid proxy to model service instance a, that is, Figure 3 the model service instance a in becomes the target service instance. In this case, since Model A has been loaded on model service instance a, model service instance a can respond to this model call request through itself.
[0108] In another case, the target service instance queries the model loading status registry and finds that it does not have the target model required by the model call request. In this case, the target service instance can route the model call request to any model service instance that has loaded the required target model. That is, the target service instance can query the model loading status registry to find one or more model service instances that have loaded the target model, and route the model call request to one of the queried model service instances. Here, when the target service instance selects which model service instance to finally route to from the queried model service instances, it can be implemented according to the load balancing mechanism or the random mechanism, etc., which is not limited here. In this way, the target service instance can route the model call request to other model service instances that have loaded the target model when it does not have the target model itself. Refer to Figure 3 , if the model call request requires Model A, since both model service instances a and b support providing Model A, therefore, the model call request may be routed by the grid proxy to model service instance b, that is, Figure 3 the model service instance b in becomes the target service instance. In this case, Model A has not been loaded on model service instance b, but model service instance b can query the model loading status registry and learn that Model A has been loaded on model service instance a, then model service instance b can continue to route the model call request to model service instance a, and finally, model service instance a will respond to this model call request.
[0109] In another case, the target service instance queries the model loading status registry and finds that there is no model service instance that has loaded the required target model. In this case, the target model is not loaded on any of the model service instances. In this regard, on the one hand, the target service instance can select any model service instance that can support the loading of the required target model as the specified service instance, and route the model call request to this specified service instance. To support the target service instance in successfully selecting the specified service instance, static attribute tags of each model service instance can be stored and managed in the model loading status registry. In practical applications, these static attribute tags can be obtained from the various runtime tag data mentioned above. In this way, the target service instance can accurately select the specified service instance that can support the loading of the required target model during the process of querying the model loading status registry. On the other hand, since the selected specified service instance has not loaded the target model, the target service instance also transmits a model loading request to the model service elastic scale-out controller deployed in the service mesh to drive the model service elastic scale-out controller to control the specified service instance to load the target model.
[0110] Among them, referring to Figure 2 , the model service elastic scale-out controller 103 can manage the loading quantity of each model and when and where to load, etc. After receiving the model loading request transmitted by the target service instance, the model service elastic scale-out controller 103 can parse from the model loading request which model needs to be loaded and to which model service instance it needs to be loaded. Furthermore, the model service elastic scale-out controller 103 can control the specified model service instance to load the target model.
[0111] In addition, in addition to being able to control model loading according to the model loading request, the model service elastic scale-out controller 103 can at least also execute the following control logics:
[0112] 1. Allocate at least two model service instances for the most recently used models. That is to say, ensure that at least two copies of all "most recently used" models are loaded.
[0113] 2. Expand the number of model service instances for models whose load exceeds a specified threshold. That is to say, if the request load of a certain model exceeds a certain threshold, the number of copies of the model can be automatically expanded. Moreover, when the load is large enough, the model can have copies in each model service instance.
[0114] 3. Actively control the unloaded models to be loaded when the loaded models are released or there is free loading space. That is to say, when there is free space or to replace a loaded model that has not been used for a long time, actively load the currently unloaded model.
[0115] 4. According to the load heuristic algorithm and the last usage time, gradually reduce the number of replicas of models with more than 2 replicas to 2, and then from 2 to 1 in a timely manner. This control logic can be preferentially implemented on the model service instance with the highest load.
[0116] Of course, the control logic in the above aspects is only exemplary. The model service elastic scaling controller 103 can also execute other aspects of control logic, and no more examples will be given here. In this way, through these automatic management mechanisms, the number of replicas and the loading location of the models can be controlled more intelligently to optimize the inference performance and resource utilization.
[0117] In this way, in this preferred continue routing mechanism, the target service instance can utilize the model loading status registry deployed in the service mesh to sense the model loading status in each model service model, and continue to route the model call request based on the model loading status to ensure that the model call request is routed to the model service instance that can provide the required target model.
[0118] In some other alternative continue routing mechanisms, the model service instances can also synchronize the model loading status with each other. In this continue routing mechanism, each model service instance can independently maintain the global information of the model loading status. When continue routing is required, the target service instance can query the global information of the model loading status maintained by itself to complete the continue routing.
[0119] It should be understood that the above several continue routing mechanisms are only exemplary, and this embodiment is not limited thereto, and no more examples will be given here.
[0120] Accordingly, in this embodiment, after receiving a model call request, the target service instance can sense the target model required by the model call request through the parsed model description information, and can also sense the real-time model loading status on each model service instance. Through these two aspects of sensing, it can be ensured that the model call request is routed to the model service instance that can provide the target model, thereby ensuring that the model call request can be correctly responded to.
[0121] Reference Figure 3 , continuing from the exemplary deployment scheme inside the model service instance in the previous text, the continue routing operation is executed by the model-specific proxy in the model service instance. Combining with the routing operation executed by the mesh proxy in this embodiment. Reference Figure 3 , the model-specific proxies of multiple model service instances can communicate with each other to support the implementation of the continue routing operation. Based on this, an exemplary routing process for a model call request generally includes:
[0122] 1. After receiving a model invocation request, the load balancer in the grid proxy routes the request to an available target service instance screened by the dynamic routing policy.
[0123] 2. In the model-specific proxy of the target service instance, the target model required can be determined by parsing model description information such as the model id in the request metadata header.
[0124] 3. The model-specific proxy in the target service instance can check whether the target model has been loaded in the model server it follows by querying the model loading status registry, etc. If it has been loaded, it triggers the model server it follows to respond to the model invocation request.
[0125] 4. If the target model has not been loaded in the model server it follows, the model-specific proxy in the target service instance can route the model invocation request to the model-specific proxy in other model service instances that can provide the target model, and this model-specific proxy will continue to forward the model invocation request to the model server it follows, so that the model invocation request can successfully reach the model server that can provide the target model.
[0126] Through this routing process, it can be ensured that the model invocation request can be effectively routed to the model service instance holding the target model in one routing process, without the need for the process to re-route, which can effectively improve the routing hit rate and thus ensure efficient model services.
[0127] Based on the technical concepts or optional implementation methods provided in the above embodiments, the model service solution provided by this application can produce the following technical improvements:
[0128] This application proposes a method and system for multi-model dynamic loading and serviceization, which is used for the management and routing of model services, and provides the processing ability in the case of high-scale, high-density and frequently changing model usage.
[0129] This application provides a method and system for generating a dynamic subset routing policy from static attribute tags, which converges the routing range of the target service instance when selecting requests and improves the request hit rate.
[0130] This application provides a routing mechanism based on the combination of grid proxy and model-specific proxy, which is used to process and manage the traffic routing related to model services. The grid proxy is used as the basic routing component. On this basis, the model-specific proxy is used as the secondary routing component. According to the model description information such as the model name in the request, by querying the data information in the model loading status registry, the request can be routed to the corresponding model service instance. This can narrow down the target service subset for querying and improve the one-time hit rate of routing.
[0131] This application proposes a mechanism for loading models on demand during use, effectively reducing resource occupancy and optimizing the use of available computing resources.
[0132] This application proposes a mechanism for decoupling the concerns of scaling model services from service-specific inference technologies to support capabilities independent of language or model service frameworks and to support customizable inference service runtimes.
[0133] This application provides a mechanism for dynamically handling situations with a high degree of model changes, providing a good plug-and-play extension mechanism, and solving the single point of failure of model services after the number of models grows to a large scale.
[0134] Figure 4 The flowchart of a model service method provided for another exemplary embodiment of this application, which can be executed based on a service mesh. Refer to Figure 4 , in the service mesh, a mesh proxy and multiple model service instances are deployed for model services. A single model service instance is used to dynamically load models. The method may include:
[0135] Step 400, the mesh proxy routes the received model call request to the target service instance;
[0136] Step 401, the target service instance parses the model description information from the model call request;
[0137] Step 402, the target service instance further routes the model call request to route the model call request to a model service instance that can provide a target model meeting the model description information.
[0138] In an optional embodiment, the mesh proxy routing the received model call request to the target service instance may include:
[0139] The mesh proxy parses the static requirements for the model service instance from the model call request;
[0140] The mesh proxy discovers candidate service instances that meet the static requirements from the multiple model service instances;
[0141] The mesh proxy selects a target service instance from the candidate service instances to route the model call request to the target service instance.
[0142] In an optional embodiment, the static requirements include static attribute tags. Model service instances with the same static attribute tags are divided into the same routing subset. The mesh proxy discovering candidate service instances that meet the static requirements from the multiple model service instances may include:
[0143] The grid proxy searches for a routing subset corresponding to a static attribute tag in the static requirement;
[0144] The model service instance contained in the found routing subset is used as the candidate service instance; the static attribute tag is used to describe the model format and other static attributes supported by the model service instance.
[0145] In an optional embodiment, multiple runtime environments are deployed for the model service in the service grid, one or more model service instances are created in a single runtime environment, and a single model service instance inherits the static attribute tags configured in the runtime environment to which it belongs.
[0146] In an optional embodiment, the grid proxy routes the received model call request to the target service instance, which may include:
[0147] The grid proxy parses the service instance entry address from the model call request;
[0148] The grid proxy routes the model call request to the model service instance corresponding to the service instance entry address.
[0149] In an optional embodiment, a model loading status registry is deployed in the service grid, and the model loading status corresponding to each of the multiple model service instances is maintained in the model loading status registry. The target service instance continues to route the model call request to route the model call request to the model service instance that can provide a target model that conforms to the model description information, which may include:
[0150] The target service instance queries the model loading status registry;
[0151] If it is determined that the required target model is not loaded, the model call request is routed to any model service instance that has loaded the required target model.
[0152] In an optional embodiment, the method may further include:
[0153] If the target service instance determines that there is no model service instance that has loaded the required target model, any model service instance that can support loading of the required target model is selected as the designated service instance;
[0154] Transmitting a model loading request to a model service elastic scaler deployed in the service grid to drive the model service elastic scaler to control the designated service instance to load the target model;
[0155] The model call request is routed to the specified service instance.
[0156] In an alternative embodiment, the method may further include:
[0157] The model service elastic scale-out device allocates at least two model service instances to the most recently used model; and / or
[0158] The model service elastic scale-out device scales out the number of model service instances for a model whose load exceeds a specified threshold; and / or
[0159] The model service elastic scale-out device actively controls an unloaded model to be loaded when a loaded model is released or there is free loading space; and / or,
[0160] The model service elastic scale-out device gradually reduces the number of replicas corresponding to a model whose number of replicas exceeds a preset threshold to a target value.
[0161] In an alternative embodiment, a single model service instance includes a model-specific proxy and a model server, where the model-specific proxy is used to perform the continue routing operation, and the model server is used to dynamically load multiple models.
[0162] In an alternative embodiment, multiple model service instances deployed for the model service in the service mesh are mapped to a unified service access interface.
[0163] It should be noted that for the technical details in the foregoing embodiments of the model service method, reference may be made to the relevant descriptions in the foregoing system embodiments. To save space, they will not be elaborated here, but this should not cause any loss to the protection scope of this application.
[0164] It should be noted that the execution subject of each step of the method provided in the foregoing embodiments may be determined according to the node devices where each component in the service mesh is located. For example, for the mesh proxy, it may be borne by node device A in the service mesh cluster, and the model service instance may be borne by the working group POD in the service mesh cluster as the workload in the working group. The execution subject of each step will not be elaborated one by one here.
[0165] In addition, in some processes described in the foregoing embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The operation numbers such as 401, 402, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.
[0166] Figure 5A schematic structural diagram of a distributed microservices system provided for another exemplary embodiment of the present application, refer to Figure 5 , the distributed microservices system may include a first node 50 and a plurality of second nodes 51.
[0167] The distributed microservices system provided in this embodiment can be applied in a cloud computing environment. Cloud computing is one of the fastest developing trends in computer technology, which involves providing hosted services through a network. Cloud computing is a service delivery model aiming to enable on-demand network access to a shared pool of configurable computing resources (such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), and these resources can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud computing environment provides computing and storage resources as services to end users, and the end users can send requests to the provided services for processing. The cloud computing environment may include one or more cloud computing nodes, and local computing devices used by cloud consumers, such as personal digital assistants (PDAs) or cellular phones, desktop computers, laptop computers, and / or automotive computer systems, can communicate with them. Nodes can communicate with each other, and they can be physically or virtually grouped (not shown) in one or more networks, such as private, community, public, or hybrid clouds, or combinations thereof. This allows the cloud computing environment to provide infrastructure, platform, and / or software as services, and cloud consumers do not need to maintain resources on local computing devices. Among them, the computing nodes and the cloud computing environment can communicate with any type of computerized device through any type of network and / or network addressable connection (such as using a web browser).
[0168] The first node and the second node in this embodiment can adopt cloud computing nodes in the cloud computing environment. In terms of physical implementation forms, the first node and the second node can be computer systems, servers, or portable electronic devices such as communication devices, and can also be container groups PODs, etc., which can be used with many other general or special computing system environments or configurations. Among them, examples of these known computing systems, environments, and / or configurations that can be used with the first node and the second node include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs (personal computers), small computer systems, large computer systems, and distributed cloud computing environments, including any of the above systems or devices, etc.
[0169] It should be understood that although a detailed description of cloud computing is disclosed here, the implementation of the distributed microservices system provided in this embodiment is not limited to the cloud computing environment. On the contrary, the distributed microservices system provided in this embodiment can be implemented in combination with any other type of computing environment known now or developed in the future.
[0170] The microservice distribution system provided in this embodiment can use a service mesh as the infrastructure. The first node 50 may include a mesh proxy 52 deployed for the model service in the service mesh. Multiple model service instances are also deployed for the model service in the service mesh, which are distributed on multiple second nodes 51. A single model service instance can be used to dynamically load the model.
[0171] On this basis, the mesh proxy 52 can be used to route the received model call request to the target service instance;
[0172] The target service instance 53 can be used to parse the model description information from the model call request; continue to route the model call request to route the model call request to a model service instance that can provide a target model that meets the model description information.
[0173] Regarding the technical details involved in the mesh proxy 52 and the target service instance 53, reference can be made to the relevant descriptions in the respective embodiments of the foregoing model service method. To save space, they will not be elaborated here, but this should not cause loss of the protection scope of this application.
[0174] Correspondingly, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executed in the foregoing method embodiment.
[0175] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0176] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows or multiple flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the processes and / or blocks Figure 1 of one or more of the processes and / or blocks Figure 1 specified in the one or more of the processes and / or blocks.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 of one or more of the processes and / or blocks Figure 1 specified in the one or more of the processes and / or blocks.
[0179] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0180] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0181] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A model service method, characterized in that, A grid proxy and multiple model service instances are deployed for a model service in a service grid, and a single model service instance is used to dynamically load a model. The method is applicable to a target service instance in the multiple model service instances, and the method includes: Receiving a model call request forwarded by the grid proxy; Parsing model description information from the model call request; Continue routing the model call request to route the model call request to a model service instance that can provide a target model that conforms to the model description information; The grid proxy routes the received model call request to the target service instance.
2. The method according to claim 1, characterized in that, The process of the grid proxy routing the received model call request to the target service instance includes: The grid proxy parses the static demand for the model service instance from the model call request; The grid agent determines a candidate service instance that meets the static requirement from the multiple model service instances; The grid proxy selects a target service instance from the candidate service instances to route the model call request to the target service instance.
3. The method according to claim 2, wherein The static requirement includes a static attribute tag, and the model service instances with the same static attribute tag are divided into the same routing subset. The process in which the grid agent discovers a candidate service instance that meets the static requirement from the multiple model service instances includes: The grid proxy searches for a routing subset corresponding to a static attribute tag in the static requirement; The model service instance contained in the found routing subset is used as the candidate service instance; the static attribute tag is used to describe the model format and other static attributes supported by the model service instance.
4. The method according to claim 3, characterized in that, Multiple runtime environments are deployed for the model service in the service grid, one or more model service instances are created in a single runtime environment, and a single model service instance inherits the static attribute tags configured in the runtime environment to which it belongs.
5. The method according to claim 1, characterized in that, The process of the grid proxy routing the received model call request to the target service instance includes: The grid proxy parses the service instance entry address from the model call request; The grid proxy routes the model call request to the model service instance corresponding to the service instance entry address; Among them, the model service instance corresponding to the service instance entry address is used as the target service instance.
6. The method according to claim 1, characterized in that, A model loading status registry is deployed in the service grid, and the model loading status corresponding to each of the multiple model service instances is maintained in the model loading status registry. The model call request is continuously routed to route the model call request to a model service instance that can provide a target model that conforms to the model description information, including: Query the model loading status registry; If it is determined that the target service instance has not loaded with the required target model, the model call request is routed to any model service instance that has loaded with the required target model.
7. The method according to claim 6, characterized in that, Also includes: If it is determined that there is no model service instance that has loaded the required target model among the target service instances, any model service instance that can support the loading of the required target model is selected as the specified service instance; A model loading request is sent to the model service elastic scaling controller deployed in the service mesh to drive the model service elastic scaling controller to control the specified service instance to load the target model; The model call request is routed to the specified service instance.
8. The method according to claim 7, characterized in that It further includes: The model service elastic scaling controller assigns at least two model service instances to the most recently used model; and / or, The model service elastic scaling controller expands the number of model service instances for a model whose load exceeds a specified threshold; and / or, The model service elastic scaling controller actively controls an unloaded model to be loaded when the loaded model is released or there is free loading space; and / or, The number of replicas corresponding to a model with the number of replicas of the model service device exceeding a preset threshold is gradually reduced to a target value.
9. The method according to claim 1, wherein A single model service instance includes a model-specific proxy and a model server. The model-specific proxy is used to perform the continued routing operation, and the model server is used to dynamically load multiple models.
10. The method according to claim 1, wherein Multiple model service instances deployed for the model service in the service mesh are mapped to a unified service access interface.
11. A distributed microservices system, characterized in that, It includes a first node and multiple second nodes. The first node contains a mesh proxy deployed for the model service in the service mesh. Multiple model service instances deployed for the model service in the service mesh are distributed on the multiple second nodes. A single model service instance is used to dynamically load a model; The mesh proxy is used to route the received model call request to the target service instance; The target service instance is used to parse model description information from the model call request; Perform continued routing on the model call request to route the model call request to a model service instance that can provide a target model that conforms to the model description information.
12. A service mesh system, characterized in that, It includes a mesh proxy and multiple model service instances deployed for the model service. A single model service instance is used to dynamically load a model. The mesh proxy and the multiple model service instances cooperate to perform the model service method according to any one of claims 1-10.
13. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the model service method according to any one of claims 1-10.
Citation Information
Cited By
Model service scheduling method and system and computing equipment
CN120896988A