Model serving method, system and storage medium

By deploying grid agents and multiple model service instances in the service mesh, and using the target service instance to analyze model description information for continuous routing, the problem of poor routing efficiency during dynamic loading of multiple models is solved, and efficient model service routing is achieved.

WO2025139255A1PCT designated stage expired Publication Date: 2025-07-03HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126193
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-10-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing model service solution has the problem of poor routing efficiency during the dynamic loading of multi-models, especially when the model loading location changes frequently, it often requires multiple re-routing to find the required model.

Method used

Deploy a grid agent and multiple model service instances for model services in the service mesh. Receive model call requests through the grid agent, and resolve model description information by the target service instance to continue routing to ensure that the request is routed to a service instance that can provide the required model.

Benefits of technology

It improves the routing hit rate, improves routing efficiency, reduces the number of rerouting times, and improves the routing efficiency in the service-oriented scenario of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126193_03072025_PF_FP_ABST
    Figure CN2024126193_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a model serving method, a system and a storage medium. It is proposed deploying in a service mesh a mesh agent and multiple model service instances for a model service to support multi-model dynamic loading. On this basis, the mesh agent can route a received model invoking request to one target service instance, and the target service instance senses which model is required by the model invoking request. In addition, the target service instance can keep routing the model invoking request, so as to route the model invoking request to a model service instance capable of providing a required model. Therefore, even if the mesh agent provided in a service gateway does not correctly hit the model service instance, rerouting is no longer required and instead, on the basis of the continuous routing capability designed for the model service instances in the embodiments, the model invoking request can be routed to a correct model service instance, thereby effectively improving the routing hit rate, and improving the routing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

A model service method, system and storage medium

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311860160.9 and application name “A Model Service Method, System and Storage Medium”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The present disclosure relates to the field of cloud computing technology, and in particular to a model service method, system, and storage medium. Background Art

[0003] Model servitization means that after the model is deployed, it provides services to the outside world through application programming interfaces (APIs) and other means for other services or clients to call.

[0004] Currently, there are a wide variety of machine learning frameworks, including PyTorch, TensorFlow, and Keras. Therefore, existing model-as-a-service solutions typically support multiple machine learning frameworks. Models trained using different machine learning frameworks often have different formats. Considering the varying user requirements for model formats, some model-as-a-service solutions have further introduced the concept of dynamic multi-model loading to control the deployment of models in different formats based on usage needs.

[0005] However, under this idea of ​​dynamic loading of multiple models, routing failures often occur because the loading location of the model changes frequently. It is necessary to re-route multiple times before the required model can be found, resulting in poor routing efficiency.

[0006] Summary of the Invention

[0007] Various aspects of the present disclosure provide a model servitization method, system, and storage medium to improve routing efficiency in model servitization scenarios.

[0008] The present disclosure provides a model servitization method, in which a grid agent and multiple model service instances are deployed for a model service in a service grid. A single model service instance is used to dynamically load a model. The method is applicable to a target service instance and includes:

[0009] receiving a model call request forwarded by the grid proxy;

[0010] The target service instance parses the model description information from the model call request;

[0011] The target service instance continues to route the model call request to route the model call request to a model service instance that can provide a target model that conforms to the model description information;

[0012] The grid proxy routes the received model call request to the target service instance.

[0013] The present disclosure also provides a distributed microservice system, comprising a first node and multiple second nodes, wherein the first node includes a grid proxy deployed for a model service in a service grid, multiple model service instances deployed for the model service in the service grid are distributed on the multiple second nodes, and a single model service instance is used to dynamically load a model;

[0014] The grid proxy is used to route the received model call request to the target service instance;

[0015] The target service instance is used to parse the model description information from the model call request; and further route the model call request to a model service instance that can provide a target model that conforms to the model description information.

[0016] An embodiment of the present disclosure also provides a service grid system, including a grid agent and multiple model service instances deployed for model services, a single model service instance is used to dynamically load models, and the grid agent and the multiple model service instances cooperate to execute the aforementioned model servitization method.

[0017] The embodiment of the present disclosure also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, causes the one or more processors to execute the aforementioned model servitization method.

[0018] An embodiment of the present disclosure also provides a computer program product, including a computer program, which implements the aforementioned model servitization method when executed by a processor.

[0019] In the embodiment of the present disclosure, a model service architecture for realizing single service and multiple models based on a service grid is proposed. A grid proxy is deployed for the model service in the service grid, and multiple model service instances are deployed to provide dynamic loading support for multiple models for the model service. On this basis, the grid proxy can route the received model call request to one of the target service instances, and it is further proposed that the target service instance parses the model description information from the model call request to accurately perceive which model is required by the model call request. Moreover, the target service instance can perform continued routing on the model call request, thereby routing the model call request to the model service instance that can provide the required model. Accordingly, even if the grid proxy provided in the service gateway does not correctly hit the model service instance, there is no need to re-route. Instead, the model call request can be routed to the correct model service instance based on the continued routing capability designed for the model service instance in this embodiment, thereby effectively improving the routing hit rate and improving routing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0021] FIG1 is a schematic structural diagram of a service grid system provided by an exemplary embodiment of the present disclosure;

[0022] FIG2 is a schematic diagram of an optional deployment solution for a service grid system provided by an exemplary embodiment of the present disclosure;

[0023] FIG3 is a logical diagram of an exemplary routing mechanism of a grid proxy provided by an exemplary embodiment of the present disclosure;

[0024] FIG4 is a flow chart of a model servitization method provided by another exemplary embodiment of the present disclosure;

[0025] FIG5 is a schematic structural diagram of a distributed microservice system provided by yet another exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions, and advantages of the present disclosure more clear, the technical solutions of the present disclosure will be clearly and completely described below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present disclosure.

[0027] Before describing the technical solutions provided by the various embodiments of the present disclosure, several technical concepts that may be involved in the present disclosure are briefly explained as follows.

[0028] Model Serving:

[0029] Model service refers to a software service based on machine learning models. It provides prediction or inference capabilities for machine learning models by encapsulating them into callable APIs or services.

[0030] Model loading:

[0031] Model loading refers to the process of loading a machine learning model from a storage medium (such as a disk, database, or cloud storage) into memory or a computing device for use in inference or other machine learning tasks.

[0032] Service Mesh:

[0033] Often used to describe the network of microservices that make up an application, it serves as the infrastructure layer that handles inter-service communication. It is responsible for reliably delivering requests by forming the complex service topology of modern cloud-native applications.

[0034] Grid Agent:

[0035] A service mesh typically consists of a control plane and a data plane. Specifically, the control plane is a set of services running in a dedicated namespace. These services perform management and control functions, including aggregating telemetry data, providing user-facing APIs, and delivering control data to data plane proxies. Together, they drive the behavior of the data plane. The data plane, in turn, consists of a series of mesh proxies running alongside each service instance.

[0036] The technical solutions provided by various embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0037] FIG1 is a schematic diagram of the structure of a service grid system provided by an exemplary embodiment of the present disclosure. This service grid system can be used to support the implementation of model servitization. It is hereinafter referred to as the "service grid."

[0038] Referring to Figure 1, this embodiment proposes a model service-based architecture that implements multiple models in a single service based on a service grid. "Single service, multiple models" can be understood as supporting multiple models within a single model service. To this end, this embodiment proposes deploying a grid proxy and multiple model service instances within the service grid for the model service. The multiple model service instances deployed for the model service share a single grid proxy, which is used to route inbound and outbound traffic for the model service.

[0039] In this embodiment, the grid proxy provides a unified way to manage and control communication between the model service and other services, which can enhance the observability, resilience, and security of communication between the model service and other services. The grid proxy is responsible for intercepting, proxying, and processing requests and responses between the model service and other services. The grid proxy is transparent to the model service and does not require any changes or code modifications to the model service. In this embodiment, the grid proxy can process model call requests from different protocols or formats and convert them into the format required by the model service. It can convert model call requests from one protocol (such as HTTP or gRPC) to a protocol supported by the model service so that the model service can correctly process the request. In addition, it also supports decompression of request data or response data to control the size of the transmitted data. The grid proxy can also collect and transmit metrics, logs, and trace data related to the model service's traffic for observation and monitoring. This helps to monitor the performance, availability, and security of the model service in real time and supports troubleshooting and performance optimization. This embodiment does not limit or list other functions of the grid proxy. In addition, this embodiment does not limit the implementation form of the grid proxy. The specific implementation and functions of the grid proxy may vary depending on the tools and frameworks of the service grid. For example, different implementation forms such as Envoy, Linkerd, and Istio can be used.

[0040] In order to support this system structure of single service and multiple service instances, an optional deployment scheme is provided in this embodiment. Figure 2 is a schematic diagram of an optional deployment scheme for a service grid system provided by an exemplary embodiment of the present disclosure. Referring to Figure 2, in this optional deployment scheme, multiple model service instances are mapped to a unified service access interface, which can be described as the model virtual service 104 shown in Figure 2. The model virtual service 104 can serve as the access entrance for all model service instances, and the grid agent 113 deployed following the model virtual service 104 can route all model call requests to the appropriate model service instance. In actual applications, by using a fixed request service name to initiate a model call request to the model virtual service 104, access and traffic to multiple models can be uniformly managed and controlled.

[0041] This way, all models expose only one service access interface (for example, a Kubernetes Service). By mapping all models to the same service access interface, a concise and consistent approach to accessing multiple models is achieved. Clients or other services can initiate requests using the request service name corresponding to the service access interface. Model call requests can include parameters to specify the desired model, effectively simplifying client integration and usage.

[0042] Referring to Figure 2, in this optional deployment scheme, multiple runtime environments can be deployed for the model virtual service 104 in the service grid, such as the runtime environment 105a and the runtime environment 105b shown in Figure 2. The runtime environment can be understood as the runtime environment used by the model service instance. In the service grid, a deployment is a resource object used to define and manage the deployment and update of applications. In this optional deployment scheme, in a single runtime environment, multiple model service instances can be deployed and managed through the deployment function. That is, multiple model service instances can use the same runtime environment. In this way, by using deployment, multiple model service instances can be easily managed and expanded, and elastic upgrades and rollbacks can be achieved.

[0043] In addition, in this optional deployment solution, static attribute tags can be configured for each of the multiple runtime environments to describe the static attributes possessed by the runtime environment. Among them, static attributes may include but are not limited to supported model formats and supported compatible protocols. It should be understood that static attributes may include various attributes that do not change dynamically and various attributes that change very infrequently. More examples of static attributes are not given here, and they can be configured as needed. Referring to Figure 2, in this optional deployment solution, the static data tags configured for the runtime environment can be maintained in the runtime tag data, such as the runtime tag data 107a and the runtime tag data 107b shown in the figure. Static attribute tags can help the grid agent perform routing, load balancing, and management. Generally, different static attribute tags can be configured for multiple runtime environments so that multiple runtime environments can be used to support different reasoning requirements.

[0044] Thus, in this optional deployment solution, multiple runtime environments 105 can be provided for the model service in the service grid, and one or more model service instances can be deployed and managed in a single runtime environment 105. It should be understood that the above deployment solution is merely exemplary, and other deployment solutions can be used in this embodiment to implement a system structure of a single service with multiple service instances, and the present invention is not limited to this. Further examples of deployment solutions are not provided here.

[0045] Based on this, in this embodiment, the grid proxy can receive the model call request and route the model call request to the target service instance. In other words, the grid proxy provided for the model service in the service grid can route the received model call request.

[0046] In this embodiment, the routing strategy within the grid proxy is not limited. Based on the model-based service concept presented in this embodiment, there is no requirement for the grid proxy's routing hit rate. In this embodiment, the model service instance that the grid proxy hits from multiple model service instances is described as the target model service instance. The routing mechanism within the grid proxy is not described in detail here, nor is it specifically limited.

[0047] In this embodiment, even if the target service instance hit by the grid proxy cannot provide the model required by the model call request, the routing will not be exited. Instead, a function of continuing routing is designed on each model service instance to continue routing through the target service instance hit by the grid proxy.

[0048] Referring to Figure 1, in this embodiment, for the target service instance, after receiving the model call request, the model description information can be parsed from the model call request. Among them, the model description information is used to point to one or more models through attribute description. The model description information may include but is not limited to the model number ID, model name, model version and model format, etc. Of course, these are only exemplary. The model description information may also contain other attribute information with identity pointing, and no further examples are given here. Through the parsed model description information, the target service instance can perceive which specific model or models are required by the model call request. For the convenience of writing, in this embodiment, the model that meets the model description information is described as the target model.

[0049] Continuing with Figure 1, after parsing the model description information from the model call request, the target service instance can then route the model call request based on the model description information to a model service instance that can provide the target model. This ensures that the routing correctly targets the model service instance. A correct hit here means that the model service instance that is hit can provide the model required by the model call request, thereby successfully responding to the model call request.

[0050] As shown in Figure 1, this embodiment utilizes a routing mechanism that combines a grid proxy with model service instances to handle and manage model service-related traffic routing. The grid proxy, as the foundational routing component, acts as a follow-up routing component if the correct model service instance is not found. The target service instance found by the grid proxy then serves as a follow-up routing component. Based on the model description information in the model call request, it routes the model call request to a model service instance that can provide the model requested. Therefore, the routing mechanism proposed in this embodiment improves the one-time hit rate of routing.

[0051] To support continued routing within a model service instance, this embodiment proposes an exemplary deployment solution within a service instance. Referring to FIG2 , in this exemplary deployment solution, a model-specific proxy and a model server can be deployed within a single model service instance, such as model-specific proxy 108 a, model-specific proxy 108 b, and model servers 106 a and 106 b shown in FIG2 .

[0052] Among them, the model-specific proxy can be used to perform the aforementioned continued routing operation. It should be understood that the model-specific proxy 108 is a proxy specifically for machine learning models proposed in this embodiment, which is used to process and manage service traffic related to the model. As a continued routing component, the model-specific proxy can perform routing based on the model description information in the model call request, ensuring that the model call request can be correctly sent to the target model. It can be understood that the aforementioned grid proxy in this embodiment is used to be responsible for routing at the service instance level, while the model-specific proxy here is responsible for routing at the model level, so that the model call request can correctly reach the required model.

[0053] Referring to Figure 2, the model server in the model service instance can be used to dynamically load multiple models. The model server can be understood as a service software that can load and manage multiple models at the same time, and is used to provide a server environment for the models used in the service instance. For example, models in the tensorflow format and the pytorch format can be loaded and managed at the same time. In addition, the model server has the ability to dynamically load models and can dynamically add, update, or delete models as needed at runtime without shutting down or restarting the server. The model-specific proxy is deployed next to the model server and is responsible for continuing to route the model call requests received by the service instance. Ultimately, the model call request will be forwarded to the model server it follows through a model-specific proxy, thereby ensuring that the model call request is correctly routed to the required model and balancing the load to improve system performance and scalability.

[0054] It should be understood that in addition to the components mentioned in the exemplary deployment scheme, the service grid in this embodiment may also deploy other components to support model servitization. Continuing with Figure 2, the service grid may also deploy:

[0055] Model Service Resource Object 101

[0056] The model service resource object 101 is represented and managed as a Kubernetes custom resource. It is a mechanism for extending native Kubernetes resources that allows you to define and manage custom resource types in Kubernetes. Specifically, it defines a new resource type through a custom resource definition (CRD) that allows you to create, update, and delete model services in a declarative manner. In this embodiment, the structure and behavior of the model service resource can be defined through the model service resource object 101, which can include at least the following aspects:

[0057] Model service description: This CRD defines the basic information of the model service, such as model name, version, label, image, container configuration, etc. This information describes the characteristics and properties of the model service.

[0058] Resource configuration: You can specify the resource configuration required by the model service, such as CPU, memory, storage, etc. This helps allocate appropriate resources to the model service to meet its performance and needs.

[0059] Extended attributes: Other custom attributes can also be included to meet specific model usage requirements. These attributes can be defined and extended based on actual conditions.

[0060] Model service resource controller 102

[0061] In this embodiment, the model service resource controller 102 may serve as a controller component for monitoring and managing the status and life cycle of the model service resource object 101 .

[0062] The model service resource controller 102 can be used to interact with the control plane of the service grid and monitor operations such as the creation, update, and deletion of model service resource objects 101. It can ensure that the state of model service resources is consistent with the expected state and take appropriate actions as needed. Specifically, in this embodiment, the functions of the model service resource controller 102 may include:

[0063] Create and update: When receiving a create or update request for a CRD, the model service resource controller ensures that the relevant resources of the model service (such as Pod, Deployment, etc.) are created or updated according to the defined requirements.

[0064] Lifecycle management: The Model Service Resource Controller is responsible for managing the entire lifecycle of Model Service resources, including starting, stopping, scaling, etc. It creates or deletes replicas of the Model Service as needed to meet load and resource requirements.

[0065] Monitoring and Adjustment: The Model Service Resource Controller monitors the status, health, and performance metrics of model service resources. Based on this monitoring data, it automatically performs load balancing, fault recovery, and horizontal scaling to ensure the availability and performance of the model service.

[0066] Error handling and rollback: When errors or failures occur in model service resources, the model service resource controller will take appropriate measures to handle the errors and rollback to ensure the stable operation of the model service.

[0067] Grid Controller 112

[0068] In a service grid, grid proxies are lightweight network proxies deployed between services, responsible for handling communication and traffic control between services. The grid controller 112 is responsible for managing and configuring grid proxies, ensuring that they work according to predefined policies and rules.

[0069] Model Repository 111

[0070] In this embodiment, model repository 111 refers to a data warehouse or system used to store and manage machine learning models. It is used to centrally manage and organize machine learning models, providing functions such as model version control, metadata management, and access control. The model repository provides model storage capabilities, including model files, code, and related resources.

[0071] Model loaders, such as model loaders 110a and 110b shown in FIG2

[0072] In this embodiment, a model loader 110 is also deployed in the model service instance. This loader is responsible for loading the machine learning model into memory and preparing it for prediction or inference. It bridges the gap between the model and the storage medium and provides the model service with the necessary resources and environment for model inference.

[0073] The model servitization concept in this embodiment is supported through the coordination of various components deployed within the service grid. For example, the model service resource object 101, the model service resource controller 102, and the model service elastic scaler 103 can coordinate with each other to support the deployment of model services and the elastic deployment of model service instances. The model loader 110 and the model repository 111 coordinate with each other to support the model loading process within the model service instances. It should be understood that the service grid may also include other components, which will be discussed in other embodiments below, and a detailed description of these components is omitted here.

[0074] In summary, in this embodiment, a model service architecture based on a service grid is proposed to realize a single service and multiple models. A grid proxy is deployed for the model service in the service grid, and multiple model service instances are deployed to provide dynamic loading support for multiple models for the model service. On this basis, the grid proxy can route the received model call request to one of the target service instances, and further proposes that the target service instance parses the model description information from the model call request to accurately perceive which model is required by the model call request. Moreover, the target service instance can perform continued routing on the model call request, thereby routing the model call request to the model service instance that can provide the required model. Accordingly, even if the grid proxy provided in the service gateway does not correctly hit the model service instance, there is no need to re-route. Instead, the model call request can be routed to the correct model service instance based on the continued routing capability designed for the model service instance in this embodiment, thereby effectively improving the routing hit rate and improving routing efficiency.

[0075] In the above or below embodiments, several preferred routing mechanisms are further proposed for the grid proxy.

[0076] In a preferred routing mechanism: the grid proxy can parse the static requirements for the model service instance from the model call request; discover candidate service instances that meet the static requirements from multiple model service instances; select the target service instance from the candidate service instances to route the model call request to the target service instance.

[0077] Among them, static requirements can be understood as requirements for template service instances in terms of static attributes. That is, static requirements can be used to point to a type of model service instance through the description of static attributes. In other words, static requirements can characterize which type of model service instance is required for the inference request. To this end, in this preferred routing mechanism, static attributes can be defined for multiple model service instances respectively. The static attributes here are the same concept as the static attributes mentioned in the runtime environment deployment part of the aforementioned embodiment. Therefore, the static attributes here can also include but are not limited to supported model formats and supported compatible protocols.

[0078] Based on the optional deployment solution provided in FIG2 in the aforementioned embodiment, a single model service instance can inherit the static attribute tags configured in the runtime environment to which it belongs to define static attributes for the model service instance. For example, if the model formats supported by a runtime environment include tensorflow and keras, then each model service instance using the runtime environment also supports these two model formats, tensorflow and keras. On this basis, in actual applications, the required static attribute tags are carried in the specified fields in the model call request. In this way, the grid agent can parse the static attribute tags from the model call request and find candidate service instances with the static attribute tags from multiple model service instances.

[0079] To further improve the routing efficiency of the grid proxy, this optimal routing mechanism also proposes a solution for discovering candidate service instances: model service instances with the same static attribute tag can be grouped into the same routing subset. This routing subset is then associated with the static attribute tag. Based on this, the grid proxy can search for the routing subset corresponding to the static attribute tag in the static requirement, and the model service instances contained in the found routing subset are then used as candidate service instances.

[0080] To support grid proxies in routing based on routing subsets, referring to Figure 2 , in this solution, a dynamic subset routing policy generator 114 can be deployed within the service grid. This dynamic subset routing policy generator 114 can be understood as a component that generates dynamic routing policies 115 for grid proxies based on the static properties of model service instances. It can collect the static properties of each model service instance from the static property description data 107 and, based on these static properties, determine the routing and traffic distribution of model call requests, thereby implementing flexible dynamic routing policies. These dynamic routing policies may include, but are not limited to, "routing model call requests to a specific model service instance," "proportionally distributing traffic to different model service instances," or performing traffic control based on other rules. The dynamic subset routing policy generator 114 is also used to apply the generated dynamic routing policies to the routing and traffic control of the model service (i.e., configured within the grid proxies 113 of the model service). This allows the model virtual service 104 to route model call requests to the corresponding model service instance according to the dynamic routing policies, and to distribute and control traffic according to the dynamic routing policies.

[0081] The "route model call requests to specific model service instances" policy maintains the static attribute tags associated with each routing subset and the model service instances contained in each routing subset. It should be understood that if the static attribute tags of a model service instance change, this solution can automatically reassign the model service instance to the appropriate routing subset.

[0082] Based on this, in this solution, the grid proxy can first route the model call request to the appropriate routing subset based on the dynamic routing strategy configured by the dynamic routing strategy generator, and then further route the model call request to a model service instance in the hit routing subset.

[0083] FIG3 is a logical diagram of an exemplary routing mechanism provided by an exemplary embodiment of the present disclosure. Referring to FIG3 , after receiving a model call request, the grid agent can parse the static requirements from the specified field of the model call request. The specified field can be the headers field in the model call request. For example, the static requirement carried in the model call request is "model format - tensorflow". Based on this, the grid agent can find a suitable routing subset for the model call request. Referring to FIG3 , the grid agent finds the routing subset on the left side for the model call request. The static attribute labels associated with the routing subset include labels that support the tensorflow model format and labels that support the keras model format. That is, the routing subset contains the model format required for the model call request - tensorflow. Referring to Figure 3, the routing subset hit by the grid proxy for the model call request contains two model service instances: model service instance a and model service instance b. These two model service instances use the same runtime environment 105a. Both model service instances support the model format - tensorflow. Therefore, they meet the static requirements of the model call request 1. These two model service instances serve as candidate service instances corresponding to the model call request.

[0084] It should be understood that in this preferred routing mechanism, the model call request carries static requirements, and the dynamic routing strategy configured in the grid agent also includes policy content related to the static requirements. On this basis, the grid agent can discover candidate service instances based on the static requirements in accordance with the preferred routing mechanism.

[0085] For example, in this preferred routing mechanism, an exemplary policy content corresponding to a static requirement in a dynamic routing policy may be configured as follows:

[0086] In this example, if the grid proxy parses the static requirements from the model call request as model format tensorflow and model format tensorflow, it can find the corresponding routing subset according to the above policy content and use the model service instance in the matching routing subset as the candidate service instance.

[0087] Furthermore, in this preferred routing mechanism, after identifying candidate service instances, the grid proxy can further select a target service instance from the candidate service instances based on load balancing-related or other policy details within the dynamic routing policy. The process for selecting a target service instance from the candidate service instances is not limited here; in some cases, the target service instance can even be selected randomly, which is not discussed in detail here.

[0088] On this basis, in this preferred routing mechanism, an exemplary continued routing process in the target service instance may generally include:

[0089] 1. After receiving a model call request, the target service instance may query the model loading status registry 109 to find out whether there is a model service instance that has loaded the target model.

[0090] Several techniques can be used to implement this lookup process, such as:

[0091] A. Use distributed cache to store the model loading status registry 109 and query it when needed.

[0092] B. Use service discovery mechanisms (e.g., Kubernetes service names, label selectors) to dynamically obtain the network location of the model service instance holding the target model.

[0093] C. Use distributed coordination services (such as ZooKeeper and etcd) to maintain the status and location information of each model service instance.

[0094] 2. If a model service instance that has loaded the target model is found, the target service instance can forward the model call request to one of the model service instances it has found.

[0095] In this way, the target service instance can dynamically find the model service instance that can provide the target model, thereby routing the model call request to the correct location to achieve an efficient routing process.

[0096] Based on this, in this preferred routing mechanism, static requirements can be parsed from model call requests and used as the basis for routing. Through methods such as routing subsets, candidate service instances that meet the static requirements can be screened. Ultimately, the grid proxy can route the model call request to a target service instance further selected from the candidate service instances. This allows the model call request to be routed to a model service instance within a reasonable range. Since the model service instances within this reasonable range all meet the static requirements of the model call request, it can effectively increase the probability that the target service instance itself can provide the model required by the model call request, reduce the probability of needing to cross model service instances during the continued routing process, and thus improve the overall efficiency of routing.

[0097] In another preferred routing mechanism: the grid proxy can parse the service instance entry address from the model call request; and route the model call request to the model service instance corresponding to the service instance entry address.

[0098] In this preferred routing mechanism, the model call request carries the service instance entry address, such as the IP address of a model service instance. In this case, the grid proxy can directly route the model call request to the model service instance corresponding to the service instance entry address.

[0099] In actual applications, in this case, the grid proxy will be configured with policy content corresponding to the service instance entry address. This policy content can also be generated by the dynamic subset routing policy generator 114 in Figure 2 and configured in the grid proxy.

[0100] For example, in this preferred routing mechanism, an exemplary policy content corresponding to the service instance entry address in the dynamic routing policy can be configured as follows:

[0101] In this example, if the grid proxy resolves the service instance entry address from the model call request, it can execute the following command according to the above policy content: $curl –H "ip:XXX" http: / / modelmesh-serving-service / modelId to route the model call request directly to the target service instance with the corresponding IP address (XXX in the command refers to the resolved IP address).

[0102] Based on this, in this preferred routing mechanism, the grid proxy can directly route the model call request to the corresponding model service instance according to the service instance entry address parsed from the model call request. Under this routing mechanism, the routing efficiency of the grid proxy layer can be effectively improved. Therefore, although the hit accuracy rate under this routing mechanism will be lower, from the perspective of the overall routing efficiency, there is still a positive improvement.

[0103] In summary, in this embodiment, several preferred routing mechanisms are proposed for the grid proxy. It should be understood that these routing mechanisms are merely exemplary and this embodiment is not limited thereto. In addition, in actual applications, as described above, the dynamic routing policy in the grid proxy can include multiple policy contents to support multiple routing mechanisms. For the grid proxy, it can select appropriate policy contents based on the parameter types carried in the model call request. For example, if the model call request carries static requirements, the first preferred routing mechanism mentioned above can be selected, while if the model call request carries the service instance entry address, the second preferred routing mechanism mentioned above can be selected. The specific selection rules can be configured as needed in the grid proxy and are not limited here. In this way, in this embodiment, the grid proxy can route the model call request to the target service instance.

[0104] In the above or below embodiments, several preferred continued routing mechanisms are also provided for the target service instance.

[0105] In a preferred continued routing mechanism, a model loading status registry can be deployed in the service grid. Referring to Figure 2 , the model loading status registry 109 can be used to store and manage the model loading status of each model service instance. The model loading status can describe the model description information corresponding to the models loaded on the model service instance, the number of loaded models, and the remaining loading space. Of course, it can also include other metadata and configuration information related to model loading, but further examples are not provided here. In addition, the model loading status registry is typically implemented using a database such as etcd to provide auxiliary functions such as storage and detection during the continued routing process of each model service instance.

[0106] Based on this, after receiving the model call request forwarded by the grid proxy, the target service instance can query the model loading status registry. After the query, the following situations may occur:

[0107] In one scenario, the target service instance queries the model loading status registry and finds that it has loaded the target model required by the model call request. In this case, the target service instance can respond to the model call request using itself, that is, using the target model it has loaded. Referring to Figure 3, if the model call request requires model A, and the model call request is routed to model service instance a by the grid proxy, then model service instance a in Figure 3 becomes the target service instance. In this case, model service instance a has already loaded model A, so model service instance a can respond to the model call request using itself.

[0108] In another scenario, the target service instance queries the model loading status registry to find that it does not have the target model required by the model call request. In this case, the target service instance can route the model call request to any model service instance that has already loaded the required target model. Specifically, the target service instance can query the model loading status registry to find one or more model service instances that have already loaded the target model and route the model call request to one of the found model service instances. Here, the target service instance can select the model service instance to be ultimately routed from the found model service instances using a load balancing mechanism or a random mechanism, which is not limited here. In this way, even if the target service instance does not have the target model loaded, it can still route the model call request to another model service instance that has already loaded the target model. Referring to Figure 3, if the model call request requires model A, since both model service instances a and b support providing model A, the model call request may be routed by the grid proxy to model service instance b. In other words, model service instance b in Figure 3 becomes the target service instance. In this case, model A is not loaded on model service instance b, but model service instance b can learn that model A has been loaded on model service instance a by querying the model loading status registry. Then model service instance b can continue to route the model call request to model service instance a, and finally, model service instance a will respond to the model call request.

[0109] In another scenario, the target service instance queries the model loading status registry and finds no model service instance that has the required target model loaded. In this case, the target model is not loaded on any model service instance. To address this issue, the target service instance can select any model service instance that supports loading the required target model as the designated service instance and route the model call request to that designated service instance. To ensure that the target service instance can successfully select the designated service instance, static attribute tags for each model service instance can be stored and managed in the model loading status registry. In practical applications, these static attribute tags can be obtained from the various runtime tag data mentioned above. This allows the target service instance to accurately select the designated service instance that supports loading the required target model when querying the model loading status registry. Furthermore, since the selected designated service instance does not have the target model loaded, the target service instance also transmits a model loading request to the model service elastic scaler deployed in the service grid, instructing the model service elastic scaler to control the designated service instance to load the target model.

[0110] Referring to Figure 2 , the model service elastic scaler 103 manages the number of models loaded, as well as the loading time and location. Upon receiving a model load request from a target service instance, the model service elastic scaler 103 parses the request to determine which model to load and which model service instance to load it to. Furthermore, the model service elastic scaler 103 controls the designated model service instance to load the target model.

[0111] In addition to controlling model loading according to model loading requests, the model service elastic scaler 103 can also execute at least the following control logic:

[0112] 1. Allocate at least two model service instances for recently used models. In other words, ensure that all recently used models are loaded with at least two copies.

[0113] 2. Scale the number of model service instances for models whose load exceeds a specified threshold. In other words, if the request load for a particular model exceeds a certain threshold, the number of model replicas can be automatically scaled. Furthermore, if the load is high enough, the model can be replicated in every model service instance.

[0114] 3. When a loaded model is released or there is free loading space, actively control the loading of unloaded models. In other words, when there is free space or to replace a loaded model that has not been used for a long time, actively load the currently unloaded model.

[0115] 4. Based on the load heuristic algorithm and the last usage time, the number of replicas of models with more than 2 replicas is gradually reduced to 2, and from 2 to 1 in a timely manner. This control logic can be implemented preferentially on the model service instance with the highest load.

[0116] Of course, the control logic described above is merely illustrative. The model service elastic scaler 103 can also perform other control logic, which are not further illustrated here. Thus, through these automatic management mechanisms, the number of model replicas and their loading locations can be more intelligently controlled to optimize inference performance and resource utilization.

[0117] In this way, in this preferred continued routing mechanism, the target service instance can use the model loading status registry deployed in the service grid to perceive the model loading status in each model service model, and continue to route the model call request based on the model loading status to ensure that the model call request is routed to the model service instance that can provide the required target model.

[0118] In other optional continued routing mechanisms, model service instances can also synchronize model loading status with each other. In this continued routing mechanism, each model service instance can independently maintain the global information of model loading status. When continued routing is required, the target service instance can query the global information of model loading status maintained by itself to complete continued routing.

[0119] It should be understood that the above-mentioned several continuing routing mechanisms are merely exemplary, and this embodiment is not limited thereto, and no further examples are given here.

[0120] Accordingly, in this embodiment, after receiving a model call request, the target service instance can perceive the target model required for the model call request through the parsed model description information, and can also perceive the real-time model loading status on each model service instance. Through these two aspects of perception, it can ensure that the model call request is routed to the model service instance that can provide the target model, thereby ensuring that the model call request can be responded to correctly.

[0121] Referring to Figure 3, continuing with the exemplary deployment scheme within the model service instance described above, the routing operation is performed by the model-specific agent within the model service instance. In conjunction with the routing operation performed by the grid agent in this embodiment, referring to Figure 3, the model-specific agents of multiple model service instances can communicate with each other to support the implementation of the routing operation. Based on this, an exemplary routing process for a model call request may generally include:

[0122] 1. After receiving a model call request, the load balancer in the grid proxy routes the request to an available target service instance that has been filtered by the dynamic routing policy.

[0123] 2. In the target service instance, the model-specific agent can determine the required target model by parsing the model description information such as the model ID in the request metadata header.

[0124] 3. The model-specific agent in the target service instance can check whether the target model has been loaded in the model server it follows by querying the model loading status registry, etc. If it has been loaded, it triggers the model server it follows to respond to the model call request.

[0125] 4. If the target model is not loaded in the model server it follows, the model-specific proxy in the target service instance can route the model call request to the model-specific proxy in other model service instances that can provide the target model. The model-specific proxy will continue to forward the model call request to the model server it follows, so that the model call request can smoothly reach the model server that can provide the target model.

[0126] This routing process ensures that model call requests can be effectively routed to the model service instance holding the target model in one routing process, without the need for process rerouting. This can effectively improve the routing hit rate and ensure efficient model service.

[0127] From the technical concepts or optional implementations provided in the above embodiments, it can be seen that the model service solution provided by the present disclosure can produce the following technical improvements:

[0128] This disclosure proposes a method and system for dynamic loading and servitization of multiple models for the management and routing of model services, providing processing capabilities for high-scale, high-density, and frequently changing model usage scenarios.

[0129] The present disclosure provides a method and system for generating a dynamic subset routing strategy from static attribute tags, selecting the routing range of a target service instance when converging requests, and improving the request hit rate.

[0130] This disclosure provides a routing mechanism based on a combination of a grid proxy and a model-specific proxy to handle and manage model service-related traffic routing. The grid proxy serves as the primary routing component, while the model-specific proxy, a secondary routing component, builds on this. Based on model description information such as the model name in the request, it queries the model loading status registry to route requests to the corresponding model service instance. This reduces the target service subset for queries and improves the first-time hit rate of routing.

[0131] This disclosure proposes a mechanism for loading models on demand during use, which effectively reduces resource usage and optimizes the use of available computing resources.

[0132] This disclosure proposes a mechanism to decouple the concerns of large-scale model servitization from service-specific reasoning technologies to support capabilities that are independent of language or model serving frameworks and support customizable reasoning service runtimes.

[0133] The present disclosure provides a mechanism for dynamically handling situations where the degree of model change is high, and provides a good plug-and-play expansion mechanism to solve the single point failure of the model service after the number of models used grows to a certain scale.

[0134] FIG4 is a flow diagram of a model servitization method provided by another exemplary embodiment of the present disclosure. The method can be performed based on a service grid. Referring to FIG4 , a grid proxy and multiple model service instances are deployed in the service grid for model services. A single model service instance is used to dynamically load models. The method may include:

[0135] Step 400: The grid proxy routes the received model call request to the target service instance;

[0136] Step 401: The target service instance parses model description information from the model call request;

[0137] Step 402: The target service instance continues to route the model call request to route the model call request to a model service instance that can provide a target model that conforms to the model description information.

[0138] In an optional embodiment, the grid proxy routing the received model call request to the target service instance may include:

[0139] The grid agent resolves static requirements for the model service instance from the model call request;

[0140] The grid agent discovers a candidate service instance that meets the static requirement from the multiple model service instances;

[0141] The grid proxy selects a target service instance from the candidate service instances to route the model call request to the target service instance.

[0142] In an optional embodiment, the static requirement includes a static attribute tag, and model service instances with the same static attribute tag are grouped into the same routing subset. The grid agent discovers a candidate service instance that meets the static requirement from the multiple model service instances, which may include:

[0143] The grid agent searches for a routing subset corresponding to a static attribute tag in the static requirement;

[0144] The model service instance contained in the found routing subset is used as the candidate service instance; the static attribute tag is used to describe the model format and other static attributes supported by the model service instance.

[0145] In an optional embodiment, multiple runtime environments are deployed for the model service in the service grid, one or more model service instances are created in a single runtime environment, and a single model service instance inherits the static attribute tags configured in the runtime environment to which it belongs.

[0146] In an optional embodiment, the grid proxy routing the received model call request to the target service instance may include:

[0147] The grid agent parses the service instance entry address from the model call request;

[0148] The grid proxy routes the model call request to the model service instance corresponding to the service instance entry address.

[0149] In an optional embodiment, a model loading status registry is deployed in the service grid, and the model loading status corresponding to each of the multiple model service instances is maintained in the model loading status registry. The target service instance continues to route the model call request to the model service instance that can provide a target model that conforms to the model description information, which may include:

[0150] The target service instance queries the model loading status registry;

[0151] If it is determined that the required target model is not loaded, the model call request is routed to any model service instance that has loaded the required target model.

[0152] In an optional embodiment, the method may further include:

[0153] If the target service instance determines that there is no model service instance that has loaded the required target model, any model service instance that can support loading the required target model is selected as the designated service instance;

[0154] Transmitting a model loading request to a model service elastic scaler deployed in the service grid to drive the model service elastic scaler to control the designated service instance to load the target model;

[0155] Routing the model call request to the specified service instance.

[0156] In an optional embodiment, the method may further include:

[0157] The model service elastic scaler allocates at least two model service instances for the most recently used model; and / or

[0158] The model service elastic scaler expands the number of model service instances for models whose load exceeds a specified threshold; and / or

[0159] The model service elastic scaler actively controls the loading of unloaded models when a loaded model is released or free loading space appears; and / or

[0160] The number of replicas corresponding to the model whose number of replicas of the model service elastic scaling device exceeds a preset threshold is gradually reduced to a target value.

[0161] In an optional embodiment, a single model service instance includes a model-specific agent and a model server, wherein the model-specific agent is used to perform the continued routing operation, and the model server is used to dynamically load multiple models.

[0162] In an optional embodiment, multiple model service instances deployed for the model service in the service grid are mapped to a unified service access interface.

[0163] It is worth noting that the technical details in each embodiment of the above-mentioned model service method can be referred to the relevant description in the aforementioned system embodiment. In order to save space, they will not be repeated here, but this should not cause any loss of the scope of protection of this disclosure.

[0164] It should be noted that the execution entities of each step of the method provided in the above embodiment can be determined based on the node devices where each component in the service grid resides. For example, a grid proxy can be hosted by node device A in the service grid cluster, while a model service instance can be hosted by a workgroup POD in the service grid cluster as the workload within the workgroup. The execution entities of each step are not described here one by one.

[0165] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The sequence numbers of the operations, such as 401 and 402, are merely used to distinguish between different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.

[0166] FIG5 is a schematic structural diagram of a distributed microservice system provided by another exemplary embodiment of the present disclosure. Referring to FIG5 , the distributed microservice system may include a first node 50 and multiple second nodes 51 .

[0167] The distributed microservices system provided in this embodiment can be applied in a cloud computing environment. Cloud computing is one of the fastest-growing trends in computer technology and involves providing managed services over a network. Cloud computing is a service delivery model that aims to provide on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services). These resources can be quickly provisioned and released with minimal management effort or interaction with the service provider. A cloud computing environment provides computing and storage resources as a service to end users, who can request the provided services for processing. A cloud computing environment may include one or more cloud computing nodes, with which local computing devices used by cloud consumers, such as personal digital assistants (PDAs) or cell phones, desktop computers, laptop computers, and / or automotive computer systems, can communicate. Nodes can communicate with each other and can be physically or virtually grouped in one or more networks (not shown), such as private, community, public, or hybrid clouds, or a combination thereof. This allows cloud computing environments to provide infrastructure, platforms, and / or software as a service without requiring cloud consumers to maintain resources on their local computing devices. Among other things, computing nodes and cloud computing environments can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0168] In this embodiment, the first node and the second node can be cloud computing nodes in a cloud computing environment. In terms of physical implementation, the first node and the second node can be computer systems, servers, or portable electronic devices such as communication devices. They can also be container groups PODs, etc., which can be used with many other general-purpose or special-purpose computing system environments or configurations. Examples of known computing systems, environments, and / or configurations that can be used with the first node and the second node include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs (personal computers), minicomputer systems, mainframe computer systems, and distributed cloud computing environments, including any of the above systems or devices.

[0169] It should be understood that although a detailed description of cloud computing is disclosed herein, the implementation of the distributed microservice system provided by this embodiment is not limited to a cloud computing environment. On the contrary, the distributed microservice system provided by this embodiment can be implemented in combination with any other type of computing environment now known or later developed.

[0170] The microservices distribution system provided in this embodiment can use a service grid as its infrastructure. A first node 50 can include a grid proxy 52 deployed for a model service in the service grid. The service grid also has multiple model service instances deployed for the model service, distributed across multiple second nodes 51. A single model service instance can be used to dynamically load a model.

[0171] On this basis, the grid proxy 52 can be used to route the received model call request to the target service instance;

[0172] The target service instance 53 can be used to parse the model description information from the model call request; and continue to route the model call request to the model service instance that can provide the target model that conforms to the model description information.

[0173] For the technical details involved in the grid agent 52 and the target service instance 53, please refer to the relevant descriptions in the aforementioned embodiments of the model service method. To save space, they will not be repeated here, but this should not cause any loss of the protection scope of this disclosure.

[0174] Accordingly, an embodiment of the present disclosure further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps performed in the above method embodiment.

[0175] Accordingly, an embodiment of the present disclosure further provides a computer program product, including a computer program, which implements the steps performed in the above method embodiment when the computer program is executed by a processor.

[0176] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0177] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0178] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0180] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0181] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0182] The foregoing is merely an embodiment of the present disclosure and is not intended to limit the present disclosure. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A model service method, wherein, A grid proxy and multiple model service instances are deployed for a model service in a service grid, and a single model service instance is used to dynamically load a model. The method is applicable to a target service instance in the multiple model service instances, and the method includes: Receiving a model call request forwarded by the grid proxy; Parsing model description information from the model call request; Continue routing the model call request to route the model call request to a model service instance that can provide a target model that conforms to the model description information; The grid proxy routes the received model call request to the target service instance.

2. The method according to claim 1, wherein, The process of the grid proxy routing the received model call request to the target service instance includes: The grid proxy parses the static demand for the model service instance from the model call request; The grid agent determines a candidate service instance that meets the static requirement from the multiple model service instances; The grid proxy selects a target service instance from the candidate service instances to route the model call request to the target service instance.

3. The method according to claim 2, wherein, The static requirement includes a static attribute tag, and the model service instances with the same static attribute tag are divided into the same routing subset. The process in which the grid agent discovers a candidate service instance that meets the static requirement from the multiple model service instances includes: The grid proxy searches for a routing subset corresponding to a static attribute tag in the static requirement; The model service instance contained in the found routing subset is used as the candidate service instance; the static attribute tag is used to describe the model format and other static attributes supported by the model service instance.

4. The method according to claim 3, wherein, Multiple runtime environments are deployed for the model service in the service grid, one or more model service instances are created in a single runtime environment, and a single model service instance inherits the static attribute tags configured in the runtime environment to which it belongs.

5. The method according to claim 1, wherein The process of the grid proxy routing the received model call request to the target service instance includes: The grid proxy parses the service instance entry address from the model call request; The grid proxy routes the model call request to the model service instance corresponding to the service instance entry address; Among them, the model service instance corresponding to the service instance entry address is used as the target service instance.

6. The method according to claim 1, wherein A model loading status registry is deployed in the service grid, and the model loading status corresponding to each of the multiple model service instances is maintained in the model loading status registry. The model call request is continuously routed to route the model call request to a model service instance that can provide a target model that conforms to the model description information, including: Query the model loading status registry; If it is determined that the target service instance has not loaded with the required target model, the model call request is routed to any model service instance that has loaded with the required target model.

7. The method according to claim 6, wherein, Also includes: If it is determined that there is no model service instance that has loaded the required target model among the target service instances, any model service instance that can support the loading of the required target model is selected as the designated service instance; A model loading request is sent to the model service elasticizer deployed in the service mesh to drive the model service elasticizer to control the designated service instance to load the target model; The model call request is routed to the designated service instance.

8. The method according to claim 7, wherein It further includes: The model service elasticizer allocates at least two model service instances for the most recently used model; and / or, The model service elasticizer expands the number of model service instances for a model whose load exceeds a specified threshold; and / or, The model service elasticizer actively controls the loading of unloaded models when the loaded model is released or there is free loading space; and / or, The number of replicas corresponding to a model whose number of replicas of the model service elasticizer exceeds a preset threshold is gradually reduced to a target value.

9. The method according to claim 1, wherein A single model service instance includes a model-specific proxy and a model server. The model-specific proxy is used to perform the continued routing operation, and the model server is used to dynamically load multiple models.

10. The method according to claim 1, wherein, Multiple model service instances deployed for the model service in the service mesh are mapped to a unified service access interface.

11. A distributed microservices system, wherein, It includes a first node and multiple second nodes. The first node includes a mesh proxy deployed for the model service in the service mesh. Multiple model service instances deployed for the model service in the service mesh are distributed on the multiple second nodes. A single model service instance is used to dynamically load models; The mesh proxy is used to route the received model call request to the target service instance; The target service instance is used to parse the model description information from the model call request; Perform continued routing on the model call request to route the model call request to a model service instance that can provide a target model that conforms to the model description information.

12. A service mesh system, wherein, It includes a mesh proxy and multiple model service instances deployed for the model service. A single model service instance is used to dynamically load models. The mesh proxy and the multiple model service instances cooperate to execute the model service method according to any one of claims 1-10.

13. A computer-readable storage medium storing computer instructions, wherein, When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the model service method according to any one of claims 1-10.

14. A computer program product, wherein, It includes a computer program that, when executed by a processor, implements the model service method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Tenant isolation method and device and electronic equipment

    CN114363254A

  • Extension method and system for agents in service grid and storage medium

    CN114579199A

  • Certificate management system, certificate management method and construction method of certificate management system

    CN114666131A

  • Traffic routing method and device based on service grid, storage medium, processor and server

    CN115865795A

  • Monitoring system, method and equipment based on service grid and storage medium

    CN116527554A