Deployment method of reasoning service, electronic device and storage medium

By automatically matching the model information configuration table and the engine information configuration table, the complexity of GPU resource allocation and inference service deployment in Kubernetes clusters is solved, and efficient model inference service deployment is achieved.

CN120631601BActive Publication Date: 2025-12-09JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511130817.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-12-09
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

In existing technologies, the allocation of GPU resources and the deployment of inference services for deep learning models in Kubernetes clusters rely on manual configuration, which leads to complex configuration, a high risk of errors, and difficulty in adapting to hardware and model updates, resulting in low deployment efficiency.

Method used

By reading the pre-saved model information configuration table and engine information configuration table, the system automatically matches resource description information, inference engine image identifier and startup parameters based on the model name, creates workload metadata, and allocates it to the target cluster nodes, thereby achieving automated deployment of the inference service.

Benefits of technology

It simplifies the configuration process of workload metadata, improves the deployment efficiency of model inference services, reduces configuration errors, adapts to hardware and model updates, and enhances deployment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631601B_ABST
    Figure CN120631601B_ABST
Patent Text Reader

Abstract

The application discloses a deployment method of an inference service, an electronic device and a storage medium. The method comprises the following steps: in response to creating the inference service, reading a model information configuration table and an engine information configuration table, and performing matching processing on the model information configuration table and the engine information table according to a model name pre-configured when the inference service is created, so as to obtain processing resources, an inference engine image identifier and inference engine start parameters that can be used to load the inference service on a container scheduling platform, and create workload metadata for carrying the inference service based on the above; and the deployment of the inference service on the container scheduling platform is realized by allocating the workload metadata to a target cluster node. Through the method, the technical problem of low deployment efficiency caused by manual configuration of parameters in workload metadata by relying on manual work in the related art is solved, and the technical effects of simplifying the deployment process of the model inference service and improving the deployment efficiency are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computers, and in particular to a reasoning service deployment method, an electronic device and a storage medium. BACKGROUND

[0002] In existing Kubernetes cluster (K8S cluster, or container scheduling platform) management, GPU (Graphics Processing Unit) resource allocation and reasoning service deployment for deep learning models mainly rely on manual configuration of multiple parameters in workload metadata. For example, when creating a reasoning service, a user needs to specify detailed GPU model, memory size and reasoning engine image. However, the above-mentioned manual configuration of parameters in workload metadata is a complex process, and is prone to configuration errors. Moreover, as GPU hardware and models are constantly updated, the code needs to be frequently modified to adapt to new devices and new models, which not only increases the deployment time of model reasoning services, but also increases the maintenance cost, resulting in the technical problem of low deployment efficiency of model reasoning services in K8S clusters. SUMMARY

[0003] The present application provides a reasoning service deployment method, an electronic device and a storage medium to at least solve the problem of low deployment efficiency of model reasoning services in K8S clusters in related technologies. According to an aspect of an embodiment of the present application, a reasoning service deployment method is provided, including: in response to the creation of a reasoning service, reading a pre-saved model information configuration table and an engine information configuration table, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one reasoning engine; based on the model name configured when the reasoning service is created, performing matching processing on the model information configuration table and the engine information configuration table to obtain resource description information, reasoning engine image identification and reasoning engine startup parameters adapted to the reasoning service; based on the resource description information, the reasoning engine image identification and the reasoning engine startup parameters, creating workload metadata for carrying the reasoning service, wherein the workload metadata is used to describe the execution strategy of a containerized application; and assigning the workload metadata to a target cluster node to deploy the reasoning service on the target cluster node.

[0004] According to another aspect of the embodiments of the present application, a deployment apparatus of an inference service is further provided, comprising: a reading unit configured to read a pre-stored model information configuration table and an engine information configuration table in response to creation of an inference service, wherein the model information configuration table comprises at least one model file, and the engine information configuration table comprises description information of at least one inference engine; a matching unit configured to perform matching processing on the model information configuration table and the engine information configuration table based on a model name configured when the inference service is created, to obtain resource description information, an inference engine image identifier and inference engine startup parameters adapted to the inference service; a first creating unit configured to create workload metadata for carrying the inference service based on the resource description information, the inference engine image identifier and the inference engine startup parameters, wherein the workload metadata is used to describe an execution strategy of a containerized application; and an allocating unit configured to allocate the workload metadata to a target cluster node, so as to deploy the inference service on the target cluster node.

[0005] According to still another aspect of the embodiments of the present application, an electronic device is further provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the steps of any of the above-described deployment methods of an inference service by using the computer program.

[0006] According to still another aspect of the embodiments of the present application, a computer readable storage medium is further provided, wherein the computer readable storage medium stores a computer program, and the computer program is configured to execute the steps of any of the above-described deployment methods of an inference service when running.

[0007] According to still another aspect of the embodiments of the present application, a computer program product or a computer program is provided, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any of the above-described deployment methods of an inference service.

[0008] According to the above embodiments provided in the application, by reading the model name in the model file applied by the user when creating the inference service, automatic matching processing is performed on the model information configuration table and the engine information configuration table, resource description information, inference engine image identifier and inference engine startup parameters that can meet the model requirements are obtained, according to the resource description information, inference engine image identifier and inference engine startup parameters, workload metadata is created, and by distributing it to the target cluster node on the container scheduling platform, the deployment of the inference service is realized. In other words, by only relying on the model name specified by the user, at least one parameter for creating workload metadata by the user can be automatically matched, which solves the low efficiency caused by manually configuring parameters in the related art to deploy the inference service, simplifies the parameter configuration process of the workload metadata, and improves the deployment efficiency of the model inference service. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application, the drawings required in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0010] Figure 1 FIG. 1 is a schematic diagram of an application scenario of a method for deploying an inference service according to an embodiment of the present application.

[0011] Figure 2 FIG. 2 is a flow diagram of an optional method for deploying an inference service according to an embodiment of the present application.

[0012] Figure 3 FIG. 3 is a schematic diagram of the overall architecture of an optional method for deploying an inference service according to an embodiment of the present application.

[0013] Figure 4 FIG. 4 is a flow diagram of an optional method for creating workload metadata according to an embodiment of the present application.

[0014] Figure 5 FIG. 5 is a schematic diagram of three optional metadata structures according to an embodiment of the present application.

[0015] Figure 6 FIG. 6 is a structural block diagram of an optional deployment device for an inference service according to an embodiment of the present application. DETAILED DESCRIPTION

[0016] With reference to the drawings and the specific embodiments described below, a better understanding of the present application can be obtained.

[0017] It should be noted that, in the description of the present application, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or further includes elements inherent in such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.

[0018] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0019] According to an aspect of the embodiments of the present application, a method for deploying a reasoning service is provided. Optionally, in the present embodiment, the above-mentioned method for deploying a reasoning service can be applied in, but not limited to, the hardware scenario as shown in Figure 1 , wherein the server device can include one or more (only one is shown in Figure 1 ) processor 102 (the processor 102 can include, but not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA) and a memory 104 for storing data, wherein the server device can further include a transmission device 106 for communication function and an input and output device 108. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned server device. For example, the server device can further include more or less components than those shown in Figure 1 , or have a different configuration from Figure 1 .

[0020] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as a computer program corresponding to the deployment method of the inference service in the embodiments of the present application. The processor 102 performs various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to a server device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0021] The transmission device 106 is used to receive or send data via a network. The specific examples of the above-mentioned network can include a wireless network provided by a communication provider of the server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF) module which is used to communicate with the Internet in a wireless manner.

[0022] The deployment method of the inference service in the embodiments of the present application can be executed by the server device, or can be executed by the server device in combination with at least one of the terminal device (which can also be understood as the input and output device 108). Wherein, the terminal device executing the deployment method of the inference service in the embodiments of the present application can also be executed by a client installed thereon.

[0023] The technical solutions in the embodiments of the present application can be but not limited to be applied in the deployment scene of processing diversified GPU cards and inference model services, and specific examples of several common application scenes are given below.

[0024] (1) Deep learning platform of cloud service provider: the technical solutions of the present application simplify the creation process of model inference services by automatically adapting different GPU resources and model files, and reduce the learning cost and technical threshold of users. Especially when facing different types and specifications of GPU cards, the cloud platform can be seamlessly connected, without the need for users to pay attention to the specific models and configurations of the GPU cards, so as to efficiently deploy and run the inference services. For example, when a user uploads a large language processing model based on a specified format and requests 80GB video memory GPU resources, the system will automatically select a compatible GPU (such as A100-PCIE-80G) and load a matching inference engine image to create and deploy the model inference service.

[0025] (2) Model inference management in high-performance computing centers: Through the cooperation between the GPU probe component and the LLM-operator (model deployment controller), the GPU resource state can be monitored in real time, and the best combination of GPU resources and inference engines can be automatically selected, thereby optimizing the performance of model inference and improving the utilization of computing resources. For example, if multiple artificial intelligence models are being tested, each model may require different sizes of video memory and specific versions of inference engines. Through the technical solution of the present application, the most suitable inference environment can be automatically created for each model, avoiding the cumbersome process of manual configuration.

[0026] (3) AI service deployment within large enterprises: By providing standardized metadata configuration and automated service creation processes, enterprise IT departments can quickly respond to business needs, reduce deployment time, and ensure service stability and high performance. For example, when deploying a model to predict user operations, only the specified model and video memory requirements need to be specified through a simple interface call, and the system can automatically complete resource matching and deployment of inference services, greatly improving the efficiency of cross-department collaboration.

[0027] The following is an example of a deployment method for inference services in the present embodiment performed by a server, Figure 2 is a flowchart of an optional deployment method for inference services according to an embodiment of the present application, as Figure 2 shown, the flow of the method can include steps S202 to S208.

[0028] Step S202, in response to the creation of an inference service, reading a pre-saved model information configuration table and an engine information configuration table, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine.

[0029] Step S204, based on the model name configured when creating the inference service, performing matching processing on the model information configuration table and the engine information configuration table to obtain resource description information, inference engine image identification, and inference engine startup parameters adapted to the inference service.

[0030] Step S206, based on the resource description information, the inference engine image identification, and the inference engine startup parameters, creating workload metadata for carrying the inference service, wherein the workload metadata is used to describe the execution strategy of a containerized application.

[0031] Step S208, assigning the workload metadata to a target cluster node to deploy the inference service on the target cluster node.

[0032] For ease of understanding, first, the professional terms involved in the embodiments of the present application are simply described.

[0033] K8S / kubernetes: an open-source container orchestration and scheduling platform, or container scheduling platform.

[0034] CMP: Cloud Manage Platform, cloud management platform, which can enable users to manage the resources of hybrid cloud and multiple data centers through a unified management platform, thereby improving work efficiency and reducing maintenance cost.

[0035] K8S master: the node in the K8S cluster where the K8S management components such as Apiserver, Kube-scheduler, Controller-manager, etc. are deployed.

[0036] K8S Node: a node in the K8S cluster that does not deploy management components and is used to run workloads.

[0037] Api-server: a module in the K8S cluster for providing external API (Application Programming Interface, Application Programming Interface) service.

[0038] List-watch: K8S unified asynchronous message processing mechanism, which can synchronize the changes of resource objects in the K8S cluster to the client in quasi-real time and ensure the reliability, sequence, performance, etc. of the message.

[0039] Service: Service provides a unified virtual IP address for a group of backend Pods (container groups). When the application needs to access this group of Pods, they can communicate through the virtual IP of Service, and the K8S cluster will automatically load balance the request to one or more instances in the group of Pods.

[0040] Among them, the containerization technology realizes the effect of "one build, run everywhere" by packaging the application and its dependent environment into a lightweight, portable and self-contained container (Container) through operating system-level virtualization.

[0041] The above deployment method of inference service can be but is not limited to being implemented on K8S cluster, and the basic processing process thereof will be simply introduced below in combination with the overall architecture diagram shown in Figure 3

[0042] ​First, a GPU detection component is deployed on at least one cluster node of the K8S cluster and runs in the form of a Daemonset on each node (node can be a server) of K8S. The detection component is mainly responsible for periodically detecting whether the node is installed with a GPU card, and sending the manufacturer, model and memory information of the GPU to the model deployment controller (LLM-Operator component); at the same time, the node is attached with a label with the GPU model, so as to facilitate application scheduling reference.

[0043] That is, by deploying a GPU detection component on the K8S cluster, the real-time GPU resource state of each node in the cluster can be obtained, and the obtained GPU resource state of each node is reported to the LLM-Operator (model deployment controller) in real time, and the model deployment controller module constructs the overall view of the GPU information of the entire K8S cluster.

[0044] It should be noted that the node (which can also be understood as a cluster node) is a working unit of the K8S cluster, responsible for running containerized applications, and each node is an independent computing resource, wherein a node can be a server, which can be a physical machine or a virtual machine, or even a cloud server.

[0045] In the process of obtaining the GPU resource information of the node through the detection component and reporting the information to the model deployment controller, the CR resource (a instance corresponding to the inference service, for example, predicting the probability of the user opening document 1) of type LLM-intance in K8S is continuously queried (List-watch). After the model deployment controller module detects the creation event of LLM-intance, the internal K8S workload construction algorithm is executed, and the key parameters (i.e. workload metadata) required for creating K8S workload are automatically inferred according to the model file applied by the user, the memory size.

[0046] Among them, the CR resource (which can also be understood as a model inference service or an inference service) created by the user includes the CR resource type kind, the type of workload workloadtype, the memory size VRAM required by the user, the model file name modelname specified by the user, and whether the user specifies the field gpuvendor of the GPU vendor, as shown in Figure 3

[0047] In this embodiment, the GPU card or GPU processor is regarded as a hardware device, and the name or identification of the GPU vendor is regarded as a device identification.

[0048] ​After the model deployment controller determines the resource description information, inference engine image identifier, and inference engine startup parameters, workload metadata is created based on these parameters. For details, refer to Figure 3 The information box shown in the automatically generated workload section includes the automatically matched GPU vendor and the two GPU cards, the model name A100, the inference engine image identifier, the inference engine image startup command command, and the startup parameters. The inference engine image identifier (which can also be understood as the network address for pulling the inference engine image) can be, but is not limited to, nvcr.io / vllm:23.09.

[0049] The so-called pulling of the inference engine image can be, but is not limited to, pulling the inference engine image matching the current inference service from the image repository according to the network address.

[0050] In the field of AI and machine learning, the inference engine image can be, but is not limited to, a Docker image or similar container image that contains a pre-installed inference engine and its dependent environment. The inference engine is software used to execute a pre-trained machine learning model for prediction or inference. It usually requires efficient mathematical operation libraries, graphics processor drivers, and specific inference frameworks (such as TensorFlow Serving, Triton Inference Server, etc.) to accelerate and optimize the inference process.

[0051] When deploying AI or machine learning models for inference services in a K8S cluster, the inference engine image plays a particularly important role. Specifically, users can select an inference engine image compatible with their own model, and then use the deployment and scheduling functions of the K8S cluster to deploy it to nodes with appropriate hardware resources (such as GPUs). In this way, the inference engine in the image can utilize these hardware resources to efficiently infer the model, while the cluster management capabilities of Kubernetes ensure high availability and scalability of the service.

[0052] Obviously, it needs to be explained that the process of obtaining the GPU resource state of each cluster node through the detection component and reporting the detected GPU resource state does not depend on whether the user creates an inference service. Only when the model deployment controller detects that the user creates an inference service, the model deployment controller will read the information in the model information configuration table, the engine information configuration table, and the model repository.

[0053] The model information configuration table can be, but is not limited to, a storage structure for associating model files with supported GPU types and corresponding inference engine images. The table is usually stored in a structured data format similar to JSON, facilitating quick lookup and parsing. Each entry in the model information configuration table represents a model or a category of models, listing the GPU types on which the model can run, the required GPU resource keys (such as nvidia.com / gpu), recommended inference engine images, and startup parameters, etc.

[0054] The engine information configuration table contains the image information of the inference engine and the default and specific parameters for starting the inference engine. Each entry represents a different version of the inference engine image, containing the default parameters, specific parameters, and necessary startup commands required when starting. This helps to automatically select the optimal inference engine configuration according to the requirements of the model and GPU resources.

[0055] It should be noted that the hardware devices in the embodiments of the present application can be, but are not limited to, GPU boards or GPU cards.

[0056] In this embodiment, the automated process from user creation of inference service request to actual running of inference service is described in detail, especially how to automatically select and configure the optimal inference engine image and startup parameters according to the characteristics of the model and GPU resources. Through the introduction of LLM-operator component and the use of model information configuration table and engine information configuration table matching processing, the entire implementation process is efficient and accurate, which can significantly reduce user manual intervention and configuration errors, while maximizing the use of GPU resources, improving the user experience of inference service and the overall performance of the system. The technical solution of the present application is especially suitable for scenarios with diverse GPU hardware device models and dynamic GPU resources, such as inference service deployment in cloud computing environments.

[0057] That is, the technical solution of the present application mainly designs a unified CRD object of type llm-instance for large model inference services using GPUs in a cloud-native way; when a user creates a large model service, only needs to create a CR resource and fill in the specified model file name and required GPU memory amount; designs a GPU detection component to detect the GPU resources existing on the cluster nodes according to the configured GPU information table and report them to the LLM-operator; designs an LLM-operator to watch the CR resource creation events of type llm-instance, automatically compares the existing GPU information in the cluster according to the model file name and memory amount applied in the CR resource, automatically calculates the required GPU card type and quantity, and completes the creation of K8S workload objects, which can shield the differences of various manufacturers' GPUs (depending on Figure 5The metadata structure shown in the design), quickly adapt to new GPU devices, simplify the user operation threshold, and have high practical value.

[0058] By reading the model name in the model file applied by the user when creating the inference service, the model information configuration table and the engine information configuration table are automatically matched, the resource description information, the inference engine image identifier and the inference engine start parameter that can meet the model demand are obtained, the working load metadata is created according to the resource description information, the inference engine image identifier and the inference engine start parameter, and the target cluster node on the container scheduling platform is allocated to realize the deployment of the inference service. In other words, at least one parameter of the user creating the working load metadata can be automatically matched by only relying on the model name specified by the user, which solves the low efficiency caused by manually configuring parameters in the related art to deploy the inference service, simplifies the parameter configuration process of the working load metadata, and improves the deployment efficiency of the model inference service.

[0059] In one exemplary embodiment, the above-mentioned response to the creation of the inference service, reading the pre-saved model information configuration table and engine information configuration table, includes: in response to the creation of the inference service, obtaining pre-configured configuration model information; based on the configuration model name and the request resource information in the configuration model information, reading the model information configuration table and the engine information configuration table, wherein the request resource information includes the request video memory quantity of the hardware device determined according to the inference service.

[0060] When the user issues a request to create an inference service (such as a large model-based inference service) through an API interface in the K8S cluster, the LLM-operator immediately responds and obtains pre-configured configuration model information related to the inference service. The configuration model information usually contains the metadata set by the user when requesting the inference service, such as the model name, the file format of the model file, the video memory size of the required GPU device, and whether the GPU manufacturer type is specified. That is, the description information of the GPU resource applied by the user provides a basis for the next resource matching and working load metadata creation.

[0061] For example, assume that the user creates an inference service named my-llm-service, requests 40GB of GPU video memory, and expects to use a GPU card provided by the first GPU manufacturer. The LLM-operator responds to this request and obtains the model name (such as a large language model), the video memory size (40GB), and the GPU manufacturer type from the request information submitted by the user.

[0062] The LLM-operator queries the model information configuration table and the engine information configuration table stored in the cluster according to the model name in the obtained configuration model information and the requested resources (mainly including the memory size and the optional GPU vendor type). The model information configuration table stores the GPU types supported by each model and the resource key, as well as the inference engine image matched therewith. The engine information configuration table records the detailed information of each inference engine image, such as the image source, the starting parameter, etc. The purpose of this process is to determine the appropriate inference engine image, GPU resource key and any specific starting parameter to meet the resource requirements and model compatibility requested by the user.

[0063] Still continuing to explain the above example, the LLM-operator uses the model name and the requested memory size of 40GB to find the model information configuration table. In the model information configuration table, an entry with the first GPU vendor name is found, and it is known that the large language model should use the vllm:23.09 version of the inference engine image on the GPU provided by the vendor, and the recommended starting parameters may include memory utilization, quantization level, tensor parallelism, etc. Then the model deployment controller checks the vllm:23.09 entry in the engine information configuration table to obtain the specific inference engine image location (which can also be understood as a website) and the starting command, to prepare for the subsequent workload creation.

[0064] The embodiment aims to extract key information from the user request and query the pre-defined configuration table based on these information to determine the GPU resource description, inference engine image and starting parameter required for creating the inference service. This process emphasizes the importance of configuration information and how the system dynamically utilizes these configuration information to automatically adapt the hardware resources, simplifies the creation process of the inference service, and improves the automation level and accuracy of resource scheduling.

[0065] In the above manner, the system can intelligently match user demand with cluster resources, not only avoiding the complexity of configuration and error rate caused by users' lack of understanding of GPU hardware details and inference engine versions, but also reducing the pressure on administrators to frequently update and maintain the system to adapt to new hardware and models, improving the resource utilization efficiency and user experience in the entire AI service ecosystem.

[0066] In an exemplary embodiment, the matching processing of the model information configuration table and the engine information configuration table based on the model name configured when the inference service is created includes: fuzzy matching of the model information configuration table and the engine information configuration table based on the model name, to obtain resource description information, inference engine image identification, and inference engine startup parameters adapted to the inference service, wherein the model information configuration table includes a mapping relationship between the model name and the device type of a hardware device, and the inference engine image identification required for running a target model corresponding to the model name, and the engine information configuration table includes the inference engine image identification and the inference engine startup parameters saved in the form of a first metadata structure.

[0067] When a user requests to create an inference service, first, based on the model name in the user request, the model information configuration table is searched or queried through a fuzzy matching algorithm to preliminarily determine the GPU type and inference engine image required for loading the model. The fuzzy algorithm allows the system to flexibly locate the closest configuration information to the model, and even if the model name is not completely consistent in the configuration table, the related GPU type and recommended inference engine image can be found. This fuzzy matching capability is crucial for handling constantly updated and changing model file names, ensuring that even in the case of model name changes, the required resources and inference engine can be correctly identified and matched.

[0068] After preliminarily determining the GPU type and inference engine image required for the model, further fuzzy matching is performed on the engine information configuration table to find the inference engine image identification matched with the model and its default startup parameters or specific startup parameters. Even if the entries in the engine information configuration table do not completely match the direct requirements of the model, the system can find the closest inference engine configuration through the algorithm to ensure that the inference service can be efficiently started and run.

[0069] It should be noted that in the embodiments of the present application, three types of metadata structures are designed as shown in Figure 5 The GPU metadata is mounted to the GPU detection component in the form of configmap, the model information (saved in the model information configuration table) and the inference engine information (saved in the engine information configuration table) are mounted to the model deployment controller component, and the data is preconfigured and initialized.

[0070] According to the model name, the model information configuration table shown in Figure 5 is fuzzy matched to obtain resource description information, which includes but is not limited to inference engine image identification (such as vllm:23.09) supported by the current model, resource name resourcekey, startup parameters, etc.

[0071] Based on the inference engine image identifier in the resource description information, the mismatch engine information configuration table is matched. If the inference engine image identifier is matched, the inference engine image is pulled from the image repository, and the inference engine startup parameters are obtained.

[0072] The inference engine is a software system responsible for executing pre-trained machine learning or deep learning models to provide inference services. This includes model loading, input data processing, model inference computation execution, output result generation, and possible model optimization and management. The inference engine can be a standalone software specifically designed for model execution.

[0073] The inference engine image is a container image that packages the inference engine and its runtime dependencies. Container images are lightweight, portable software packages that contain all the files, libraries, environment settings, and configurations needed to run an application. In modern cloud-native environments, especially in clusters managed using Kubernetes, container images are the standard form of deploying and running applications. The inference engine image contains the binary code of the inference engine, model files if pre-loaded, drivers, library files, and runtime environments specific to GPUs or other hardware accelerators.

[0074] The inference engine is a functional software responsible for model inference computation, while the inference engine image is a containerized data package that contains all dependencies and configurations required for the inference engine to run, and can be directly scheduled and run by Kubernetes or other container management systems.

[0075] The inference engine image is the specific implementation form of the inference engine in a containerized environment, which allows the inference engine to be easily deployed to different servers or nodes without the need for separate installation and configuration at each location.

[0076] In a Kubernetes cluster, when a user requests to create a model inference service, the system will select the appropriate inference engine image based on the requested GPU type and inference requirements, and start a container on a node with corresponding GPU resources, thereby providing efficient inference services. This approach improves the flexibility of deployment and the utilization of resources, simplifying the complexity of operation and maintenance.

[0077] In an exemplary embodiment, the fuzzy matching of the model information configuration table and the engine information configuration table based on the model name comprises: calling a model repository based on the model name; obtaining model metadata information of the target model from the model repository, wherein the inference service is a service in the target model; determining a target device identifier of a hardware device providing a computing resource for the inference service based on the model metadata information; determining the resource description information adapted to the inference service based on the target device identifier, the resource description information being used to describe target computing resources required for loading the target model; and determining the inference engine image identifier and the inference engine startup parameter from the engine information configuration table based on the target device identifier.

[0078] When receiving the request for creating an inference service, the first step is to initiate a call to a model repository based on the model name contained in the request to obtain the metadata information of the model. The model repository is a centralized storage area for storing model files of various models and providing an access interface to these files. By directly calling the model repository through the model name, the specific information of the model, such as the file format and the minimum hardware requirements for model running, can be efficiently obtained.

[0079] The model repository (which can also be understood as a model repository module or a model repository component on a container scheduling platform) can provide model file storage and access services to the outside in the form of NFS (Network File System), but is not limited thereto.

[0080] The model repository returns the metadata information of the target model after receiving the call request. This step is based on the exact matching of the model name to ensure that the obtained information is directly related to the inference service to be created, avoiding resource scheduling errors caused by inaccurate or incomplete model information.

[0081] According to the model metadata information, especially the size of the model and the recommended GPU memory requirement, the target device identifier (the identifier of the GPU manufacturer) of the hardware device (GPU card) in the K8S cluster that can currently provide sufficient computing resources is determined. This process involves querying the model information configuration table to find the most suitable GPU type and the corresponding resourcekey for the model.

[0082] After determining the target device identifier, the system further refines the resource description information to describe the specific computing resources required for loading the target model, such as the number and model of GPU cards. This description information is crucial for Kubernetes, as it determines how to allocate and schedule GPU resources in the cluster to meet the needs of the inference service.

[0083] Based on the target device identifier, the system retrieves the most suitable inference engine image identifier and startup parameters from the engine information configuration table. This step ensures that the selection and configuration of the inference engine match the target hardware device, thereby optimizing the performance and efficiency of model inference.

[0084] This embodiment aims to call the model repository by model name, obtain metadata information, and then determine the target device identifier and resource description information, and finally match the best inference engine image and parameters. In this way, the creation of the inference service is not only based on the accurate identification of the model name, but also fully considers the actual running requirements of the model and the hardware resource status within the cluster, optimizes the allocation and utilization of resources, reduces the operation complexity, and improves the accuracy and efficiency of resource scheduling.

[0085] In an exemplary embodiment, the above-mentioned determination of the target device identifier of the hardware device providing computing resources for the inference service based on the model metadata information includes: based on the model file format and the model name in the model metadata information, querying the model information configuration table to obtain a hardware resource list currently available for loading the inference service in the container scheduling platform; and determining the target device identifier from at least one device identifier included in the hardware resource list.

[0086] When a user requests to create an inference service, the system first queries the model information configuration table according to the model file format and the model name in the model metadata. As shown in Figure 5 The model information configuration table is a file in the form of a pre-defined metadata structure, which records in detail the specific requirements of different models for GPU hardware, including supported GPU manufacturers, video memory size, and recommended resourcekey and inference engine image. This step is the basis of the entire resource matching process, ensuring that the system can accurately identify the model's preference for hardware and determine the candidate GPU resource list based on it.

[0087] After determining the preliminary screening of the hardware resource list, the target device identifier most suitable for the current inference service requirement is selected from the hardware resource list. This decision-making process considers the user's specific requirements (such as whether the GPU manufacturer is specified), the current available state of the GPU resource, and the model's requirements for GPU performance. The system analyzes the characteristics of each GPU in the hardware resource list, including its manufacturer, video memory size, and current utilization, to determine which GPU is the best choice, thereby determining the target device identifier.

[0088] Through the query of the model information configuration table, based on the model file format and the model name, the hardware resource list currently available for loading the inference service in the container scheduling platform (which can also be understood as a K8S cluster) is efficiently screened out. Then the best target device identifier, i.e., the GPU vendor, is selected from the list for matching the computing resources of the inference service.

[0089] That is, through deep analysis of the model metadata information and efficient query of the model information configuration table, the precise matching of the inference service and the GPU resources in the container scheduling platform is realized. This scheme automatically identifies the file format and name of the model, generates a compatible hardware resource list, and optimally selects the target device identifier, i.e., the most suitable GPU vendor and model, from the list, thereby improving the intelligence and efficiency of resource scheduling. Moreover, users do not need to deeply understand the details of the GPU, and can create, load and use high-speed and stable inference services. At the same time, the utilization rate of system resources is optimized, the operation and maintenance cost is reduced, and strong technical support is provided for the rapid deployment and efficient operation of AI application scenarios.

[0090] In an exemplary embodiment, the determination of the target device identifier from the at least one device identifier contained in the hardware resource list comprises: in the case that the first field in the model metadata information is not configured with a default device identifier, determining the target device identifier from the at least one device identifier according to a preset priority.

[0091] By querying the first field in the model metadata information, it is determined whether the user configures a default device identifier when creating the inference service. For details, please refer to the user creation of CR resources in Figure 3 The first field can be but is not limited to gpuvendor. If the field is configured with a device identifier, it is determined that the user specifies a GPU vendor; otherwise, the user does not specify a GPU vendor.

[0092] In the case that the user does not specify a GPU vendor according to the first field, the optimal GPU vendor can be but is not limited to being automatically selected according to a preset priority. The setting standard of the priority can be but is not limited to the performance of the GPU, the sorting of each device identifier in the hardware resource list, the usage frequency of each device identifier, etc.

[0093] In this embodiment, by introducing the mechanism of dynamically selecting the GPU vendor, when the user does not explicitly specify a default device identifier in the model metadata information requested by the user, the system automatically selects the optimal target from the compatible GPU vendors according to the preset priority order to load the inference service.

[0094] The above screening mechanism not only enhances the flexibility of the system, but also improves the efficiency of resource allocation, especially in a multi-GPU hardware environment, it can ensure that the system can still complete the resource scheduling intelligently and quickly when the user has no preference or the preference is ambiguous, and guarantee the high-performance operation of the inference service. In addition, the dynamic optimization mechanism reduces the user's dependence on GPU professional knowledge and simplifies the process of creating an inference service.

[0095] In other words, in the case where the user does not specify the GPU manufacturer, the GPU resources are automatically optimized by default priority, which improves the intelligence and flexibility of the container scheduling platform. This mechanism ensures that the inference service can automatically match the GPU manufacturer with the best performance and the most suitable resources according to the model characteristics, and can achieve efficient deployment even in a variable hardware environment.

[0096] In an exemplary embodiment, the above-mentioned determination of the target device identifier from the at least one device identifier according to the default priority includes at least one of the following: determining the default arrangement order of the at least one device identifier as the default priority, and determining the device identifier with the highest priority as the target device identifier; configuring the default priority of the at least one device identifier based on the usage frequency of each device identifier in the at least one device identifier, wherein the usage frequency is proportional to the priority of the at least one device identifier; and determining the target device identifier with the highest priority from the at least one device identifier according to the default priority.

[0097] In this embodiment, the priority of at least one device identifier in the hardware resource list can be set in two ways, for example, but not limited to. The first way is to set the default arrangement order of the GPU manufacturer in the hardware resource list as the default priority. According to this default order, the GPU manufacturer with the highest priority (the first in the order) is selected as the target device identifier to load the inference service.

[0098] For example, assuming that the GPU manufacturers defined in the model information configuration table are arranged in the order of the first manufacturer, the second manufacturer, and the third manufacturer, when the user does not explicitly specify the GPU manufacturer, the system will first try to match the GPU resources of the first manufacturer as the target device identifier. If the GPU resources of the first manufacturer do not meet the requirements, the system will check the GPU resources of the second manufacturer and the third manufacturer in order until the most suitable GPU resources are determined.

[0099] The second way is to configure a preset priority based on the usage frequency of the GPU vendor in the historical inference service. In this mechanism, the GPU vendor with higher usage frequency is given higher priority. Therefore, when the GPU vendor is not specified in the user request, the system automatically selects the GPU vendor with the highest usage frequency and the highest priority as the target device identifier according to the usage frequency priority of the GPU vendor.

[0100] The above-mentioned first static default arrangement order and the second dynamic usage frequency priority setting method can be flexibly switched according to the actual situation, ensuring that the inference service can always run on the optimal GPU resource.

[0101] Among them, the default sorting mechanism is suitable for scenarios lacking historical data or with relatively fixed user preferences. It directly points to the preferred GPU vendor through the predefined order, simplifying the decision-making process. The priority setting method based on usage frequency can dynamically adjust the priority based on actual usage, preferentially selecting the GPU resource that performs outstandingly and frequently in historical data, further improving the efficiency of resource scheduling and the quality of inference service. The combination of the two mechanisms not only considers the agility of the system, but also enhances its adaptability and intelligence. It has significant technical effects on optimizing GPU resource scheduling, improving inference service performance, and reducing user intervention.

[0102] In an exemplary embodiment, the above-mentioned determination of the target device identifier from at least one device identifier contained in the hardware resource list further comprises: in the case that the first field in the model metadata information is configured with a default device identifier, determining the default device identifier as the target device identifier.

[0103] As can be known from the above description in the embodiments, first, it is determined according to the first field whether the user has configured a default GPU vendor when creating the inference service. If the first field is empty, it means that the user has not specified a GPU vendor.

[0104] If the default device identifier (i.e., the user-specified GPU vendor) has been configured, there is no need for further GPU vendor screening or dynamic adjustment. Instead, the default device identifier is directly taken as the target device identifier, and the inference engine image ID configured by the model in the specified vendor is directly filtered out according to the GPU vendor indicated by the target device identifier.

[0105] The above-mentioned mechanism focuses on the personalized needs of users, ensuring that the inference service can run on the GPU resource preferred by the user, thereby providing the best performance and meeting specific computing needs. At the same time, since the user-specified GPU vendor is directly matched, the system avoids additional resource exploration and matching process, further reducing the delay of service deployment and improving the overall response speed.

[0106] By default device identification in model metadata information, the user's preference for a specific GPU vendor is directly met, realizing efficient and personalized resource scheduling. The complex selection and matching process is avoided, the delay of inference service creation is reduced, and the user experience is enhanced. At the same time, since the user-specified GPU resources are accurately matched, the inference service can exert optimal performance at runtime, meeting the efficiency and stability in high-computing-demand scenarios, and embodying the system's ability to quickly respond to user needs and optimize resource allocation.

[0107] In an exemplary embodiment, based on the target device identification, the resource description information adapted to the inference service is determined; the pre-set request memory quantity and request computing resource are obtained; based on the target device identification, the device model list of the target hardware device currently available in the container scheduling platform detected by the detection component is obtained, wherein the target device model of the target hardware device corresponds to the target device identification; in the case that the remaining memory quantity of the first hardware device of the first device model in the device model list is greater than or equal to the request memory quantity, the target computing resource is determined based on the remaining memory quantity and the request computing resource.

[0108] After determining the target GPU vendor according to the method in the above embodiment, the VRAM memory size applied by the user is further compared with the GPU information reported by the GPU detection component, and the GPU model is determined according to the comparison result.

[0109] Specifically, when the user initiates an inference service creation request, the system first parses the model metadata information to extract the required request memory quantity and computing resource demand. This data is the basis for the subsequent GPU resource matching process, ensuring that the inference service can be created with the required GPU configuration.

[0110] According to the determined target device identification, the information reported by the GPU detection component is queried to obtain a list of all available GPU device models matching the target device identification in the current container scheduling platform. This step ensures that the system can know which GPU resources can be used to meet a specific user request.

[0111] After obtaining the model list of the target hardware device, the system further compares the remaining memory size or memory quantity of each GPU model, such as finding the first GPU model whose remaining memory quantity is greater than or equal to the user request memory quantity. After determining the GPU model that meets the condition, the target computing resource required to create the inference service can be accurately calculated based on this model and its remaining computing resources.

[0112] Through the above steps, the GPU resource accurate matching based on the target device identification is realized. First, the system captures the user's request for the amount of video memory and computing resources through the model metadata information, ensuring that the subsequent matching is accurate; then according to the real-time information detected by the GPU detection component, the current available GPU model list matching the user's request device identification is accurately obtained; finally, by comparing the remaining video memory of the GPU model, the first GPU model meeting the user's video memory requirement is located, and the target computing resource required for creating the inference service is calculated.

[0113] Obviously, the target device identification can include multiple GPU models, that is, a GPU manufacturer can provide multiple GPU boards of different models.

[0114] By obtaining the user's request for the amount of video memory and computing resource demand, and accurately matching the GPU resource with the target device identification, efficient use of resources during the creation of the inference service is ensured. This mechanism can intelligently filter out GPU models that meet the user's video memory requirements and computing power, effectively reducing resource waste and speeding up the deployment of model inference services.

[0115] In an exemplary embodiment, the above method further comprises: in the case where the remaining video memory is less than the requested video memory, determining a second device model from the device model list; in the case where the remaining video memory of the second hardware device of the second device model is greater than or equal to the requested video memory, determining the target computing resource based on the remaining video memory and the requested computing resource.

[0116] First, check whether the remaining video memory of the currently available GPU meets the user's request for the amount of video memory. If it is found that the remaining video memory of a certain GPU model (first device model) is lower than the user's request, it indicates that the current GPU may not be sufficient to independently support the inference service to be created.

[0117] That is, search for the next GPU model (second device model) in the device model list that can meet the video memory requirement. This step reflects the flexibility and resource allocation ability of the system, which can quickly find alternative solutions when the first choice fails.

[0118] After determining the second device model, it is verified again whether the remaining video memory is greater than or equal to the user's request for the amount of video memory, and based on this and the user's request for the computing resource, the final target computing resource is determined. This process ensures that even in the case of initial selection failure, the system can still provide sufficient GPU resources for the inference service, avoiding the performance bottleneck of the inference service caused by insufficient resources.

[0119] For example, assume that the inference service requested by the user to create requires 20GB of video memory, but the GPU model provided by the first GPU vendor is currently only left with 18GB of video memory for the A100-PCIE-40G GPU, and the system will determine that the remaining video memory is insufficient.

[0120] In this case, other GPU board cards of different models are detected, for example, a GPU board card of the L20-PCIE-48G model is detected, until a GPU model with at least 20GB of remaining video memory is found.

[0121] In the case where it is determined that the GPU board card of the L20-PCIE-48G model has sufficient remaining video memory, the number of GPU board cards required to create the inference service is calculated according to its characteristics and the user's request for computing resources, and the target computing resources are determined.

[0122] In the above manner, in the case where the remaining video memory is insufficient, the replacement GPU model is intelligently determined, the resources are effectively supplemented and optimized, and even if the initial resource allocation does not meet the requirements, the GPU resources that meet the conditions can be quickly adjusted and found, avoiding the risk of delay and performance degradation of service creation.

[0123] In an exemplary embodiment, the determining, based on the target device identifier, the inference engine image identifier and the inference engine startup parameter from the engine information configuration table includes: querying a second field from the model information configuration table based on the target device identifier; in the case where the second field is found from the engine information configuration table, determining an inference engine containing the second field as a target inference engine; and obtaining the inference engine image identifier and the inference engine startup parameter based on engine description information of the target inference engine.

[0124] When the user requests to create an inference service, first, a second field suitable for the target device (i.e., the GPU vendor and model specified by the user) is queried from the model information configuration table. The second field here refers to the model file format and related metadata compatible with a specific GPU device, which is an important basis for selecting an inference engine.

[0125] The second field may, but is not limited to, be vllm:23.09 as shown in the table. Figure 5 According to the second field, the engine information configuration table is matched, and if the same field is matched from the engine description information in the engine information configuration table, the target inference engine is determined, and the inference engine image identifier and the inference engine startup parameter in the target inference engine are obtained, which are used to create the workload metadata.

[0126] In the foregoing manner, automatic selection and parameter configuration of the inference engine based on the target device identifier are achieved. First, the model file format and metadata information are determined through the model information configuration table, which is the basis for selecting the inference engine. Then, the inference engine image compatible with a specific format is located in the engine information configuration table, ensuring the matching of the inference engine and the target device. Finally, the key parameters are extracted from the inference engine description information, completing the preparation work for creating the workload metadata.

[0127] In the foregoing manner, not only is the parameter determination process for creating workload metadata by the user simplified, but also the efficiency of creating workload metadata is improved. By automatically selecting the most suitable inference engine image and starting parameters, the system can quickly respond to user needs, accelerate the deployment of inference services, ensure high-quality operation of the services, and improve the flexibility of the entire container scheduling platform.

[0128] In an exemplary embodiment, the foregoing method further includes: in the case where the second field is not found from the engine information configuration table, determining a default inference engine as the target inference engine; and based on the engine description information of the default inference engine, obtaining the inference engine image identifier and the inference engine starting parameters.

[0129] When the field corresponding to the model file format requested by the user (e.g., the second field) cannot be found in the engine information configuration table, or for some reason the suitable inference engine cannot be determined, the default inference engine mechanism will be enabled.

[0130] Specifically, reference can be made to the default engine shown in the engine information configuration table in FIG. 8. Figure 5 The system will read the inference engine image identifier and the engine starting parameters from the description information of the default engine as the basis for creating the inference service workload. Through this process, it is ensured that even in the absence of specific configurations, the deployment of inference services can be achieved using the default configured inference engine.

[0131] In the foregoing manner, it can be ensured that in the case where the target inference engine is not matched, the service creation process can still proceed smoothly through the default inference engine mechanism. Delay or failure of the deployment of inference services caused by the absence of model or GPU description information is avoided, and the flexibility of the system and the service experience of the user are improved. In other words, through the pre-set default inference engine, the system can automatically select and configure the inference engine, ensuring that even in the absence of specific configurations, the model inference service can still be efficiently and reliably deployed and run.

[0132] In order to more clearly understand the overall implementation process of determining the creation of workload metadata, the following will be combined with FIG. 8. Figure 4The flowchart is further described as shown.

[0133] S402, according to the model name, calling the model warehouse, obtaining the model metadata information.

[0134] Among them, the model metadata information includes but is not limited to model file format, model name, etc.

[0135] S404, according to the model name and the model file format, querying the model information configuration table, matching the GPU list supported by the model, and the resourcekey and inference engine image name corresponding to each GPU.

[0136] Among them, the GPU list includes but is not limited to the device identification composed of the GPU manufacturers supported by the model, because the GPU manufacturers supported by different model inference services may be different, and the user may also specify the GPU manufacturer when creating the model inference service, therefore, the GPU manufacturer needs to be determined first.

[0137] S406, judge whether the user specifies the GPU manufacturer.

[0138] Specifically, through the user creates CR resource part as shown in Figure 3 , the value on the gpuvendor field is used to determine whether the user specifies the GPU manufacturer.

[0139] If specified, step S408 is executed; otherwise, jump to step S414.

[0140] S408, if the user specifies the GPU manufacturer, the inference engine image ID (inference engine image identifier) configured by the model in the specified manufacturer is directly filtered out.

[0141] S410, according to the determined manufacturer, querying the engine information configuration table to obtain the start parameters of the inference engine.

[0142] Among them, the start parameters include but are not limited to default start parameters and specific start parameters.

[0143] S412, based on the inference engine image identifier and the inference engine start parameters, constructing workload metadata, thereby creating a workload.

[0144] S414, according to the manufacturer order in the model information configuration table, selecting the first manufacturer from top to bottom;

[0145] S416, according to the VRAM video memory size applied by the user, and the GPU information reported by the GPU detection component.

[0146] S418, obtaining the GPU model list available for A manufacturer in K8S cluster;

[0147] S420, determine whether the remaining available GPU memory size of Model 1 is greater than the VRAM memory size applied by the user.

[0148] If yes, execute step S410; otherwise, execute step S422.

[0149] S422, query whether there are other GPU boards of other models provided by A manufacturer in the cluster.

[0150] If yes, execute step S418; otherwise, jump to step S414 to select other GPU manufacturers.

[0151] In an exemplary embodiment, the above-mentioned creating, based on the resource description information, the inference engine image identifier, and the inference engine startup parameter, workload metadata for carrying the inference service includes: automatically packaging the resource description information, the inference engine image identifier, and the inference engine startup parameter in a preset format to obtain the workload metadata.

[0152] After confirming the target GPU device and the corresponding inference engine, an automatic packaging method is adopted to format the obtained resource description information (including the model, number, and memory size of the GPU board), inference engine image identifier, and inference engine startup parameter, and integrate them into workload metadata that is easy for Kubernetes to understand and execute. This packaging process ensures the integrity, consistency, and accuracy of the information.

[0153] The form of the packaged workload metadata can be, but is not limited to, as shown in Figure 3 , including the type kind of the workload, the GPU manufacturer and the number, the model name A100, the inference engine image identifier (which can also be understood as the network address of the inference engine image) vllm:23.09, the startup command command of the inference engine, and the startup parameters, etc.

[0154] By automatically packaging the resource description information, the inference engine image identifier, and the startup parameter, standardized workload metadata is formed, simplifying the complexity of resource management and inference service deployment. Automatic packaging ensures accurate information transmission and efficient scheduling of Kubernetes, reduces manual intervention, avoids configuration errors, and speeds up service online.

[0155] In this way, the user can focus more on business logic and model optimization without worrying about the details of the underlying infrastructure, improving the efficiency of deploying model inference services on a container scheduling platform, simplifying the user operation process, and optimizing resource scheduling.

[0156] In an example embodiment, before reading the pre-stored model information configuration table and engine information configuration table in response to the creation of the inference service, the above method further comprises: setting at least one probe component on each cluster node in the container scheduling platform; reading node resource information on the each cluster node detected by the at least one probe component at a preset time interval; and generating a resource information configuration table based on the node resource information.

[0157] At least one GPU probe component is deployed on each node in the container scheduling platform, i.e., the Kubernetes cluster. GPU resource information on the node is read and reported by the probe component at a preset time interval. This information includes the model, number, and memory size of the GPU, which forms the basis for resource information updates.

[0158] For example, the probe component executes an LSPCI command every 5 minutes to obtain PCI device information on the node. Then, according to the GPU information configuration table, the detailed specifications of each GPU card are parsed, including PCI-ID, VRAM size, etc. These data are then encapsulated and sent to the LLM-operator through the REST interface, ensuring the real-time and accuracy of the resource information.

[0159] The system summarizes the collected node resource information to generate or update the resource information configuration table, which records the latest status of all GPU resources in the cluster, including manufacturer, model, memory size, and available quantity, providing detailed data basis for the selection of inference engines and the scheduling of workloads.

[0160] By setting the probe component on each cluster node, continuous monitoring of node resource information and dynamic updating of the resource information configuration table are achieved. This ensures that the system always has the latest GPU resource status, enhancing the adaptability of the container scheduling platform to heterogeneous environments and the flexibility of resource scheduling. Through real-time monitoring and configuration table updating, the system can more accurately match GPU resources with inference service requirements, avoiding resource waste and improving computing efficiency and user satisfaction.

[0161] In an example embodiment, reading the node resource information on the each cluster node detected by the at least one probe component at a preset time interval comprises: executing a target probe command at a preset time interval; in response to the target probe command, probing the cluster node with the probe component to obtain the node resource information, and storing the node resource information in the resource information configuration table in the form of a second metadata structure.

[0162] As described in the above embodiments, the detection component is assumed to execute the LSPCI command once at a preset time interval to obtain the PCI device information on the cluster nodes. Then, according to the GPU information configuration table, the detailed specifications of each GPU are parsed, including PCI-ID, VRAM size, etc. These data are then packaged and sent to the LLM-operator through the REST interface, ensuring the real-time and accuracy of the resource information.

[0163] After detecting the node resource information of each cluster node, the node resource information is saved to the GPU information configuration table in the form of a second metadata structure as shown in Figure 5 , which is a standardized data format for the system to understand and process.

[0164] The second metadata structure not only includes the model, total number and memory size of the GPU, but also contains the available memory of each model of GPU, so that the resource information configuration table can more comprehensively reflect the GPU resource state of the cluster.

[0165] By periodically executing the detection command, the GPU resource information of the cluster nodes is collected and updated, ensuring the real-time and accuracy of the resource information configuration table. This dynamic monitoring and updating mechanism enables the system to respond to changes in GPU resources in a timely manner, such as fluctuations in memory usage, increases or decreases in GPU models, etc., so that more intelligent and accurate resource scheduling can be achieved during the inference service creation process, avoiding resource waste and improving computing efficiency, providing users with a more stable and efficient service experience. At the same time, the standardized second metadata structure simplifies the information processing process, enhancing the maintainability and scalability of the system.

[0166] In an exemplary embodiment, the above method further comprises: reporting the node resource information to a model deployment controller in real time, wherein the node resource information includes device identification, device model and memory information of hardware devices providing computing resources. The container scheduling platform is deployed with at least one of the detection component, the model deployment controller and the model repository.

[0167] As shown in Figure 3 , after detecting the GPU resource information of the cluster nodes through the detection component, the information is immediately reported to the model deployment controller through a preset communication mechanism. The reported information includes but is not limited to the identification (PCI-ID), model and memory information of the GPU device and other key data.

[0168] For example, taking a Kubernetes cluster as an example, assume that the probe component detects that 2 GPUs of the target model are added to one of the cluster nodes, each with 80GB of memory. The probe component generates node resource information in JSON format in real time and reports it to the model deployment controller (LLM-operator) through the REST interface.

[0169] It should be noted that in the present embodiment, the container scheduling platform, i.e., the Kubernetes cluster, is deployed with the probe component, the model deployment controller (LLM-operator), and the model repository. These three components work together to ensure efficient creation and operation of inference services.

[0170] For example, in a K8S cluster, the probe component is responsible for monitoring GPU resources, the model repository is responsible for storing and providing model files, and the model deployment controller is responsible for analyzing user requirements, matching GPU resources, selecting inference engines, and finally creating workloads. This integrated platform design realizes real-time feedback of resource information and promotes the automation and intelligent management of inference services.

[0171] Through the probe component, node resource information is reported to the model deployment controller in real time, realizing timely feedback and integrated management of GPU resource status. This improves the accuracy and timeliness of resource scheduling, ensuring that inference services can automatically select the most suitable GPU devices and memory specifications based on the latest resource availability, thereby speeding up the deployment process of inference services and optimizing resource utilization.

[0172] In an exemplary embodiment, the above method further comprises: in the case where the target resource information exists on one of the cluster nodes is detected by the target probe component, adding a node label to the one of the cluster nodes, wherein the node label contains the device model of the hardware device in the target resource information.

[0173] As can be known from the above description of the embodiments, in the case where the target resource information exists on a specific node is detected by the target probe component, the system automatically adds a specific node label to the node. This node label contains the specific model information of the GPU device. This makes it easier for the system to identify nodes with specific GPU resources, thereby optimizing resource scheduling.

[0174] By dynamically adding node tags through the target detection component, and collaborating with the model deployment controller and other components, an intelligent and flexible resource management and scheduling platform is built. In the Kubernetes cluster, the dynamic node tagging mechanism automatically updates node identifiers based on real-time changes in GPU resource information, ensuring the model deployment controller's immediate awareness of GPU resource status. This simplifies the inference service creation process; users no longer need to manually specify the GPU model, as the system automatically identifies and allocates the most suitable resources, improving resource utilization, reducing maintenance costs, and enhancing platform scalability and user experience.

[0175] For example, suppose a user plans to create a large model inference service but lacks information about the underlying GPU hardware model and other details. Using the technical solution in this application, the user only needs to specify the required GPU memory size when creating the inference service, such as "VRAM=40GB". After receiving the application, the model deployment controller automatically selects nodes with sufficient memory and matching GPU models for scheduling based on real-time node tag information. For instance, the system may identify the node tag and automatically select the corresponding GPU card model as the optimal option based on memory size and model compatibility, thereby efficiently completing the creation and deployment of the inference service.

[0176] By automatically adding node tags containing device model information to nodes after GPU resources are identified by the detection component, dynamic association and identification of resource information are achieved. This greatly optimizes the intelligence and efficiency of resource scheduling, ensuring that the inference service can accurately match nodes with suitable GPU devices, avoiding ineffective resource searches and waste, and improving system response speed.

[0177] As can be seen from the description of the above embodiments, the main technical solutions of this application include the following points.

[0178] (1) Design of GPU detection components.

[0179] This module is deployed within the Kubernetes cluster and runs as a daemonset on each Kubernetes node. The module is mainly responsible for periodically detecting whether a GPU is installed on the node and sending the GPU's manufacturer, model, and memory information to the LLM-operator component. It also attaches a tag with the GPU model to the node for application scheduling reference.

[0180] The module detects the existing PCI device (herein assumed to be a GPU board) information on the node through an LSPCI command, obtains the GPU model and memory size information corresponding to the PCI-ID according to the GPU information configuration table, and then sends the node information and the GPU information on the node to the LLM-operator module through a rest interface for summarization, so as to construct the overall view of the GPU information of the entire K8S cluster by the LLM-operator module.

[0181] Through the GPU detection module, the GPU model and memory information on each cluster node can be detected, a special model label is added to the node, and the GPU resource information is reported in real time for the LLM-operator module to make decisions on the use of the GPU model.

[0182] (2) Design of the model warehouse module.

[0183] The module is mainly used for storing large model weight files, and provides an nfs directory to multiple workloads for mounting and accessing simultaneously in an NFS file server mode. The model files are not specially modified, for example, the model files downloaded from huggingface can be uploaded and used.

[0184] (3) Design of the LLM-operator module.

[0185] The module is designed in a cloud-native operator architecture and deployed in the K8S cluster in a deployment mode to continuously list and watch the CR resources of the type llm-instance. The module depends on two tables of model file information configuration and inference engine information configuration. When the module watches the creation event of the llm-instance, it executes the internal K8S workload construction algorithm to automatically infer the key parameters required for creating the K8S workload according to the user-applied model file and memory size, and finally creates the workload instance bearing the large model inference service.

[0186] The key attributes of the CR resource include the model name, the user-applied VRAM size, the user-specified GPU manufacturer (optional), the workload type (deployment, statefulset) and the corresponding parameters. The determination of each parameter in the workload metadata can refer to the description of the above embodiment. Figure 4

[0187] ​The LLM-operator module is deployed in the form of a stateless load within the K8S cluster, multiple replicas can be deployed, the replicas perform master selection work, and it is ensured that only one replica performs detection at the same time; internally, according to the metadata configuration, the model file applied by the user and the GPU memory size are read, the specific inference engine image is automatically matched, and the actual GPU model and number are automatically calculated, thereby simplifying the complexity of creating an inference service.

[0188] (4) GPU adaptive metadata structure design.

[0189] Mainly includes three categories of metadata as shown in Figure 5 : GPU information, model file information, and inference engine information; the metadata is stored in the form of configmap, allowing dynamic addition and update, and the LLM-operator and GPU detection module will calculate and match GPU resources according to the content set in the metadata; the data structure design of the three categories of metadata information is as shown in Figure 5 .

[0190] By designing the above three categories of metadata structure, new inference engine images, new model files, and new GPU models can be extended in a configuration form, and the entire process does not need to modify the business code.

[0191] The overall implementation process of the GPU adaptive large model inference service creation and deployment method proposed in the embodiments of the present application is as follows: using the LLM-operator of K8S and the self-designed rule template, the administrator is supported to quickly configure new model GPUs and inference engine rule resources, the user is supported to automatically parse the key information of the model file when creating an inference service, the GPU resources are automatically matched, the appropriate inference engine for different GPUs is automatically loaded, the creation of the work load is completed, and the method has high practical value. The overall implementation of the scheme will be described below with reference to the overall architecture diagram as shown in Figure 3 .

[0192] (1) Design a GPU detection component.

[0193] The main execution logic inside the GPU detection component includes the following contents.

[0194] 1) Get the PCI device list of the current node.

[0195] 2) According to the PCI device identifier, query the GPU metadata information to obtain the GPU model, number, and memory size of the current node.

[0196] 3) Attach the GPU model information to the current node in the form of a node label.

[0197] 4) Report the detected GPU model, number and memory size of the node to the LLM-operator module.

[0198] (2) Develop LLM-operator module and supporting CRD resources, the main internal logic of which includes the following.

[0199] 1) Create CRD resources.

[0200] 2) Monitor CR resource creation events.

[0201] 3) According to the field information in the CR resource, construct the necessary metadata of the workload according to the GPU adaptive inference workload construction algorithm, mainly including GPU model, resourcekey and number of layers, inference engine image, inference engine startup parameters, etc.

[0202] 4) Call K8S interface service to create workload.

[0203] (3) Design metadata structure in the form of configmap to mount GPU metadata information to the gpu detection component, model information and inference engine information to the llm-operator component, and complete data preloading and initialization.

[0204] (4) Develop model repository module to provide model file storage and access services externally in NFS mode.

[0205] (5) Complete the deployment of the above modules in the K8S cluster, wherein the GPU detection module (or detection component) is deployed in daemonset mode, the LLM-operator is deployed in deployment mode, the metadata is deployed in configmap mode, and the model repository is deployed in NFS server mode.

[0206] It should be noted that, Figure 3 The increase of node labels, vendor1-device-plugn and workload metadata shown is a K8S native resource object, except that the difference between the technical solution of the present application and the related art lies in the way of obtaining and configuring multiple parameters in the workload metadata, which mainly relies on manual configuration in the related art, and the technical solution of the present application focuses on automatic adaptation for parameter configuration.

[0207] Figure 3The user creates a CR resource, a GPU detection component, an LLM-operator (model deployment controller), and a model warehouse shown in the figure are entity components designed in the technical solution of the present application, and the GPU information configuration (or GPU information configuration table), the model information configuration (or model information configuration table), and the engine information configuration (or engine information configuration table) are data structures, i.e., metadata structures, designed in the technical solution of the present application.

[0208] Through the above embodiments provided in the present application, the model deployment controller of K8S and the self-designed rule template can be used to support administrators to quickly configure new models of GPUs and inference engine rule resources, support users to automatically parse the key information of a model file when creating an inference service, automatically match GPU resources, automatically load a version of an inference engine suitable for different GPUs, and complete the creation of a workload. Not only can the use process of the user be simplified, but also the overall efficiency and usability of the system can be improved, which has important theoretical and practical significance.

[0209] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a general hardware platform as required, of course, but also through hardware, but in many cases, the former is a better embodiment.

[0210] According to another aspect of the embodiments of the present application, a deployment device of an inference service is also provided, and a structure diagram of the system is shown in the figure. Figure 6 As shown, the device includes the following modules: a reading unit 602, configured to read a pre-stored model information configuration table and an engine information configuration table in response to the creation of an inference service, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine; a matching unit 604, configured to perform matching processing on the model information configuration table and the engine information configuration table based on a model name configured when the inference service is created, to obtain resource description information, an inference engine image identifier, and inference engine startup parameters adapted to the inference service; a first creation unit 606, configured to create workload metadata for carrying the inference service based on the resource description information, the inference engine image identifier, and the inference engine startup parameters, wherein the workload metadata is used to describe an execution strategy of a containerized application; and an allocation unit 608, configured to allocate the workload metadata to a target cluster node, to deploy the inference service on the target cluster node.

[0211] The specific execution steps involved in various computing processes and dynamic optimization of storage spaces in the above modules can refer to the description in the above embodiments, which will not be described here.

[0212] It is obvious that the picture display device described above can be used to implement the deployment method of the inference service provided in the above embodiments, which has been described and will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware, or a combination of software and hardware can also be implemented and conceived.

[0213] It should be noted that the reading unit 602 in this embodiment can be used to perform the above step S202, the matching unit 604 in this embodiment can be used to perform the above step S204, the first creating unit 606 in this embodiment can be used to perform the above step S206, and the allocating unit 608 in this embodiment can be used to perform the above step S208.

[0214] In one exemplary embodiment, the reading unit 602 described above includes: a first obtaining module, configured to obtain pre-configured configuration model information in response to creation of the inference service; and a reading module, configured to read the model information configuration table and the engine information configuration table based on a configuration model name in the configuration model information and request resource information, wherein the request resource information includes a request amount of video memory of a hardware device determined according to the inference service.

[0215] In one exemplary embodiment, the matching unit 604 described above includes: a matching module, configured to perform fuzzy matching on the model information configuration table and the engine information configuration table based on the model name to obtain resource description information, an inference engine image identifier, and inference engine startup parameters adapted to the inference service, wherein the model information configuration table includes a mapping relationship between the model name and a device type of a hardware device and the inference engine image identifier required for running a target model corresponding to the model name, and the engine information configuration table includes the inference engine image identifier and the inference engine startup parameters saved in a first meta structure.

[0216] In an example embodiment, the matching module comprises: a first calling submodule configured to call a model repository based on the model name; a first obtaining submodule configured to obtain model metadata information of the target model from the model repository, wherein the inference service is a service in the target model; a first processing submodule configured to determine a target device identifier of a hardware device providing a computing resource for the inference service based on the model metadata information; a second processing submodule configured to determine the resource description information adapted to the inference service based on the target device identifier, the resource description information being used to describe target computing resources required for loading the target model; and a third processing submodule configured to determine the inference engine image identifier and the inference engine startup parameter from the engine information configuration table based on the target device identifier.

[0217] In an example embodiment, the matching module comprises: a first querying submodule configured to query the model information configuration table based on the model file format in the model metadata information and the model name to obtain a hardware resource list currently available for loading the inference service in a container scheduling platform; and a fourth processing submodule configured to determine the target device identifier from at least one device identifier included in the hardware resource list.

[0218] In an example embodiment, the matching module further comprises: a fifth processing submodule configured to determine the target device identifier from the at least one device identifier according to a preset priority in a case where a first field in the model metadata information is not configured with a default device identifier.

[0219] In an example embodiment, the matching module comprises at least one of: a sixth processing submodule configured to determine a default arrangement order of the at least one device identifier as the preset priority, and determine a device identifier with a high ranking as the target device identifier; a first configuration submodule configured to configure the preset priority of the at least one device identifier based on a usage frequency of each device identifier in the at least one device identifier, wherein the usage frequency is directly proportional to the priority of the at least one device identifier; and a seventh processing submodule configured to determine the target device identifier with the highest priority from the at least one device identifier according to the preset priority.

[0220] In an example embodiment, the matching module comprises: a second obtaining sub-module, configured to obtain a requested memory quantity and a requested computing resource; a third obtaining sub-module, configured to obtain, based on the target device identifier, a list of device models of target hardware devices currently available in the container scheduling platform detected by the detection component, wherein a target device model of the target hardware device corresponds to the target device identifier; an eighth processing sub-module, configured to, in a case where a remaining memory quantity of a first hardware device of a first device model in the list of device models is greater than or equal to the requested memory quantity, determine the target computing resource based on the remaining memory quantity and the requested computing resource.

[0221] In an example embodiment, the matching module comprises: a ninth processing sub-module, configured to, in a case where the remaining memory quantity is less than the requested memory quantity, determine a second device model from the list of device models; and a tenth processing sub-module, configured to, in a case where a remaining memory quantity of a second hardware device of the second device model is greater than or equal to the requested memory quantity, determine the target computing resource based on the remaining memory quantity and the requested computing resource.

[0222] In an example embodiment, the matching module comprises: a second querying sub-module, configured to query a second field from the model information configuration table based on the target device identifier; an eleventh processing sub-module, configured to, in a case where the second field is found from the engine information configuration table, determine an inference engine containing the second field as a target inference engine; and a fourth obtaining sub-module, configured to obtain an inference engine image identifier and inference engine startup parameters based on engine description information of the target inference engine.

[0223] In an example embodiment, the apparatus further comprises: a first processing unit, configured to, in a case where the second field is not found from the engine information configuration table, determine a default inference engine as the target inference engine; and a first obtaining unit, configured to obtain the inference engine image identifier and the inference engine startup parameters based on engine description information of the default inference engine.

[0224] In an example embodiment, the first creating unit 606 comprises: an encapsulating module, configured to automatically encapsulate the resource description information, the inference engine image identifier, and the inference engine startup parameters in a preset format to obtain the workload metadata.

[0225] In an example embodiment, the apparatus further comprises: a setting unit configured to set at least one probe component on each cluster node in the container scheduling platform before reading the pre-stored model information configuration table and the engine information configuration table in response to the creation of the inference service; a second processing unit configured to read node resource information on the each cluster node detected by the at least one probe component according to a preset time interval; and a third processing unit configured to generate a resource information configuration table based on the node resource information.

[0226] In an example embodiment, the second processing unit comprises: an execution module configured to execute a target probe command according to a preset time interval; and a probe module configured to probe the cluster node on which the probe component is arranged in response to the target probe command, to obtain the node resource information, and to store the node resource information in the resource information configuration table in the form of a second metadata structure.

[0227] In an example embodiment, the apparatus further comprises: a reporting unit configured to report the node resource information to a model deployment controller in real time, wherein the node resource information comprises device identification, device model, and display memory information of a hardware device providing computing resources, and wherein the container scheduling platform is arranged with at least one probe component, a model deployment controller, and a model repository.

[0228] In an example embodiment, the apparatus further comprises: an adding unit configured to add a node label to one of the each cluster node in a case where target resource information exists on the one of the each cluster node detected by a target probe component, wherein the node label comprises a device model of a hardware device in the target resource information.

[0229] It should be noted that each of the above modules can be implemented by software or hardware, and for the latter, the following implementation manners can be used, but are not limited thereto: all of the above modules are located in the same processor; or the above modules are located in different processors in any combination.

[0230] According to another aspect of the embodiments of the present application, an electronic device is provided, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described inference service deployment method embodiments.

[0231] According to another aspect of the embodiments of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is configured to perform the steps in any of the above-described inference service deployment method embodiments when running.

[0232] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0233] According to still another aspect of the embodiments of the present application, a computer program product is also provided, which includes a computer program. The computer program, when executed by a processor, implements the steps in any of the above-described inference service deployment method embodiments.

[0234] The embodiments of the present application also provide another computer program product, which includes a non-volatile computer readable storage medium. The non-volatile computer readable storage medium stores a computer program. The computer program, when executed by a processor, implements the steps in any of the above-described inference service deployment method embodiments.

[0235] The skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in a general manner. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0236] The above describes in detail a method for deploying an inference service provided by the present application. The principles and implementation modes of the present application are described herein by applying specific examples. The above description of the examples is only to help understand the method of the present application and its core idea. It should be noted that for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application. These improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for deploying an inference service, comprising: in response to creation of the inference service, reading a pre-stored model information configuration table and an engine information configuration table, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine; based on a model name configured when the inference service is created, performing matching processing on the model information configuration table and the engine information configuration table to obtain resource description information, an inference engine image identifier, and inference engine startup parameters that are adapted to the inference service; based on the resource description information, the inference engine image identifier, and the inference engine startup parameters, creating workload metadata for carrying the inference service, wherein the workload metadata is used to describe an execution strategy of a containerized application; and assigning the workload metadata to a target cluster node to deploy the inference service on the target cluster node; wherein the matching processing on the model information configuration table and the engine information configuration table based on the model name configured when the inference service is created comprises: based on the model name, performing fuzzy matching on the model information configuration table and the engine information configuration table to obtain the resource description information, the inference engine image identifier, and the inference engine startup parameters that are adapted to the inference service, wherein the model information configuration table includes a mapping relationship between the model name and a device type of a hardware device, and the inference engine image identifier required for running a target model corresponding to the model name, and the engine information configuration table includes the inference engine image identifier and the inference engine startup parameters saved in a first metadata structure.

2. The method of claim 1, wherein: the reading of the pre-stored model information configuration table and the engine information configuration table in response to the creation of the inference service comprises: in response to the creation of the inference service, obtaining pre-configured configuration model information; and based on a configuration model name in the configuration model information and request resource information, reading the model information configuration table and the engine information configuration table, wherein the request resource information includes a requested GPU memory amount of a hardware device determined according to the inference service.

3. The method of claim 1, wherein: the fuzzy matching on the model information configuration table and the engine information configuration table based on the model name comprises: based on the model name, calling a model repository; obtaining model metadata information of the target model from the model repository, wherein the inference service is a service in the target model; based on the model metadata information, determining a target device identifier of the hardware device that provides a computing resource for the inference service; and based on the target device identifier, determining the resource description information that is adapted to the inference service, the resource description information being used to describe a target computing resource required for loading the target model. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ Determine the inference engine image identifier and the inference engine start parameter from the engine information configuration table based on the target device identifier.

4. The method of claim 3, wherein, The target device identifier of the hardware device providing the computing resource for the inference service is determined based on the model metadata information, comprising: Based on the model file format and the model name in the model metadata information, query the model information configuration table to obtain a hardware resource list currently available for loading the inference service in the container scheduling platform; Determine the target device identifier from at least one device identifier included in the hardware resource list.

5. The method of claim 4, wherein, The target device identifier is determined from at least one device identifier included in the hardware resource list, comprising: In the case where the first field in the model metadata information is not configured with a default device identifier, determine the target device identifier from the at least one device identifier according to a preset priority.

6. The method of claim 5, wherein, The target device identifier is determined from the at least one device identifier according to a preset priority, comprising at least one of: The default arrangement order of the at least one device identifier is determined as the preset priority, and a device identifier with a high ranking is determined as the target device identifier; The preset priority of the at least one device identifier is configured based on the usage frequency of each device identifier in the at least one device identifier, wherein the usage frequency is directly proportional to the priority of the at least one device identifier; The target device identifier with the highest priority is determined from the at least one device identifier according to the preset priority.

7. The method of claim 4, wherein, The target device identifier is determined from at least one device identifier included in the hardware resource list, further comprising: In the case where the first field in the model metadata information is configured with a default device identifier, the default device identifier is determined as the target device identifier.

8. The method of claim 3, wherein, The resource description information adapted to the inference service is determined based on the target device identifier, comprising: Obtain a requested GPU amount and a requested computing resource; Based on the target device identifier, obtain a device model list of target hardware devices currently available in the container scheduling platform detected by a detection component, wherein the target device model of the target hardware device corresponds to the target device identifier; In the case where the remaining GPU amount of the first hardware device of the first device model in the device model list is greater than or equal to the requested GPU amount, determine the target computing resource based on the remaining GPU amount and the requested computing resource.

9. The method of claim 8, wherein, The method further comprises: In the case where the remaining GPU amount is less than the requested GPU amount, determine a second device model from the device model list. In a case where the remaining GPU memory quantity of the second hardware device of the second device model is greater than or equal to the requested GPU memory quantity, the target computing resource is determined based on the remaining GPU memory quantity and the requested computing resource.

10. The method of claim 3, wherein, the determining, based on the target device identifier, the inference engine image identifier and the inference engine startup parameter from the engine information configuration table comprises: querying a second field from the model information configuration table based on the target device identifier; in a case where the second field is found from the engine information configuration table, determining an inference engine containing the second field as a target inference engine; obtaining the inference engine image identifier and the inference engine startup parameter based on engine description information of the target inference engine.

11. The method of claim 10, wherein, the method further comprises: in a case where the second field is not found from the engine information configuration table, determining a default inference engine as the target inference engine; obtaining the inference engine image identifier and the inference engine startup parameter based on engine description information of the default inference engine.

12. The method of claim 1, wherein, the creating workload metadata for carrying the inference service based on the resource description information, the inference engine image identifier and the inference engine startup parameter comprises: automatically packaging the resource description information, the inference engine image identifier and the inference engine startup parameter in a preset format to obtain the workload metadata.

13. The method of any one of claims 1 to 12, wherein, before the reading the pre-stored model information configuration table and the engine information configuration table in response to the creation of the inference service, the method further comprises: setting at least one detection component on each cluster node in a container scheduling platform; reading node resource information of the each cluster node detected by the at least one detection component at a preset time interval; generating a resource information configuration table based on the node resource information.

14. The method of claim 13, wherein, the reading the node resource information of the each cluster node detected by the at least one detection component at a preset time interval comprises: executing a target detection command at a preset time interval; in response to the target detection command, detecting a cluster node on which the detection component is set to obtain the node resource information, and storing the node resource information in the resource information configuration table in the form of a second metadata structure.

15. The method of claim 13, wherein, the method further comprises: reporting the node resource information to a model deployment controller in real time, wherein the node resource information includes device identifier, device model and GPU memory information of a hardware device providing a computing resource, and the container scheduling platform is deployed with the at least one detection component, the model deployment controller and a model storage library.

16. The method of claim 13, wherein: the method further comprises: in a case where it is detected by the target detection component that there is target resource information on one of the cluster nodes, adding a node label to the one of the cluster nodes, wherein the node label contains a device model of a hardware device in the target resource information.

17. An electronic device, comprising: comprises: a memory for storing a computer program; a processor for implementing the steps of the method for deploying an inference service according to any one of claims 1-16 when executing the computer program.

18. A computer-readable storage medium, characterized in that, a computer program stored in the computer readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the method for deploying an inference service according to any one of claims 1-16.

19. A computer program product comprising a computer program, characterized in that, the computer program, when executed by a processor, implements the steps of the method for deploying an inference service according to any one of claims 1-16.

Citation Information

Patent Citations

  • Inference service deployment method, device and equipment and storage medium

    CN111625245A

  • Model reasoning system, method and equipment

    CN114881236A