Deployment method of reasoning service, electronic equipment and storage medium
By automatically matching the model information configuration table and the engine information configuration table, the complexity of GPU resource allocation and inference service deployment in the Kubernetes cluster is solved, efficient inference service deployment is achieved, the configuration process is simplified, and deployment efficiency is improved.
Patent Information
- Application Number
- CN202511130817.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-13
AI Technical Summary
In existing technologies, GPU resource allocation and inference service deployment for deep learning models in Kubernetes clusters rely on manual configuration, which makes the configuration complex and error-prone, increases deployment time and maintenance costs, and reduces deployment efficiency.
By reading the pre-saved model information configuration table and engine information configuration table, automatically matching resource description information, inference engine image identifier and startup parameters based on the model name, creating workload metadata, and distributing it to the target cluster node, the automated deployment of the inference service is achieved.
It simplifies the configuration process of workload metadata, improves the deployment efficiency of model inference services, reduces user manual intervention and configuration errors, and improves deployment efficiency.
Smart Images

Figure CN120631601A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a method for deploying an inference service, an electronic device, and a storage medium. Background Art
[0002] In existing Kubernetes cluster (K8S cluster, or container scheduling platform) management, GPU (Graphics Processing Unit) resource allocation and inference service deployment for deep learning models primarily rely on manually configuring multiple parameters in workload metadata. For example, when creating an inference service, users must specify the detailed GPU model, video memory size, and inference engine image. However, this manual configuration of workload metadata parameters is complex and prone to configuration errors. Furthermore, with the continuous updates of GPU hardware and models, code needs to be frequently modified to adapt to new devices and models. This not only increases the deployment time of model inference services, but also increases maintenance costs, resulting in low deployment efficiency of model inference services in K8S clusters. Summary of the Invention
[0003] The present application provides a method for deploying an inference service, an electronic device, and a storage medium to at least solve the problem of low deployment efficiency of model inference services in K8S clusters in related technologies. According to one aspect of an embodiment of the present application, a method for deploying an inference service is provided, comprising: in response to the creation of an inference service, reading a pre-saved model information configuration table and an engine information configuration table, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine; based on the model name configured when creating the inference service, matching the model information configuration table and the engine information configuration table to obtain resource description information, inference engine image identifier, and inference engine startup parameters adapted to the inference service; based on the resource description information, the inference engine image identifier, and the inference engine startup parameters, creating workload metadata for carrying the inference service, wherein the workload metadata is used to describe the execution strategy of the containerized application; allocating the workload metadata to the target cluster node to deploy the inference service on the target cluster node.
[0004] According to another aspect of an embodiment of the present application, a deployment device for an inference service is also provided, including: a reading unit, for reading a pre-saved model information configuration table and an engine information configuration table in response to the creation of an inference service, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine; a matching unit, for matching the model information configuration table and the engine information configuration table based on the model name configured when creating the inference service, to obtain resource description information, inference engine image identifier and inference engine startup parameters adapted to the inference service; a first creation unit, for creating workload metadata for carrying the inference service based on the resource description information, the inference engine image identifier and the inference engine startup parameters, wherein the workload metadata is used to describe the execution strategy of the containerized application; an allocation unit, for allocating the workload metadata to a target cluster node to deploy the inference service on the target cluster node.
[0005] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the steps of any of the above-mentioned methods for deploying an inference service through the computer program.
[0006] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned methods for deploying an inference service at runtime.
[0007] According to another aspect of an embodiment of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the steps of the aforementioned method for deploying an inference service.
[0008] By adopting the above-mentioned embodiment provided by the present application, by reading the model name in the model file requested by the user when creating the inference service, the model information configuration table and the engine information configuration table are automatically matched, and resource description information, inference engine image identifier and inference engine startup parameters that can meet the model requirements are obtained. Based on the resource description information, inference engine image identifier and inference engine startup parameters, workload metadata is created, and the workload metadata is allocated to the target cluster node on the container scheduling platform to achieve the deployment of the inference service. In other words, by automatically matching at least one parameter of the workload metadata created by the user only based on the model name specified by the user, the low efficiency caused by manually configuring parameters in the related art is solved, and the parameter configuration process of the workload metadata is simplified, thereby improving the deployment efficiency of the model inference service. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0010] Figure 1 This is a schematic diagram of an application scenario of a method for deploying an inference service according to an embodiment of the present application.
[0011] Figure 2 This is the process intention of an optional method for deploying an inference service according to an embodiment of the present application.
[0012] Figure 3 This is a schematic diagram of the overall architecture of an optional method for deploying an inference service according to an embodiment of the present application.
[0013] Figure 4 This is an optional flowchart for creating workload metadata according to an embodiment of the present application.
[0014] Figure 5 This is a schematic diagram of three optional metadata structures according to an embodiment of the present application.
[0015] Figure 6 This is a structural block diagram of an optional inference service deployment device according to an embodiment of the present application. DETAILED DESCRIPTION
[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0018] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] According to one aspect of the embodiment of the present application, a method for deploying an inference service is provided. Optionally, in this embodiment, the above-mentioned method for deploying an inference service can be applied to, but is not limited to, Figure 1 In the hardware scenario shown, the server device may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a field programmable gate array FPGA and other processing devices) and a memory 104 for storing data. The server device may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0020] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the deployment method of the reasoning service in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the server device via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0021] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communication provider of the server device. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0022] The method for deploying the inference service of the embodiment of the present application can be executed by a server device, or by the server device in conjunction with at least one of the terminal devices (also understood as the input / output device 108). The method for deploying the inference service of the embodiment of the present application can also be executed by a client installed on the terminal device.
[0023] The technical solutions in the embodiments of the present application can be applied to, but are not limited to, deployment scenarios for processing diverse GPU cards and inference model services. Specific examples of several common application scenarios are given below.
[0024] (1) Deep learning platform for cloud service providers: The technical solution of this application simplifies the creation process of model inference services by automatically adapting to different GPU resources and model files, reducing users' learning costs and technical barriers. In particular, when facing GPU cards of different types and specifications, the cloud platform can seamlessly connect and efficiently deploy and run inference services without users having to pay attention to the specific model and configuration of the GPU card. For example, when a user uploads a large language processing model based on a specified format and requests GPU resources with 80GB of video memory, the system will automatically select a compatible GPU (such as A100-PCIE-80G) and load the matching inference engine image to create and deploy the model inference service.
[0025] (2) Model inference management in high-performance computing centers: Through the collaboration between the GPU detection component and the LLM-operator (model deployment controller), it is possible to monitor the GPU resource status in real time and automatically select the best combination of GPU resources and inference engines, thereby optimizing the performance of model inference and improving the utilization of computing resources. For example, assuming that multiple artificial intelligence models are being tested, each model may require a different size of video memory and a specific version of the inference engine. The technical solution of this application can automatically create the most suitable inference environment for each model, avoiding the tedious process of manual configuration.
[0026] (3) Deployment of AI services within large enterprises: By providing standardized metadata configuration and automated service creation processes, enterprise IT departments can quickly respond to business needs, reduce deployment time, and ensure service stability and high performance. For example, when deploying a model to predict user operations, simply call the specified model and graphics memory requirements through a simple interface, and the system will automatically complete resource matching and inference service deployment, significantly improving the efficiency of cross-departmental collaboration.
[0027] The following takes the deployment method of executing the inference service in this embodiment by the server as an example: Figure 2 This is a flow chart of an optional method for deploying an inference service according to an embodiment of the present application. Figure 2 As shown, the process of the method may include steps S202 to S208.
[0028] Step S202: In response to the creation of the inference service, read the pre-saved model information configuration table and engine information configuration table, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine.
[0029] Step S204: Based on the model name configured when creating the inference service, the model information configuration table and the engine information configuration table are matched to obtain resource description information, inference engine image identifier and inference engine startup parameters adapted to the inference service.
[0030] Step S206: Create workload metadata for hosting the inference service based on the resource description information, the inference engine image identifier, and the inference engine startup parameters, wherein the workload metadata is used to describe the execution strategy of the containerized application.
[0031] Step S208: distribute the workload metadata to a target cluster node to deploy the inference service on the target cluster node.
[0032] To facilitate understanding, the professional terms involved in the embodiments of this application are first briefly described.
[0033] K8S / kubernetes: an open source container orchestration and scheduling platform, or container scheduling platform.
[0034] CMP: Cloud Manage Platform, a cloud management platform that enables users to manage hybrid clouds and resources across multiple data centers through a unified management platform, thereby improving work efficiency and reducing maintenance costs.
[0035] K8S master: The node where K8S management components are deployed in the K8S cluster, such as Apiserver, Kube-scheduler, Controller-manager, etc.
[0036] K8S Node: A node that does not deploy management components in a K8S cluster and is used to run workloads.
[0037] Api-server: A module in the K8S cluster that provides external API (Application Programming Interface) services.
[0038] List-watch: K8S's unified asynchronous message processing mechanism can synchronize changes in resource objects in the K8S cluster to the client in near real time, and ensure the reliability, sequence, and performance of the messages.
[0039] Service: A Service provides a unified virtual IP address for a group of backend Pods (container groups). When applications need to access this group of Pods, they can communicate through the Service's virtual IP address, and the K8S cluster will automatically load balance the request to one or more instances in the group of Pods.
[0040] Among them, containerization technology uses operating system-level virtualization to package applications and their dependent environments into lightweight, portable, self-contained containers to achieve the effect of "build once, run everywhere".
[0041] The deployment method of the above inference service can be implemented on a K8S cluster, but is not limited to it. Figure 3 The overall architecture diagram shown briefly introduces its basic processing process.
[0042] First, deploy a GPU detection component on at least one node in the Kubernetes cluster and run it as a DaemonSet on each Kubernetes node (which can be a server). The detection component periodically detects whether a GPU is installed on the node and sends the GPU manufacturer, model, and memory information to the model deployment controller (LLM-Operator). It also tags the node with the GPU model for easy reference in application scheduling.
[0043] In other words, by deploying the GPU detection component on the K8S cluster, the GPU resource status of each node in the cluster can be obtained in real time, and the obtained GPU resource status of each node can be reported to the LLM-Operator (model deployment controller) in real time. The model deployment controller module builds an overall view of the GPU information of the entire K8S cluster.
[0044] It should be noted that a node (also known as a cluster node) is a working unit of a K8S cluster, responsible for running containerized applications, and each node is an independent computing resource. A node can be a server, which can be a physical machine, a virtual machine, or even a cloud server.
[0045] While the detection component acquires node GPU resource information and reports this information to the model deployment controller, it continuously queries (List-watch) CR resources in Kubernetes of the LLM-intance type (an instance corresponding to the inference service, for example, predicting the probability of a user opening Document 1). After the model deployment controller module detects the creation event of the LLM-intance, it executes the internal Kubernetes workload construction algorithm and automatically infers the key parameters (i.e., workload metadata) required to create the Kubernetes workload based on the model file and GPU memory size requested by the user.
[0046] Among them, the CR resources created by users (which can also be understood as model inference services or reasoning services) are as follows Figure 3 As shown, it includes the CR resource type kind, the workload type workloadtype, the user-required video memory size VRAM, the user-specified model file name modelname, and whether the user has specified the GPU vendor field gpuvendor.
[0047] In this embodiment, a GPU card or GPU processor is regarded as a hardware device, and the name or logo of a GPU manufacturer is regarded as a device logo.
[0048] After the model deployment controller determines the resource description information, inference engine image identifier, and inference engine startup parameters, workload metadata is created based on these parameters. Figure 3 The automatically generated workload information box shows the automatically matched GPU vendor, two GPU cards, the model name A100, the inference engine image ID, the inference engine image startup command, and startup parameters. The inference engine image ID (which can also be understood as the network address for pulling the inference engine image) can be, but is not limited to, nvcr.io / vllm:23.09.
[0049] The so-called pulling of the inference engine image may include, but is not limited to, pulling the inference engine image that matches the current inference service from the image repository according to the above network address.
[0050] In the field of AI and machine learning, an inference engine image can refer to, but is not limited to, a Docker image or similar container image that contains a pre-installed inference engine and its dependencies. An inference engine is software used to execute pre-trained machine learning models for prediction or inference. It typically requires efficient mathematical operation libraries, graphics processing unit (GPU) drivers, and specific inference frameworks (such as TensorFlow Serving and Triton Inference Server) to accelerate and optimize the inference process.
[0051] When deploying AI or machine learning models for inference services in a Kubernetes cluster, inference engine images are particularly useful. Specifically, users can select an inference engine image compatible with their model and then use the Kubernetes cluster's deployment and scheduling capabilities to deploy it to nodes with appropriate hardware resources (such as GPUs). The inference engine in the image can then leverage these hardware resources to efficiently infer the model, while Kubernetes' cluster management capabilities ensure high availability and scalability.
[0052] Obviously, it should be noted that the process of obtaining the GPU resource status of each cluster node through the detection component and reporting the detected GPU resource status does not depend on whether the user creates an inference service. Only when the model deployment controller detects that the user has created an inference service, the model deployment controller will read the model information configuration table, engine information configuration table and information in the model warehouse.
[0053] The model information configuration table can be, but is not limited to, a storage structure used to associate model files with supported GPU types and corresponding inference engine images. This table is typically stored in a structured data format similar to JSON, facilitating quick search and parsing. Each entry in the model information configuration table represents a model or model category, listing the GPU types that the model can run on, the required GPU resource key (such as nvidia.com / gpu), the recommended inference engine image, and startup parameters.
[0054] The engine information configuration table contains the inference engine image information and the default and specific parameters for starting the inference engine. Each entry represents a different version of the inference engine image and includes the default and specific parameters required for startup, as well as the necessary startup commands. This helps automatically select the optimal inference engine configuration based on the model and GPU resource requirements.
[0055] It should be noted that the hardware devices in the various embodiments of the present application may be, but are not limited to, GPU boards or GPU cards.
[0056] In this embodiment, the automated process from the user creating an inference service request to the actual operation of the inference service is described in detail, especially how to automatically select and configure the optimal inference engine image and startup parameters based on the characteristics of the model and GPU resources. Through the introduction of the LLM-operator component and the matching processing of the model information configuration table and the engine information configuration table, the entire implementation process is made efficient and accurate, which can significantly reduce the user's manual intervention and configuration errors, while maximizing the use of GPU resources, improving the user experience of the inference service and the overall system performance. The technical solution of this application is particularly suitable for scenarios with diversified GPU hardware device models and dynamic changes in GPU resources, such as inference service deployment in cloud computing environments.
[0057] In other words, the technical solution of this application mainly designs a unified CRD object of type llm-instance for large model inference services using GPUs in a cloud-native way; when users create large model services, they only need to create CR resources, fill in the specified model file name and the required GPU memory quantity; design a GPU detection component, detect the GPU resources existing on the cluster nodes according to the configured GPU information table, and report it to the LLM-operator; design an LLM-operator, watch the CR resource creation event of type llm-instance, automatically compare the existing GPU information in the cluster according to the model file name and memory quantity applied for in the CR resource, automatically calculate the required GPU card model and quantity, and complete the creation of the K8S workload object, which can shield the differences between GPUs of various manufacturers (depending on Figure 5The metadata structure shown in the figure is designed), which can quickly adapt to new GPU devices, simplify user operation thresholds, and has high practical value.
[0058] By adopting the above-mentioned embodiment provided by the present application, by reading the model name in the model file requested by the user when creating the inference service, the model information configuration table and the engine information configuration table are automatically matched, and resource description information, inference engine image identifier and inference engine startup parameters that can meet the model requirements are obtained. Based on the resource description information, inference engine image identifier and inference engine startup parameters, workload metadata is created, and the workload metadata is allocated to the target cluster node on the container scheduling platform to achieve the deployment of the inference service. In other words, by automatically matching at least one parameter of the workload metadata created by the user only based on the model name specified by the user, the low efficiency caused by manually configuring parameters in the related art is solved, and the parameter configuration process of the workload metadata is simplified, thereby improving the deployment efficiency of the model inference service.
[0059] In an exemplary embodiment, the above-mentioned response to the creation of the inference service, reading the pre-saved model information configuration table and the engine information configuration table, includes: in response to the creation of the inference service, obtaining the pre-configured configuration model information; based on the configuration model name and requested resource information in the configuration model information, reading the model information configuration table and the engine information configuration table, wherein the requested resource information includes the requested video memory amount of the hardware device determined according to the inference service.
[0060] When a user requests an inference service (such as a large model-based inference service) through an API in a Kubernetes cluster, the LLM operator immediately responds and retrieves the pre-configured configuration model information associated with the inference service. This configuration model information typically includes metadata specified when the user requests the inference service, such as the model name, the model file format, the required GPU memory size, and whether the GPU vendor type is specified. This descriptive information about the GPU resources requested by the user provides the basis for subsequent resource matching and workload metadata creation.
[0061] For example, suppose a user creates an inference service named my-llm-service, requests 40GB of GPU memory, and wishes to use a GPU card from a leading GPU vendor. The LLM-operator responds to this request by obtaining the model name (e.g., Large Language Model), memory size (40GB), and GPU vendor type from the user's request information.
[0062] Based on the model name and requested resources (primarily including video memory size and optional GPU vendor type) in the obtained configuration model information, the LLM operator queries the model information configuration table and engine information configuration table stored in the cluster. The model information configuration table stores the GPU types and resource keys supported by each model, as well as the corresponding inference engine images. The engine information configuration table records detailed information for each inference engine image, such as the image source and startup parameters. The purpose of this process is to determine the appropriate inference engine image, GPU resource key, and any specific startup parameters to meet the resource requirements and model compatibility requested by the user.
[0063] Continuing with the above example, the LLM-operator uses the model name and the requested memory size of 40GB to search the model information configuration table. It finds the entry with the first GPU vendor name in the model information configuration table and learns that the large language model should use the inference engine image vllm:23.09 on the vendor-provided GPU. Recommended startup parameters may include memory utilization, quantization level, and tensor parallelism. The model deployment controller then checks the vllm:23.09 entry in the engine information configuration table to obtain the specific inference engine image location (also known as a URL) and startup command, preparing for subsequent workload creation.
[0064] This embodiment aims to extract key information from user requests and, based on this information, query a predefined configuration table to determine the GPU resource description, inference engine image, and startup parameters required to create an inference service. This process emphasizes the importance of configuration information and how the system dynamically utilizes this information to automatically adapt hardware resources, simplifying the inference service creation process and improving the automation and accuracy of resource scheduling.
[0065] Through the above approach, the system can intelligently match user needs with cluster resources. This not only avoids the configuration complexity and error rate caused by users' lack of understanding of GPU hardware details and inference engine versions, but also reduces the pressure on administrators to frequently update and maintain the system to adapt to new hardware and models, thereby improving resource utilization efficiency and user experience in the entire AI service ecosystem.
[0066] In an exemplary embodiment, the above-mentioned matching processing of the model information configuration table and the engine information configuration table based on the model name configured when creating the inference service includes: based on the model name, fuzzy matching of the model information configuration table and the engine information configuration table to obtain resource description information, inference engine image identifier and inference engine startup parameters adapted to the inference service, wherein the model information configuration table includes the mapping relationship between the model name and the device type of the hardware device, and the inference engine image identifier required to run the target model corresponding to the model name, and the engine information configuration table includes the inference engine image identifier and the inference engine startup parameters saved in the form of a first-ary structure.
[0067] When a user requests to create an inference service, the system first searches or queries the model information configuration table based on the model name in the user's request using a fuzzy matching algorithm to preliminarily determine the GPU type and inference engine image required to load the model. The use of a fuzzy algorithm allows the system to flexibly locate the configuration information that most closely matches the model. Even if the model name is not exactly the same as in the configuration table, the relevant GPU type and recommended inference engine image can be found. This fuzzy matching capability is crucial for handling constantly updated and changing model file names, ensuring that the required resources and inference engine are correctly identified and matched even when the model name changes.
[0068] After initially determining the GPU type and inference engine image required by the model, the system performs a fuzzy match on the engine information configuration table to find the inference engine image identifier and its default or specific startup parameters that match the model. Even if the entries in the engine information configuration table don't exactly match the model's direct requirements, the system's algorithm can find the closest inference engine configuration, ensuring that the inference service can start and run efficiently.
[0069] It should be noted that in the embodiments of the present application, the following Figure 5 The three types of metadata structures shown in the figure are: GPU metadata is mounted to the GPU detection component in the form of configmap, model information (stored in the model information configuration table) and inference engine information (stored in the engine information configuration table) are mounted to the model deployment controller component, and data is pre-set and initialized.
[0070] Among them, according to the model name, fuzzy matching Figure 5 The model information configuration table shown in the figure obtains the resource description information, which includes but is not limited to the inference engine image identifier (such as vllm:23.09) supported by the current model, the resource name resourcekey, the startup parameters, etc.
[0071] Based on the inference engine image identifier in the resource description information, the engine information configuration table is matched. If the inference engine image identifier is matched, the inference engine image is pulled from the image repository and the inference engine startup parameters are obtained at the same time.
[0072] An inference engine is a software system responsible for executing pre-trained machine learning or deep learning models and providing inference services. This includes loading the model, processing input data, executing model inference calculations, generating output results, and possibly optimizing and managing the model. The inference engine can be standalone software designed specifically for executing models.
[0073] An inference engine image is a container image that packages the inference engine and its runtime dependencies. A container image is a lightweight, portable software package that contains all the files, libraries, environment settings, and configurations required to run an application. In modern cloud-native environments, especially clusters managed by Kubernetes, container images are the standard form of deploying and running applications. An inference engine image contains the inference engine binary code, model files (if pre-loaded), drivers, libraries, and a runtime environment specific to GPUs or other hardware accelerators.
[0074] The inference engine is functional software responsible for the model's inference calculations. The inference engine image is a containerized data package that contains the inference engine and all the dependencies and configurations required for its operation. It can be directly scheduled and run by Kubernetes or other container management systems.
[0075] The inference engine image is the specific implementation form of the inference engine in a containerized environment. It allows the inference engine to be easily deployed on different servers or nodes without having to be installed and configured separately in each location.
[0076] In a Kubernetes cluster, when a user requests to create a model inference service, the system selects the appropriate inference engine image based on the requested GPU type and inference requirements, and launches the container on a node with the corresponding GPU resources, thereby providing efficient inference services. This approach improves deployment flexibility and resource utilization, simplifying operations and maintenance.
[0077] In an exemplary embodiment, the above-mentioned fuzzy matching of the model information configuration table and the engine information configuration table based on the model name includes: calling the model repository based on the model name; obtaining the model metadata information of the target model from the model repository, wherein the inference service is a service in the target model; determining the target device identifier of the hardware device that provides computing resources for the inference service based on the model metadata information; determining the resource description information adapted to the inference service based on the target device identifier, the resource description information is used to describe the target computing resources required to load the target model; determining the inference engine image identifier and the inference engine startup parameters from the engine information configuration table based on the target device identifier.
[0078] When a request to create an inference service is received, the first step is to call the model repository based on the model name included in the request to obtain the model's metadata. The model repository is a centralized storage area that stores model files for various models and provides an interface to access these files. By directly calling the model repository using the model name, you can efficiently obtain specific model information, such as the file format and the minimum hardware requirements required to run the model.
[0079] The model repository (which can also be understood as the model repository module or model repository component on the container scheduling platform) can provide model file storage and access services to the outside world in the form of, but not limited to, NFS (Network File System).
[0080] After receiving a call request, the Model Repository returns the target model's metadata. This step ensures that the obtained information is directly relevant to the inference service created by the request based on an exact match of the model name, avoiding resource scheduling errors caused by inaccurate or incomplete model information.
[0081] Based on the model metadata, particularly the model size and recommended GPU memory requirements, the target device ID (GPU vendor ID) of the hardware device (GPU card) in the Kubernetes cluster that can provide sufficient computing resources is determined. This process involves querying the model information configuration table to find the GPU type and corresponding resource key that best suits the model.
[0082] After identifying the target device, the system further refines the resource description to describe the specific computing resources required to load the target model, such as the number and model of GPU cards. This description is crucial for Kubernetes, as it determines how GPU resources are allocated and scheduled within the cluster to meet the needs of the inference service.
[0083] Based on the target device ID, the system retrieves the most suitable inference engine image ID and startup parameters from the engine information configuration table. This step ensures that the inference engine selection and configuration match the target hardware device, thereby optimizing the performance and efficiency of model inference.
[0084] This embodiment aims to call the model repository by model name, obtain metadata information, and then determine the target device identification and resource description information, ultimately matching the optimal inference engine image and parameters. This ensures that the creation of the inference service is not only based on the accurate identification of the model name, but also fully considers the actual operation requirements of the model and the hardware resource status within the cluster, achieving optimal resource allocation and utilization, reducing operational complexity, and improving the accuracy and efficiency of resource scheduling.
[0085] In an exemplary embodiment, the above-mentioned determination of the target device identifier of the hardware device that provides computing resources for the inference service based on the model metadata information includes: based on the model file format and the model name in the model metadata information, querying the model information configuration table to obtain a list of hardware resources currently available for loading the inference service in the container scheduling platform; and determining the target device identifier from at least one device identifier contained in the hardware resource list.
[0086] When a user requests to create an inference service, the system first queries the model information configuration table based on the model file format and model name in the model metadata. Figure 5 As shown, the model information configuration table is a predefined metadata structure that details the specific GPU hardware requirements of different models, including supported GPU vendors, memory size, and recommended resource keys and inference engine images. This step is the foundation of the entire resource matching process, ensuring that the system can accurately identify the model's hardware preferences and determine a list of candidate GPU resources based on this information.
[0087] After determining the hardware resource list obtained through preliminary screening, the target device ID that best suits the current inference service requirements is selected from the hardware resource list. This decision-making process takes into account the user's specific requirements (such as whether a GPU vendor has been specified), the current availability of GPU resources, and the model's GPU performance requirements. The system analyzes the characteristics of each GPU in the hardware resource list, including its vendor, memory size, and current utilization, to determine the optimal GPU and thus determine the target device ID.
[0088] By querying the model information configuration table and based on the model file format and name, we efficiently filter the list of hardware resources currently available for loading inference services on the container scheduling platform (also known as the Kubernetes cluster). We then select the optimal target device identifier (GPU vendor) from this list to match the inference service's computing resources.
[0089] In other words, by deeply analyzing model metadata and efficiently querying the model information configuration table, a precise match is achieved between the inference service and the GPU resources in the container scheduling platform. This solution automatically identifies the file format and name of the model, generates a list of compatible hardware resources, and selects the target device identifier, that is, the most suitable GPU manufacturer and model, from which it improves the intelligence and efficiency of resource scheduling. Moreover, users do not need to deeply understand the details of the GPU to create, load, and use high-speed and stable inference services. At the same time, system resource utilization is optimized, operation and maintenance costs are reduced, and strong technical support is provided for the rapid deployment and efficient operation of AI application scenarios.
[0090] In an exemplary embodiment, the above-mentioned determining the target device identifier from at least one device identifier contained in the hardware resource list includes: when the first field in the model metadata information is not configured with a default device identifier, determining the target device identifier from the at least one device identifier according to a preset priority.
[0091] By querying the first field in the model metadata information, determine whether the user configured the default device identifier when creating the inference service. For details, please refer to Figure 3 In the user-created CR resource section, the first field may be, but is not limited to, gpuvendor. If the field is configured with a device identifier, it is determined that the user has specified a GPU vendor; otherwise, the GPU vendor is not specified.
[0092] If it is determined based on the first field that the user has not specified a GPU manufacturer, the optimal GPU manufacturer may be automatically selected based on, but not limited to, a preset priority. Priority setting criteria may include, but are not limited to, GPU performance, the ranking of each device identifier in the hardware resource list, and the frequency of use of each device identifier.
[0093] In this embodiment, by introducing a mechanism for dynamically selecting GPU manufacturers, when the default device identifier is not explicitly specified in the model metadata information requested by the user, the system automatically selects the optimal target from compatible GPU manufacturers according to a preset priority order to load the inference service.
[0094] This screening mechanism not only enhances system flexibility but also improves resource allocation efficiency. Especially in multi-GPU environments, it ensures that even when users have no or ambiguous preferences, the system can still intelligently and quickly schedule resources, ensuring high-performance inference services. Furthermore, the dynamic selection mechanism reduces user reliance on GPU expertise and simplifies the process of creating inference services.
[0095] In other words, even if the user doesn't specify a GPU vendor, GPU resources are automatically selected based on pre-set priorities, enhancing the intelligence and flexibility of the container scheduling platform. This mechanism ensures that the inference service automatically matches the GPU vendor with the best performance and resources based on model characteristics, enabling efficient deployment even in changing hardware environments.
[0096] In an exemplary embodiment, the above-mentioned determining the target device identifier from the at least one device identifier according to the preset priority includes at least one of the following: determining the default arrangement order of the at least one device identifier as the preset priority, and determining a device identifier with a higher order as the target device identifier; configuring the preset priority of the at least one device identifier based on the frequency of use of each device identifier in the at least one device identifier, wherein the frequency of use is proportional to the priority of the at least one device identifier; and determining the target device identifier with the highest priority from the at least one device identifier according to the preset priority.
[0097] In this embodiment, the priority of at least one device identifier in the hardware resource list can be set in two ways, but is not limited to the following. The first way is to use the default order of GPU vendors in the hardware resource list as the preset priority. Based on this default order, the GPU vendor with the highest priority (ranked first) is selected as the target device identifier for loading the inference service.
[0098] For example, assuming that the GPU manufacturers defined in the model information configuration table are arranged as the first manufacturer, the second manufacturer, and the third manufacturer, when the user does not explicitly specify the GPU manufacturer, the system will first try to match the GPU resources of the first manufacturer as the target device identifier; if the GPU resources of the first manufacturer do not meet the requirements, the GPU resources of the second and third manufacturers will be checked in order until the most suitable GPU resources are determined.
[0099] The second method is to configure a preset priority based on the GPU vendor's historical usage frequency in inference services. In this mechanism, GPU vendors with higher usage frequency are given higher priority. Therefore, if the user request does not specify a GPU vendor, the system automatically selects the GPU vendor with the highest usage frequency and priority as the target device identifier based on the GPU vendor's usage frequency priority.
[0100] By setting priorities through the first static default sorting order and the second dynamic usage frequency, you can flexibly switch according to actual conditions, ensuring that the inference service always runs on the optimal GPU resources.
[0101] Among them, the default sorting mechanism is suitable for scenarios where there is a lack of historical data or user preferences are relatively fixed. It directly points to the preferred GPU manufacturer through a pre-defined order, simplifying the decision-making process. The method of setting priorities based on frequency of use can dynamically adjust priorities based on actual usage, giving priority to GPU resources that have outstanding performance and frequent use in historical data, further improving the efficiency of resource scheduling and the quality of inference services. The integrated application of these two mechanisms not only takes into account the agility of the system, but also enhances its adaptability and intelligence, and has significant technical effects on optimizing GPU resource scheduling, improving inference service performance, and reducing user intervention.
[0102] In an exemplary embodiment, the above-mentioned determination of the target device identifier from at least one device identifier contained in the hardware resource list also includes: when the first field in the model metadata information has been configured with a default device identifier, determining the default device identifier as the target device identifier.
[0103] In combination with the description in the above embodiment, it can be seen that first, it is determined based on the first field whether the user configured a default GPU manufacturer when creating the inference service. If the first field is empty, it means that the user did not specify a GPU manufacturer.
[0104] If the default device ID (that is, the user-specified GPU vendor) has been configured, no further GPU vendor filtering or dynamic adjustment is required. Instead, the default device ID is used as the target device ID, and the inference engine image ID configured for the model in the specified vendor is directly filtered based on the GPU vendor indicated by the target device ID.
[0105] This mechanism focuses on personalized user needs, ensuring that inference services can run on the user's preferred GPU resources, providing optimal performance and meeting specific computing requirements. Furthermore, by directly matching the user's specified GPU vendor, the system avoids additional resource exploration and matching processes, further reducing service deployment latency and improving overall responsiveness.
[0106] By using the default device identifier in the model metadata, users' preferences for specific GPU vendors are directly met, enabling efficient and personalized resource scheduling. This avoids the complex selection and matching process, reduces latency in creating inference services, and enhances the user experience. Furthermore, by precisely matching user-specified GPU resources, inference services can achieve optimal performance at runtime, ensuring high efficiency and stability in scenarios with high computational demands, demonstrating the system's rapid response to user needs and its ability to optimize resource allocation.
[0107] In an exemplary embodiment, the above-mentioned method determines the resource description information adapted to the inference service based on the target device identifier; obtains a preset requested video memory quantity and requested computing resources; based on the target device identifier, obtains a device model list of currently available target hardware devices in the container scheduling platform detected by the detection component, wherein the target device model of the target hardware device corresponds to the target device identifier; when the remaining video memory quantity of the first hardware device of the first device model in the device model list is greater than or equal to the requested video memory quantity, determines the target computing resource based on the remaining video memory quantity and the requested computing resource.
[0108] After determining the target GPU manufacturer according to the method in the above embodiment, the VRAM memory size requested by the user is further compared with the GPU information reported by the GPU detection component, and the GPU model is determined based on the comparison result.
[0109] Specifically, when a user initiates an inference service creation request, the system first parses the model metadata to extract the requested amount of video memory and computing resource requirements. This data forms the basis for the subsequent GPU resource matching process, ensuring that the inference service is configured with the appropriate GPU configuration when it is created.
[0110] Based on the target device ID, the system queries the information reported by the GPU detection component to obtain a list of all available GPU device models in the current container scheduling platform that match the target device ID. This step ensures that the system understands which GPU resources can be used to meet specific user requests.
[0111] After obtaining a list of target hardware models, the system further compares the remaining memory size or amount of each GPU model. For example, it finds the first GPU model with a remaining memory amount equal to or greater than the user's requested amount. Once the GPU model that meets the requirements is identified, the system can then accurately calculate the target computing resources required to create the inference service based on this model and its remaining computing resources.
[0112] Through the above steps, precise matching of GPU resources based on the target device ID is achieved. First, the system captures the user's request for video memory and computing resources through model metadata information, ensuring the accuracy of the foundation for subsequent matching. Then, based on real-time information detected by the GPU detection component, it accurately obtains a list of currently available GPU models that match the user's requested device ID. Finally, by comparing the remaining video memory capacity of GPU models, it locates the first GPU model that meets the user's video memory requirements and uses this to calculate the target computing resources required to create the inference service.
[0113] Obviously, the target device identifier may include multiple GPU models, that is, a GPU manufacturer may provide multiple GPU boards of different models.
[0114] By obtaining the user's requested video memory and computing resource requirements and combining them with the target device identifier to accurately match GPU resources, we ensure efficient resource utilization when creating inference services. This mechanism intelligently selects GPU models that meet the user's video memory requirements and have matching computing power, effectively reducing resource waste and accelerating the deployment of model inference services.
[0115] In an exemplary embodiment, the above method also includes: when the remaining video memory amount is less than the requested video memory amount, determining a second device model from the device model list; when the remaining video memory amount of the second hardware device of the second device model is greater than or equal to the requested video memory amount, determining the target computing resource based on the remaining video memory amount and the requested computing resource.
[0116] First, check whether the remaining video memory of the currently available GPU meets the user's requested video memory requirement. If the remaining video memory of a certain GPU model (the first device model) is found to be less than the user's requested amount, this indicates that the current GPU may not be sufficient to independently support the inference service to be created.
[0117] This means searching the device model list for the next GPU model (the second device model) that can meet the video memory requirements. This step demonstrates the system's flexibility and resource allocation capabilities, enabling it to quickly find alternative solutions when the first choice fails.
[0118] After determining the second device model, the system verifies again whether the remaining graphics memory is greater than or equal to the user's requested graphics memory. Based on this information and the user's requested computing resources, the system determines the final target computing resources. This process ensures that even if the initial selection fails, the system can still provide sufficient GPU resources for the inference service, avoiding performance bottlenecks caused by insufficient resources.
[0119] For example, suppose the inference service a user requests to create requires 20GB of video memory, but the A100-PCIE-40G GPU provided by the first GPU manufacturer currently has only 18GB of video memory left. The system will determine that the remaining video memory is insufficient.
[0120] In this case, detect other models of GPU boards, for example, detect the GPU board model L20-PCIE-48G, until a GPU model with at least 20 GB of remaining video memory is found.
[0121] After determining that the remaining video memory of the L20-PCIE-48G GPU board is sufficient, calculate the number of GPU boards required to create the inference service based on its characteristics and the computing resource requirements requested by the user, and determine the target computing resources.
[0122] This approach enables efficient resource replenishment and optimized configuration by intelligently identifying alternative GPU models when remaining video memory is insufficient. This ensures that even if initial resource allocation doesn't meet demand, matching GPU resources can be quickly adjusted and found, avoiding service creation delays and the risk of performance degradation.
[0123] In an exemplary embodiment, the above-mentioned determining the inference engine image identifier and the inference engine startup parameters from the engine information configuration table based on the target device identifier includes: querying the second field from the model information configuration table based on the target device identifier; when the second field is found from the engine information configuration table, determining the inference engine containing the second field as the target inference engine; and obtaining the inference engine image identifier and the inference engine startup parameters based on the engine description information of the target inference engine.
[0124] When a user requests to create an inference service, the target device identifier (i.e., the user-specified GPU manufacturer and model) is first used to query the model information configuration table for the second field applicable to that device. This second field refers to the model file format and related metadata compatible with the specific GPU device, which is an important basis for inference engine selection.
[0125] The second field can be but is not limited to: Figure 5 The vllm:23.09 shown in the figure matches the engine information configuration table according to the second field. If the same field is matched from the engine description information in the engine information configuration table, the target inference engine is determined, and the inference engine image identifier and inference engine startup parameters in the target inference engine are obtained. These parameters are used to create workload metadata.
[0126] Through this approach, we achieve automated inference engine selection and parameter configuration based on the target device's identification. First, the model information configuration table determines the model file format and metadata, which serve as the basis for selecting an inference engine. Then, within the engine information configuration table, we locate an inference engine image compatible with the specified format, ensuring a match between the inference engine and the target device. Finally, we extract key parameters from the inference engine description, completing the preparations for workload metadata creation.
[0127] The above approach not only simplifies the parameter determination process for users to create workload metadata, eliminating the need for users to have in-depth understanding of GPU compatibility and inference engine details, but also improves the efficiency of workload metadata creation. By automatically selecting the most suitable inference engine image and startup parameters, the system can quickly respond to user needs and accelerate the deployment of inference services, while ensuring high-quality service operation and improving the flexibility of the entire container scheduling platform.
[0128] In an exemplary embodiment, the above method also includes: when the second field is not found in the engine information configuration table, determining the default reasoning engine as the target reasoning engine; based on the engine description information of the default reasoning engine, obtaining the reasoning engine image identifier and the reasoning engine startup parameters.
[0129] When a field (such as the second field) corresponding to the model file format requested by the user cannot be found in the engine information configuration table, or a suitable inference engine cannot be determined for some reason, the default inference engine mechanism will be enabled.
[0130] Specifically, you can refer to Figure 5 If no match is found for the default engine shown in the Engine Information Configuration table, the system will use the inference engine image identifier and engine startup parameters from its description information as the basis for creating the inference service workload. This process ensures that the inference service can be deployed using the default inference engine even without specific configuration.
[0131] This approach ensures that even if the target inference engine is not matched, the default inference engine mechanism ensures the smooth operation of the service creation process. This avoids delays or failures in inference service deployment caused by missing model or GPU description information, improving system resilience and user service experience. In other words, through the pre-set default inference engine, the system can automatically select and configure the inference engine, ensuring that even in the absence of specific configuration, the model inference service can still be deployed and run efficiently and reliably.
[0132] In order to more clearly understand the overall implementation process of determining and creating workload metadata, the following Figure 4The flowchart shown further describes this.
[0133] S402: Call the model repository according to the model name to obtain model metadata information.
[0134] The model metadata information includes but is not limited to the model file format, model name, etc.
[0135] S404: According to the model name and model file format, query the model information configuration table to obtain a list of GPUs supported by the model, as well as the resource key and inference engine image name corresponding to each GPU.
[0136] The GPU list includes but is not limited to device identifiers consisting of GPU vendors supported by the model. Since different model inference services may support different GPU vendors, users may also specify a GPU vendor when creating a model inference service. Therefore, you need to determine the GPU vendor first.
[0137] S406: Determine whether the user has specified a GPU manufacturer.
[0138] Specifically through Figure 3 The user-created CR resource section shown uses the value of the gpuvendor field to determine whether the user has specified a GPU vendor.
[0139] If specified, execute step S408; otherwise, jump to step S414.
[0140] S408: If the user specifies a GPU manufacturer, directly filter out the inference engine image ID (inference engine image identifier) configured for the model in the specified manufacturer.
[0141] S410: According to the determined manufacturer, query the engine information configuration table to obtain the startup parameters of the inference engine.
[0142] The startup parameters include but are not limited to default startup parameters and specific startup parameters.
[0143] S412: Build workload metadata based on the inference engine image identifier and the inference engine startup parameters, thereby creating a workload.
[0144] S414, selecting the first manufacturer from top to bottom according to the manufacturer order in the model information configuration table;
[0145] S416 , comparing the VRAM video memory size requested by the user with the GPU information reported by the GPU detection component.
[0146] S418: Get the list of GPU models available to manufacturer A in the K8S cluster.
[0147] S420 , determining whether the remaining available video memory size of model 1 is greater than the VRAM video memory size requested by the user.
[0148] If yes, execute step S410; otherwise, execute step S422.
[0149] S422: Check whether there are other models of GPU boards provided by manufacturer A in the cluster.
[0150] If yes, execute step S418; otherwise, jump to step S414 and select another GPU manufacturer.
[0151] In an exemplary embodiment, the above-mentioned creation of workload metadata for carrying the inference service based on the resource description information, the inference engine image identifier and the inference engine startup parameters includes: automatically packaging the resource description information, the inference engine image identifier and the inference engine startup parameters according to a preset format to obtain the workload metadata.
[0152] After confirming the target GPU device and corresponding inference engine, an automated packaging process is implemented to format the acquired resource description information (including the GPU card model, number of GPUs, and memory size), the inference engine image identifier, and the inference engine startup parameters, and integrate them into workload metadata that is easy for Kubernetes to understand and execute. This packaging process ensures the integrity, consistency, and accuracy of the information.
[0153] The form of the encapsulated workload metadata can be, but is not limited to, Figure 3 As shown, it includes the workload type kind, GPU manufacturer and quantity, model name A100, inference engine image identifier (which can also be understood as the network address of the inference engine image) vllm:23.09, inference engine startup command command and startup parameters, etc.
[0154] By automatically encapsulating resource descriptions, inference engine image identifiers, and startup parameters, standardized workload metadata is formed, simplifying resource management and inference service deployment. Automated encapsulation ensures accurate information delivery and efficient Kubernetes scheduling, reducing manual intervention, avoiding configuration errors, and accelerating service rollout.
[0155] In this way, users can focus more on business logic and model optimization without having to worry about the details of the underlying infrastructure. This improves the efficiency of deploying model inference services on the container scheduling platform, simplifies user operation processes, and optimizes resource scheduling.
[0156] In an exemplary embodiment, before reading the pre-saved model information configuration table and engine information configuration table in response to the creation of the inference service, the above method also includes: setting at least one detection component on each cluster node in the container scheduling platform; reading the node resource information on each cluster node detected by the at least one detection component at a preset time interval; and generating a resource information configuration table based on the node resource information.
[0157] At least one GPU detection component is deployed on each node in the Kubernetes cluster, the container scheduling platform. At preset intervals, this detection component regularly reads and reports GPU resource information on the node. This information, including the GPU model, number of GPUs, and memory size, forms the basis for resource information updates.
[0158] For example, the detection component executes an LSPCI command every five minutes to obtain PCI device information on the node. It then parses the GPU information configuration table to obtain the detailed specifications of each GPU card, including PCI ID and VRAM size. This data is then packaged and sent to the LLM operator via the REST interface, ensuring the real-time and accurate resource information.
[0159] The system summarizes the collected node resource information and generates or updates the resource information configuration table. This table records the latest status of all GPU resources in the cluster, including manufacturer, model, memory size, and available quantity, providing detailed data basis for subsequent inference engine selection and workload scheduling.
[0160] By setting up a detection component on each cluster node, we achieve continuous monitoring of node resource information and dynamic updates of the resource information configuration table. This ensures that the system always has the latest GPU resource status, enhancing the container scheduling platform's adaptability to heterogeneous environments and the flexibility of resource scheduling. Through real-time monitoring and configuration table updates, the system can more accurately match GPU resources with inference service requirements, avoiding resource waste and improving computing efficiency and user satisfaction.
[0161] In an exemplary embodiment, the above-mentioned reading of the node resource information on each cluster node detected by the at least one detection component according to a preset time interval includes: executing a target detection command according to a preset time interval; in response to the target detection command, detecting the cluster node provided with the detection component to obtain the node resource information, and storing the node resource information in the resource information configuration table in the form of a second metadata structure.
[0162] As described in the preceding example, the detection component executes LSPCI commands at preset intervals to obtain PCI device information on cluster nodes. It then parses the GPU configuration table to obtain the detailed specifications of each GPU, including PCI ID and VRAM size. This data is then packaged and sent to the LLM operator via the REST interface, ensuring the real-time and accurate delivery of resource information.
[0163] After detecting the node resource information of each cluster node, Figure 5 The second metadata structure shown saves the node resource information to the GPU information configuration table, which is a standardized data format that is easy for the system to understand and process.
[0164] The second metadata structure includes not only the GPU model, total number and memory size, but also the available memory for each GPU model, so that the resource information configuration table can more comprehensively reflect the GPU resource status of the cluster.
[0165] By regularly executing probe commands to collect and update GPU resource information for cluster nodes, the real-time and accuracy of the resource information configuration table is ensured. This dynamic monitoring and update mechanism enables the system to promptly respond to changes in GPU resources, such as fluctuations in video memory usage and the addition or removal of GPU models. This enables more intelligent and accurate resource scheduling during the inference service creation process, avoiding resource waste, improving computing efficiency, and providing users with a more stable and efficient service experience. Furthermore, the standardized secondary metadata structure simplifies the information processing process and enhances the system's maintainability and scalability.
[0166] In an exemplary embodiment, the method further includes: reporting the node resource information to a model deployment controller in real time, wherein the node resource information includes the device identifier, device model, and video memory information of the hardware device providing the computing resources. At least one of the detection components, the model deployment controller, and the model repository are deployed on a container scheduling platform.
[0167] like Figure 3 As shown in the figure, after the detection component detects the GPU resource information of the cluster nodes, it will immediately report this information to the model deployment controller through the preset communication mechanism. The reported information includes but is not limited to key data such as the GPU device identifier (PCI-ID), model, and video memory information.
[0168] For example, in a Kubernetes cluster, suppose the detection component detects the addition of two new GPUs of the target model, each with 80GB of memory, to a cluster node. The detection component generates node resource information in a JSON-like format in real time and reports it to the model deployment controller (LLM-operator) via a REST API.
[0169] It should be noted that in this embodiment, the container scheduling platform, namely the Kubernetes cluster, deploys a detection component, a model deployment controller (LLM-operator), and a model repository. These three components work together to ensure the efficient creation and operation of the inference service.
[0170] For example, in a Kubernetes cluster, the detection component monitors GPU resources, the model repository stores and serves model files, and the model deployment controller interprets user requirements, matches GPU resources, selects inference engines, and ultimately creates workloads. This integrated platform design enables real-time feedback of resource information and promotes automated and intelligent management of inference services.
[0171] By using the detection component to report node resource information to the model deployment controller in real time, this system enables timely feedback and integrated management of GPU resource status. This improves the accuracy and timeliness of resource scheduling, ensuring that the inference service automatically selects the most suitable GPU device and memory specifications based on the latest resource availability, thereby accelerating the deployment process of the inference service and optimizing resource utilization.
[0172] In an exemplary embodiment, the above method also includes: when the target detection component detects that target resource information exists on one of the cluster nodes, adding a node label to one of the cluster nodes, wherein the node label includes the device model of the hardware device in the target resource information.
[0173] As described in the preceding embodiments, when the target detection component detects the presence of target resource information on a specific node, the system automatically adds a specific node label to that node. This node label contains the specific model information of the GPU device. This makes it easier for the system to identify nodes with specific GPU resources, thereby optimizing resource scheduling.
[0174] By dynamically adding node labels through the target detection component and working in conjunction with the model deployment controller and other components, an intelligent and flexible resource management and scheduling platform is built. In a Kubernetes cluster, the dynamic node labeling mechanism automatically updates node identifiers based on real-time changes in GPU resource information, ensuring that the model deployment controller is constantly aware of GPU resource status. This simplifies the process of creating inference services, eliminating the need for users to manually specify GPU models. The system automatically identifies and allocates the most appropriate resources, improving resource utilization, reducing maintenance costs, and enhancing the platform's scalability and user experience.
[0175] For example, suppose a user plans to create a large model inference service, but cannot understand the model and other details of the underlying GPU hardware. Using the technical solution in the embodiment of the present application, the user only needs to specify the required GPU memory size when creating the inference service, such as "VRAM=40GB". After receiving the application, the model deployment controller will automatically filter out nodes with sufficient video memory and matching GPU models for scheduling based on real-time node label information. For example, the system may recognize the node label and automatically select the corresponding model of GPU board as the optimal option based on the video memory size and model compatibility, thereby efficiently completing the creation and deployment of the inference service.
[0176] By automatically adding node tags containing device model information to nodes after GPU resources are identified by the detection component, dynamic resource association and identification is achieved. This significantly improves the intelligence and efficiency of resource scheduling, ensuring that inference services are accurately matched to nodes with suitable GPU devices, avoiding ineffective resource searches and waste, and improving system responsiveness.
[0177] From the description of the above embodiments, it can be seen that the main technical solutions of the technical solution of this application include the following points.
[0178] (1) Design of GPU detection component.
[0179] This module is deployed within the K8S cluster and runs on each K8S node in the form of a daemonset. This module is mainly responsible for regularly detecting whether the node has a GPU installed and sending the GPU manufacturer, model, and video memory information to the LLM-operator component. It also adds a label with the GPU model to the node for convenient application scheduling reference.
[0180] This module uses the LSPCI command to detect the information of the existing PCI devices on the node (here it is assumed that the GPU board is a PCI device), and obtains the GPU model and video memory size information corresponding to the PCI-ID according to the GPU information configuration table; then the node information and the GPU information on the node are sent to the lLLM-operator module through the rest interface for aggregation. The LLM-operator module constructs an overall view of the GPU information of the entire K8S cluster.
[0181] The GPU detection module can detect the GPU model and video memory information on each cluster node, add a special model label to the node, and interact with the LLM-operator module to report GPU resource information in real time for the LLM-operator to decide on the GPU model to use.
[0182] (2) Design of model warehouse module.
[0183] This module is mainly used to store large model weight files. It uses the NFS file server to provide an nfs directory for simultaneous mounting and access by multiple workloads. The model files themselves are not modified in any special way. For example, model files downloaded from huggingface can be uploaded and used.
[0184] (3) Design of LLM-operator module.
[0185] This module adopts the cloud-native operator architecture design and is deployed in the K8S cluster in the deployment mode. It continuously lists and watches the CR resources of type llm-instance. This module relies on two tables: model file information configuration and inference engine information configuration. Whenever the module watches the creation event of llm-instance, it will execute the internal K8S workload construction algorithm, automatically infer the key parameters required to create the K8S workload based on the model file and video memory size requested by the user, and finally create a workload instance that carries the large model inference service.
[0186] The key attributes of CR resources include model name, VRAM size requested by the user, GPU manufacturer specified by the user (optional), workload type (deployment, statefulset) and corresponding parameters. The determination method of each parameter in the workload metadata can refer to the above embodiment. Figure 4 Description.
[0187] The LLM-operator module is deployed in the K8S cluster as a stateless load. Multiple replicas can be deployed and a master is elected between the replicas to ensure that only one replica is performing detection at the same time. Internally, based on the metadata configuration, it reads the model file and GPU memory size requested by the user, automatically matches it to the specific inference engine image, and automatically calculates the actual GPU model and number, simplifying the complexity of creating the inference service.
[0188] (4) GPU-adaptive metadata structure design.
[0189] Mainly include Figure 5 The three types of metadata shown are: GPU information, model file information, and inference engine information; the metadata is stored in configmap format, allowing dynamic addition and update. The LLM-operator and GPU detection module will calculate and match GPU resources based on the content set in the metadata; the data structure design of the three types of metadata information is as follows Figure 5 shown.
[0190] By designing the above three types of metadata structures, it is possible to support the expansion of new inference engine images, new model files, and new GPU models in the form of configuration without modifying the business code.
[0191] The overall implementation process of a GPU adaptive large model reasoning service creation and deployment method proposed in the embodiment of this application is: using K8S's LLM-operator and custom-designed rule templates to support administrators to quickly configure new models of GPUs and reasoning engine rule resources, support users to automatically parse key information of model files when creating reasoning services, automatically match GPU resources, and automatically load the appropriate version of the reasoning engine for different GPUs to complete the creation of workloads, which has high practical value. Figure 3 The overall architecture diagram shown describes the overall implementation of the solution.
[0192] (1) Design GPU detection component.
[0193] The main execution logic within the GPU detection component includes the following:
[0194] 1) Get the PCI device list of this node.
[0195] 2) Based on the PCI device ID, query the GPU metadata information to obtain the GPU model, number of GPUs, and memory size of this node.
[0196] 3) Attach the GPU model information to this node as a node label.
[0197] 4) Report the detected GPU model, number of GPUs, and memory size of the node to the LLM-operator module.
[0198] (2) Develop the LLM-operator module and supporting CRD resources. Its internal main logic includes the following contents.
[0199] 1) Create a CRD resource.
[0200] 2) Monitor CR resource creation events.
[0201] 3) Based on the field information in the CR resource, the necessary workload metadata is constructed according to the GPU adaptive inference workload construction algorithm, mainly including GPU model, resource key and number of frames, inference engine image, inference engine startup parameters, etc.
[0202] 4) Call the K8S interface service and create a workload.
[0203] (3) Design the metadata structure, mount the GPU metadata information to the gpu detection component in the form of configmap, mount the model information and inference engine information to the llm-operator component, and complete the data pre-setting and initialization.
[0204] (4) Develop a model warehouse module to provide model file storage and access services to the outside world via NFS.
[0205] (5) Complete the deployment of the above modules in the K8S cluster, where the GPU detection module (or detection component) is deployed in daemonset mode, the LLM-operator is deployed in deployment mode, the metadata is deployed in configmap mode, and the model warehouse is deployed in NFS server mode.
[0206] It should be noted that Figure 3 The added node labels, vendor1-device-plugn and workload metadata shown are K8S native resource objects. The only difference between the technical solution of this application and the related technology lies in the different methods of obtaining and configuring multiple parameters in the workload metadata. The related technology mainly relies on manual configuration, while the technical solution of this application focuses on parameter configuration through automatic adaptation.
[0207] Figure 3The user-created CR resources, GPU detection components, LLM-operator (model deployment controller) and model warehouse shown are entity components designed in the technical solution of this application. The GPU information configuration (or GPU information configuration table), model information configuration (or model information configuration table) and engine information configuration (or engine information configuration table) are data structures designed in the technical solution of this application, namely metadata structures.
[0208] Through the above-mentioned embodiments provided in this application, it is possible to utilize the K8S model deployment controller and custom-designed rule templates to support administrators in quickly configuring new GPU models and inference engine rule resources. This allows users to automatically parse key information from model files when creating inference services, automatically match GPU resources, and automatically load the appropriate inference engine version for different GPUs to complete workload creation. This not only simplifies the user experience but also improves the overall efficiency and usability of the system, possessing important theoretical and practical significance.
[0209] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0210] According to another aspect of the embodiment of the present application, a deployment device for an inference service is also provided. The structural diagram of the system is shown in FIG. Figure 6 As shown, it includes the following modules: a reading unit 602, which is used to read a pre-saved model information configuration table and an engine information configuration table in response to the creation of an inference service, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine; a matching unit 604, which is used to match the model information configuration table and the engine information configuration table based on the model name configured when creating the inference service, and obtain resource description information, inference engine image identifier and inference engine startup parameters adapted to the inference service; a first creation unit 606, which is used to create workload metadata for carrying the inference service based on the resource description information, the inference engine image identifier and the inference engine startup parameters, wherein the workload metadata is used to describe the execution strategy of the containerized application; an allocation unit 608, which is used to allocate the workload metadata to the target cluster node to deploy the inference service on the target cluster node.
[0211] The specific execution steps involved in the various calculation processes in the above modules and the dynamic optimization of storage space can be referred to the description in the above embodiments and will not be repeated here.
[0212] Obviously, the above-mentioned screen display device can be used to implement the deployment method of the reasoning service provided in the above-mentioned embodiment, and the details that have been explained will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0213] It should be noted that the reading unit 602 in this embodiment can be used to execute the above step S202, the matching unit 604 in this embodiment can be used to execute the above step S204, the first creation unit 606 in this embodiment can be used to execute the above step S206, and the allocation unit 608 in this embodiment can be used to execute the above step S208.
[0214] In an exemplary embodiment, the above-mentioned reading unit 602 includes: a first acquisition module, used to obtain pre-configured configuration model information in response to the creation of the inference service; a reading module, used to read the model information configuration table and the engine information configuration table based on the configuration model name and requested resource information in the configuration model information, wherein the requested resource information includes the requested video memory amount of the hardware device determined according to the inference service.
[0215] In an exemplary embodiment, the matching unit 604 includes: a matching module for performing fuzzy matching on the model information configuration table and the engine information configuration table based on the model name to obtain resource description information, inference engine image identifier and inference engine startup parameters adapted to the inference service, wherein the model information configuration table includes a mapping relationship between the model name and the device type of the hardware device, and the inference engine image identifier required to run the target model corresponding to the model name, and the engine information configuration table includes the inference engine image identifier and the inference engine startup parameters saved in the form of a first-ary structure.
[0216] In an exemplary embodiment, the above-mentioned matching module includes: a first calling submodule, used to call the model repository based on the model name; a first acquisition submodule, used to obtain the model metadata information of the target model from the model repository, wherein the inference service is a service in the target model; a first processing submodule, used to determine the target device identifier of the hardware device that provides computing resources for the inference service based on the model metadata information; a second processing submodule, used to determine the resource description information adapted to the inference service based on the target device identifier, and the resource description information is used to describe the target computing resources required to load the target model; a third processing submodule, used to determine the inference engine image identifier and the inference engine startup parameters from the engine information configuration table based on the target device identifier.
[0217] In an exemplary embodiment, the above-mentioned matching module includes: a first query sub-module, which is used to query the model information configuration table based on the model file format and the model name in the model metadata information, and obtain a list of hardware resources in the container scheduling platform that can currently be used to load the inference service; a fourth processing sub-module, which is used to determine the target device identifier from at least one device identifier contained in the hardware resource list.
[0218] In an exemplary embodiment, the matching module further includes: a fifth processing submodule for determining the target device identifier from the at least one device identifier according to a preset priority when the first field in the model metadata information is not configured with a default device identifier.
[0219] In an exemplary embodiment, the above-mentioned matching module includes at least one of the following: a sixth processing sub-module, used to determine the default arrangement order of the at least one device identification as the preset priority, and determine the device identification with the highest ranking as the target device identification; a first configuration sub-module, used to configure the preset priority of the at least one device identification based on the frequency of use of each device identification in the at least one device identification, wherein the frequency of use is proportional to the priority of the at least one device identification; a seventh processing sub-module, used to determine the target device identification with the highest priority from the at least one device identification according to the preset priority.
[0220] In an exemplary embodiment, the above-mentioned matching module includes: a second acquisition sub-module, used to obtain a pre-set requested video memory quantity and requested computing resources; a third acquisition sub-module, used to obtain a list of device models of target hardware devices currently available in the container scheduling platform detected by the detection component based on the target device identifier, wherein the target device model of the target hardware device corresponds to the target device identifier; an eighth processing sub-module, used to determine the target computing resource based on the remaining video memory quantity and the requested computing resource when the remaining video memory quantity of the first hardware device of the first device model in the device model list is greater than or equal to the requested video memory quantity.
[0221] In an exemplary embodiment, the matching module includes: a ninth processing sub-module for determining a second device model from the device model list when the remaining video memory amount is less than the requested video memory amount; and a tenth processing sub-module for determining the target computing resource based on the remaining video memory amount and the requested computing resource when the remaining video memory amount of the second hardware device of the second device model is greater than or equal to the requested video memory amount.
[0222] In an exemplary embodiment, the above-mentioned matching module includes: a second query sub-module, used to query the second field from the model information configuration table based on the target device identifier; an eleventh processing sub-module, used to determine the inference engine containing the second field as the target inference engine when the second field is found from the engine information configuration table; and a fourth acquisition sub-module, used to obtain the inference engine image identifier and the inference engine startup parameters based on the engine description information of the target inference engine.
[0223] In an exemplary embodiment, the above-mentioned device also includes: a first processing unit, used to determine the default reasoning engine as the target reasoning engine when the second field is not found in the engine information configuration table; a first acquisition unit, used to obtain the reasoning engine image identifier and the reasoning engine startup parameters based on the engine description information of the default reasoning engine.
[0224] In an exemplary embodiment, the first creation unit 606 includes: a packaging module for automatically packaging the resource description information, the inference engine image identifier and the inference engine startup parameters according to a preset format to obtain the workload metadata.
[0225] In an exemplary embodiment, the above-mentioned device also includes: a setting unit, which is used to set at least one detection component on each cluster node in the container scheduling platform before reading the pre-saved model information configuration table and engine information configuration table in response to the creation of the inference service; a second processing unit, which is used to read the node resource information on each cluster node detected by the at least one detection component at a preset time interval; and a third processing unit, which is used to generate a resource information configuration table based on the node resource information.
[0226] In an exemplary embodiment, the above-mentioned second processing unit includes: an execution module for executing a target detection command at a preset time interval; a detection module for detecting cluster nodes provided with a detection component in response to the target detection command, obtaining the node resource information, and storing the node resource information in the resource information configuration table in the form of a second metadata structure.
[0227] In an exemplary embodiment, the above-mentioned device also includes: a reporting unit, which is used to report the node resource information to the model deployment controller in real time, wherein the node resource information includes the device identification, device model and video memory information of the hardware device providing computing resources, wherein at least one of the detection components, model deployment controller and model repository is deployed on the container scheduling platform.
[0228] In an exemplary embodiment, the above-mentioned device also includes: an adding unit, which is used to add a node label to one of the cluster nodes when the target resource information is detected on one of the cluster nodes through the target detection component, wherein the node label includes the device model of the hardware device in the target resource information.
[0229] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0230] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the method for deploying an inference service.
[0231] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the method for deploying an inference service at runtime.
[0232] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0233] According to another aspect of the embodiments of the present application, a computer program product is further provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned embodiments of the method for deploying an inference service are implemented.
[0234] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned embodiments of the method for deploying an inference service are implemented.
[0235] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0236] The above is a detailed introduction to a deployment method of an inference service provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A method for deploying an inference service, characterized by: include: In response to the creation of the inference service, read a pre-saved model information configuration table and an engine information configuration table, wherein the model information configuration table includes at least one model file, and the engine information configuration table includes description information of at least one inference engine; Based on the model name configured when creating the inference service, the model information configuration table and the engine information configuration table are matched to obtain resource description information, inference engine image identifier and inference engine startup parameters adapted to the inference service; Creating workload metadata for hosting the inference service based on the resource description information, the inference engine image identifier, and the inference engine startup parameters, wherein the workload metadata is used to describe an execution strategy of the containerized application; The workload metadata is distributed to a target cluster node to deploy the inference service on the target cluster node.
2. The method according to claim 1, characterized in that The step of reading a pre-saved model information configuration table and an engine information configuration table in response to creation of the inference service includes: In response to the creation of the inference service, obtaining pre-configured configuration model information; Based on the configuration model name and requested resource information in the configuration model information, the model information configuration table and the engine information configuration table are read, wherein the requested resource information includes the requested video memory amount of the hardware device determined according to the inference service.
3. The method according to claim 1, characterized in that The matching process of the model information configuration table and the engine information configuration table based on the model name configured when creating the inference service includes: Based on the model name, the model information configuration table and the engine information configuration table are fuzzy matched to obtain the resource description information, the inference engine image identifier and the inference engine startup parameters adapted to the inference service, wherein the model information configuration table includes a mapping relationship between the model name and the device type of the hardware device, and the inference engine image identifier required to run the target model corresponding to the model name, and the engine information configuration table includes the inference engine image identifier and the inference engine startup parameters saved in the form of a first-ary structure.
4. The method according to claim 3, characterized in that The performing fuzzy matching on the model information configuration table and the engine information configuration table based on the model name includes: Based on the model name, calling a model repository; Acquiring model metadata information of the target model from the model repository, wherein the inference service is a service in the target model; Determining, based on the model metadata information, a target device identifier of the hardware device that provides computing resources for the inference service; Determining, based on the target device identifier, the resource description information adapted to the inference service, where the resource description information is used to describe the target computing resources required to load the target model; Based on the target device identifier, the inference engine image identifier and the inference engine startup parameters are determined from the engine information configuration table.
5. The method according to claim 4, characterized in that The determining, based on the model metadata information, a target device identifier of the hardware device that provides computing resources for the inference service includes: Based on the model file format and the model name in the model metadata information, query the model information configuration table to obtain a list of hardware resources currently available for loading the inference service in the container scheduling platform; The target device identifier is determined from at least one device identifier included in the hardware resource list.
6. The method according to claim 5, characterized in that The determining the target device identifier from at least one device identifier included in the hardware resource list includes: In a case where the first field in the model metadata information is not configured with a default device identifier, the target device identifier is determined from the at least one device identifier according to a preset priority.
7. The method according to claim 6, characterized in that Determining the target device identifier from the at least one device identifier according to a preset priority includes at least one of the following: Determine the default arrangement order of the at least one device identification as the preset priority, and determine a device identification with a higher order as the target device identification; configuring the preset priority of the at least one device identifier based on a usage frequency of each device identifier in the at least one device identifier, wherein the usage frequency is proportional to the priority of the at least one device identifier; According to the preset priority, the target device identifier with the highest priority is determined from at least one device identifier.
8. The method according to claim 5, characterized in that The step of determining the target device identifier from at least one device identifier included in the hardware resource list further includes: In a case where the first field in the model metadata information has been configured with a default device identifier, the default device identifier is determined as the target device identifier.
9. The method according to claim 4, characterized in that The determining, based on the target device identifier, the resource description information adapted to the reasoning service includes: Get the preset requested video memory amount and requested computing resources; Based on the target device identifier, obtaining a device model list of target hardware devices currently available in the container scheduling platform detected by the detection component, wherein the target device model of the target hardware device corresponds to the target device identifier; When the remaining video memory amount of the first hardware device of the first device model in the device model list is greater than or equal to the requested video memory amount, the target computing resource is determined based on the remaining video memory amount and the requested computing resource.
10. The method according to claim 9, characterized in that The method further comprises: When the remaining video memory amount is less than the requested video memory amount, determining a second device model from the device model list; When the remaining video memory amount of the second hardware device of the second device model is greater than or equal to the requested video memory amount, the target computing resource is determined based on the remaining video memory amount and the requested computing resource.
11. The method according to claim 4, characterized in that The determining, based on the target device identifier, the inference engine image identifier and the inference engine startup parameters from the engine information configuration table includes: Based on the target device identifier, querying a second field from the model information configuration table; In the case where the second field is found in the engine information configuration table, determining the inference engine containing the second field as the target inference engine; Based on the engine description information of the target inference engine, the inference engine image identifier and the inference engine startup parameters are obtained.
12. The method according to claim 11, characterized in that The method further comprises: If the second field is not found in the engine information configuration table, determining a default inference engine as the target inference engine; Based on the engine description information of the default inference engine, the inference engine image identifier and the inference engine startup parameters are obtained.
13. The method according to claim 1, wherein The creating workload metadata for carrying the inference service based on the resource description information, the inference engine image identifier, and the inference engine startup parameters includes: The resource description information, the inference engine image identifier and the inference engine startup parameters are automatically packaged according to a preset format to obtain the workload metadata.
14. The method according to any one of claims 1 to 13, characterized in that Before reading the pre-saved model information configuration table and engine information configuration table in response to the creation of the inference service, the method further includes: At least one detection component is set on each cluster node in the container scheduling platform; Reading node resource information on each cluster node detected by the at least one detection component at a preset time interval; Based on the node resource information, a resource information configuration table is generated.
15. The method according to claim 14, characterized in that The reading, at a preset time interval, the node resource information on each cluster node detected by the at least one detection component includes: Execute target detection commands according to preset time intervals; In response to the target detection command, the cluster node provided with the detection component is detected to obtain the node resource information, and the node resource information is stored in the resource information configuration table in the form of a second metadata structure.
16. The method according to claim 14, characterized in that The method further comprises: The node resource information is reported to the model deployment controller in real time, wherein the node resource information includes the device identification, device model and video memory information of the hardware device providing computing resources, and the at least one detection component, the model deployment controller and the model repository are deployed on the container scheduling platform.
17. The method according to claim 14, characterized in that The method further comprises: When the target detection component detects that target resource information exists on one of the cluster nodes, a node label is added to the one cluster node, wherein the node label includes a device model of the hardware device in the target resource information.
18. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for deploying an inference service according to any one of claims 1 to 17 when executing the computer program.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for deploying an inference service according to any one of claims 1 to 17 are implemented.
20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for deploying an inference service according to any one of claims 1 to 17 are implemented.
Citation Information
Patent Citations
Online model reasoning system
CN111414233A
Inference service deployment method, device and equipment and storage medium
CN111625245A
Inference service configuration method and device, electronic equipment and storage medium
CN112015521A
Mask detection and deployment system and method based on image recognition
CN112085010A
Model reasoning system, method and equipment
CN114881236A
Cited By
Model deployment method, electronic equipment, storage medium and program product
CN121560346A
Creation method and device of game assistant model, terminal equipment and storage medium
CN121570817A
Inference service copy pool management method, electronic equipment and storage medium
CN122287909A
Inference service replica pool management method, electronic device, and storage medium
CN122287909B