Inference service management system, device, management method, medium, and program product
By building an inference service management system that includes storage, inference, and control modules, flexible expansion of container images and non-intrusive management of inference services are achieved, solving the problem of high coupling between container image scalability and code environment, and improving management convenience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, container images have poor scalability and flexibility, and the inference code is highly coupled with the environment, making inference service management inconvenient.
An inference service management system is constructed, including a storage module, an inference module, and a control module. The storage module stores the entry scripts of working components and the model files of pre-trained language models. The inference module deploys container groups and mounts image containers. The control module manages instructions, realizing the decoupling and non-intrusive management of business code and inference environment.
It improves the scalability and flexibility of container images, provides non-intrusive and scalable inference service management, and enhances the user experience.
Smart Images

Figure CN121433812B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a reasoning service management system, device, management method, medium and program product. BACKGROUND
[0002] In related technologies, a management framework / reasoning framework can implement management of reasoning services of large language models. These frameworks usually adopt an architecture of "management layer + reasoning engine layer", where the management layer is responsible for request routing, model management and node management, and the reasoning engine layer is responsible for loading models and providing reasoning services.
[0003] However, the above framework needs to build a customized image containing framework code and reasoning engine environment, resulting in poor scalability and flexibility of the container image, and the management framework and the reasoning engine are tightly coupled in code and environment, which is not conducive to the maintenance and management of reasoning services. SUMMARY
[0004] The present application provides a reasoning service management system, device, management method, medium and program product to at least solve the problem of poor scalability and flexibility of the container image in related technologies, and high coupling of reasoning code and environment, resulting in inconvenient management of reasoning services.
[0005] The present application provides a reasoning service management system, comprising: a storage module for storing an entry script of a work component and a model file of a pre-trained language model; a reasoning module for responding to a deployment instruction of the pre-trained language model, deploying at least one first container group inside the reasoning module, deploying at least one image container inside the first container group, mounting between the storage module and the first container group, and starting the work component through the entry script; the work component is mounted in the image container, the work component starts a reasoning engine in the image container, the reasoning engine loads a pre-trained language model corresponding to the model file, to provide reasoning services through the pre-trained language model, and the work component responds to a management instruction of the reasoning services to manage the reasoning services; a control module, the control module has a second container group deployed inside, and a management component is deployed inside the second container group, the management component sends a deployment instruction to the reasoning module, and the management component sends a management instruction to the work component.
[0006] The present application also provides an electronic device comprising the reasoning service management system described above.
[0007] The application further provides a reasoning service management method applied to the management component of the reasoning service management system, and the method comprises the following steps: identifying a user interaction action of an interaction interface; determining at least one of a deployment request and a management request of a user according to the user interaction action, generating a deployment instruction according to the deployment request, and generating a management instruction according to the management request; sending the deployment instruction to a reasoning module, deploying at least one first container group in the reasoning module in response to the deployment instruction, deploying at least one image container in the first container group, mounting the storage module and the first container group, starting a work component through an entry script, mounting the work component in the image container, starting a reasoning engine in the image container through the work component, loading a pre-training language model corresponding to a model file of the reasoning engine through the pre-training language model, and providing a reasoning service through the pre-training language model; and sending the management instruction to the work component, responding to the management instruction of the reasoning service through the work component, and managing the reasoning service.
[0008] The application further provides a nonvolatile computer readable storage medium, and the nonvolatile computer readable storage medium stores a computer program.
[0009] The application further provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the steps of the reasoning service management method.
[0010] Because the inference framework in the related art needs to build a customized image containing framework code and an inference engine environment, the scalability and flexibility of the container image are poor, which is not conducive to the maintenance and management of the reasoning service. Therefore, the embodiment of the application builds a reasoning service management system, which comprises a storage module, a reasoning module and a control module. The storage module stores an entry script of a work component and a model file of a pre-training language model. The reasoning module deploys a first container group and deploys an image container in the first container group. The storage module and the first container group are mounted. The work component is started through the entry script in the storage module. The entry script and the image container are decoupled. The specific entry script is the business code of the image container. The image container provides an inference environment of the reasoning service. The decoupling of the business code and the inference environment is achieved through the mounting mode. When the image container is expanded, the influence of the entry script on the expansion does not need to be considered. The flexible expansion of the container image is achieved, and the scalability of the container image is improved.
[0011] Meanwhile, by mounting the work component in the image container, non-intrusive management between the work component and the image container can be realized. The work component can start the inference engine in the image container, the inference engine loads the pre-trained language model corresponding to the model file to realize the inference service. The work component can also manage the inference service, and the control module is deployed with a second container group, the second container group is internally deployed with a management component. The management component can send a deployment instruction to the inference module and send a management instruction to the work component. Through the dual management between the management component and the inference module and the work component, flexible management of the image container and the inference service is realized, thereby providing non-intrusive and scalable management of the inference service and improving the user experience.
[0012] Therefore, the technical problem that the inference service is inconvenient to manage due to the poor scalability and flexibility of the container image and the high coupling between the inference code and the environment in the related art can be solved, the code and the environment are decoupled, the flexibility and scalability of the container image are improved, and the convenience of the inference service management is improved. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0014] Figure 1 A schematic diagram of the inference service management system according to an embodiment of the present application;
[0015] Figure 2 An architecture schematic diagram of the work component according to an embodiment of the present application;
[0016] Figure 3 An architecture diagram of the inference service management system according to an embodiment of the present application;
[0017] Figure 4 A timing diagram of the inference service management according to an embodiment of the present application;
[0018] Figure 5 A flowchart of the inference service management according to an embodiment of the present application;
[0019] Figure 6 A flowchart of the inference service management method according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0021] It should be noted that, in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0022] Before introducing the scheme of the present application, some application environment architectures and related technologies involved in the scheme of the present application will be introduced to assist the understanding of the scheme of the present application.
[0023] A container orchestration system is used for efficient scheduling and management of containerized applications deployed on multiple servers in a cloud environment. The platform can create multiple container groups (Pods), each of which runs at least one container, and these containers are responsible for carrying various business processes. With the help of the built-in load balancing mechanism, Kubernetes realizes unified management, automatic service discovery and access control of the container cluster.
[0024] The inference framework is specially designed to simplify and standardize the deployment and operation process of AI models. It mainly targets the current mainstream inference engines such as vLLM, SGLang, llama.cpp and Ollama for unified packaging and management. It provides a standardized interface integration of these engines and implements complete life cycle management, including model loading, unloading, model readiness detection, resource allocation, information statistics and other functions.
[0025] The version of the inference engine will be updated with its own functions and support for open source models. Among them, the version update and bug fixing for adapting to the constantly released open source models (such as DeepSeek-R1-0528, GLM-4.5, Kimi-K2-Instruct, etc.) account for the majority. If you want to always deploy the latest open source model, you need to constantly update the version of the inference engine.
[0026] Currently, large language model inference engines represented by vLLM, SGLang, etc. have been widely applied. In order to manage these inference services, some management / inference frameworks such as Xinference, FastChat, OpenLLM, etc. have appeared in the industry. These frameworks usually adopt the architecture of "management layer + inference engine layer", and the management layer is responsible for request routing, model management and node management, and the inference engine layer is responsible for loading models and providing inference services.
[0027] However, the related art generally has the following problems:
[0028] (1) Mirror image is highly customized: each management / inference framework needs to make a special Docker image that integrates the framework code and the environment dependencies of the inference engine, resulting in bloated images and complex management.
[0029] (2) Environment dependency is strongly coupled: the management framework and the inference engine are tightly coupled in code and environment. When the inference engine is upgraded or replaced, the framework code often needs to be modified and the image needs to be rebuilt, which is high in maintenance cost.
[0030] (3) In container orchestration systems such as Kubernetes, Kubernetes resource management and management / inference framework resource management are independent of each other, lacking standardized resource isolation and management capabilities under the Kubernetes system.
[0031] (4) Poor deployment flexibility: Since it is not natively supported in Kubernetes cloud-native scenarios, it is usually necessary to start the management component, then start the worker component node / Pod, and then select the worker component to load the model and start the service through the API (Application Programming Interface, Application Programming Interface) of the management component. The above solution needs to manage the model health status by itself and cannot be synchronized to the container orchestration system (for example, the worker component considers that the container group is ready after the worker component is started, but at this time the model in the container group has not been specified and loaded by the management component, and cannot provide services, only tells the management component that a new node / container group has joined the management / inference framework cluster, which overlaps with the functions of the container orchestration system and is out of sync), which makes it difficult to fully utilize the elasticity and fault self-healing capabilities of the container orchestration system.
[0032] In order for those skilled in the art to better understand the present application scheme, the present application will be further described in detail below in conjunction with the drawings and specific embodiments.
[0033] Embodiments of the present application provide an inference service management system, which is described in detail in conjunction with the components of the inference service management system.
[0034] In particular, Figure 1 A schematic diagram of a reasoning service management system provided according to an embodiment of the present application.
[0035] As Figure 1 shown, the reasoning service management system 10 comprises a storage module 11, a reasoning module 12 and a control module 13.
[0036] The storage module 11 is configured to store an entry script of a work component and a model file of a pre-trained language model; the reasoning module 12 is configured to, in response to a deployment instruction of the pre-trained language model, deploy at least one first container group inside the reasoning module 12, deploy at least one image container inside the first container group, mount the storage module 11 and the first container group, and start the work component through the entry script; the work component is mounted in the image container, the work component starts an inference engine in the image container, the inference engine loads a pre-trained language model corresponding to the model file, so as to provide an inference service through the pre-trained language model, and the work component manages the inference service in response to a management instruction of the inference service; and the control module 13 is configured to deploy a second container group inside the control module 13, deploy a management component inside the second container group, send a deployment instruction to the reasoning module 12 by the management component, and send a management instruction to the work component by the management component.
[0037] It should be noted that pre-training is a strategy for training a deep learning model, and its core is to preliminarily train the model using a large-scale data set, so that the model learns a general feature representation. This process is similar to the basic learning stage before a human learns new knowledge, through extensive reading, observation and experience accumulation.
[0038] The pre-trained language model generally refers to designing a language model training task based on a large-scale corpus (including, for example, language training materials such as sentences, paragraphs, etc.), training a large-scale neural network algorithm structure to learn and implement, and finally obtaining a large-scale neural network algorithm structure and parameters, which is the pre-trained language model. Subsequent other tasks can perform feature extraction or task fine-tuning based on the model to achieve specific task purposes. The idea of pre-training is to first train a task to obtain a set of model parameters, and then initialize the network model parameters using the set of model parameters, and then train other tasks using the initialized network model to obtain a model adapted to other tasks. By pre-training on a large-scale corpus, the neural language representation model can learn powerful language representation capabilities and can extract rich syntactic and semantic information from text. The pre-trained language model can provide word and sentence-level features containing rich semantic information for downstream tasks, or directly fine-tune the pre-trained model for downstream tasks, to obtain a downstream exclusive model conveniently and quickly.
[0039] The neural network algorithm structure trained by the pre-trained language model can be a CNN (Convolutional Neural Network), an RNN (Recurrent Neural Network), an LSTM (Long Short Term Memory), or the like, or a model constructed by an attention network, such as a Transformer, a BERT (Bidirectional Encoder Representations from Transformers), a GPT (Generative Pre-trained Transformer), and the like, which are not limited in the present application. The attention network refers to a network model trained by using an attention mechanism, which extracts more important feature information in an input sequence by assigning different weights to each part of the input sequence, so that the model finally obtains more accurate output.
[0040] It can be understood that the embodiment of the present application constructs an inference service management system 10, which includes a storage module 11, an inference module 12, and a control module 13. The storage module 11 stores an entry script of a work component and a model file of a pre-trained language model. The inference module 12 deploys a first container group and deploys an image container inside the first container group. The storage module 11 and the first container group are mounted. The work component is started by the entry script in the storage module 11. The decoupling of the entry script and the image container is achieved. The specific entry script refers to the business code of the image container. The image container provides an inference environment for the inference service. The decoupling of the business code and the inference environment is achieved by mounting. When the image container is expanded, the influence of the entry script on the expansion does not need to be considered. The flexible expansion of the container image is achieved. The scalability of the container image is improved. At the same time, by mounting the work component in the image container, the non-intrusive management of the work component and the image container is achieved. The work component can start the inference engine in the image container. The inference engine loads the pre-trained language model corresponding to the model file. The inference service is achieved. The work component can also manage the inference service. The control module 13 deploys a second container group. The management component is deployed inside the second container group. The management component can send a deployment instruction to the inference module 12 and send a management instruction to the work component. Through the dual management between the management component and the inference module 12 and the work component, the flexible management of the image container and the inference service is achieved. The non-intrusive and scalable management of the inference service is provided. The use experience of the user is improved.
[0041] The working component of the embodiment of the present application is a set of intelligent scripts, which enters the image container in a shared storage mounting manner, thereby realizing non-intrusive start and management of the inference engine. The entry script is the core file for starting the working component and is also the key carrier for managing the inference service. The pre-trained language model can be a large language model. The model file is the data carrier of the pre-trained language model, including the weight parameters and structure configuration of the model, and is the basis for the inference engine to generate inference services according to instructions. The inference engine is the inference core of the pre-trained language model pre-installed in the image container, such as vLLM and SGLang, which is used to load the model file and provide inference services.
[0042] The first container group of the embodiment of the present application can be a Pod on a container orchestration system, such as a Kubernetes cluster. The image container is a container instance created based on an official inference engine image, which is deployed in the first container group and is the running carrier of the inference engine and the working component. Essentially, it is a standardized running environment loaded with an inference engine environment. The first container group provides the image container with network, storage, resource and other basic configurations. The second container group is located in the control module, and the management component sends a deployment instruction to the inference module and sends a management instruction to the working component.
[0043] In addition, it should be noted that the first container group of the embodiment of the present application is a native resource object on the container orchestration system. In engineering practice, resource configuration is usually used to realize business deployment, traffic management and automatic scaling with SVC (service), ISVC (entry service) or KSVC (Knative service) resource objects. The execution entity of the final business code is located in the container group, and ultimately any of the above resource objects will use a Pod as the underlying implementation.
[0044] Further, in some embodiments of the present application, the control module 13 internally deploys a first communication interface of an interactive interface, the management component identifies a user interaction action of the interactive interface through the first communication interface, determines at least one of a deployment request and a management request of the user according to the user interaction action, generates a deployment instruction according to the deployment request, and generates a management instruction according to the management request.
[0045] The first communication interface is used to receive a model deployment or management request of a user, and the request includes request parameters, such as a model name, an inference engine type, an inference engine image name, etc.
[0046] It can be understood that the control module 13 of the embodiment of the present application deploys a first communication interface of an interactive interface, the management component identifies a user interaction action of the interactive interface through the first communication interface, and determines a deployment request and a management request of the user according to the user interaction action, and generates a corresponding control instruction according to the corresponding request, without the need for the user to perform additional operations.
[0047] Further, it needs to be explained that the management component in the embodiments of the present application can be deployed separately in the control module 13, and not deployed together with the inference module 12, thereby realizing decoupling of management and inference.
[0048] Further, in some embodiments of the present application, the management component is configured to extract deployment parameters in the deployment request, determine resource configuration and resource objects of the first container group according to the deployment parameters, and generate deployment instructions according to the resource configuration and the resource objects.
[0049] The deployment parameters can include a model name, an inference engine type, an inference engine image name, etc.; the resource objects include first container group resource objects and dependent resource objects of the first container group, and the core is the resource objects, and the dependent resource objects are necessary conditions for the inference service to run normally.
[0050] It can be understood that the management component in the embodiments of the present application can extract deployment parameters in the deployment request, and determine resource configuration and resource objects of the first container group according to the deployment parameters, and generate deployment instructions according to the resource configuration and the resource objects.
[0051] The resource configuration in the embodiments of the present application includes: passing through an environment variable starting with ICHAT_ when starting the work component, and covering the default entry point / command parameter of the image to start the instruction of the entry script, and avoiding directly starting the inference engine by default by covering the default instruction of the image, so as to manage the inference service by the work component subsequently.
[0052] The dependent resource objects in the embodiments of the present application can include service and entry rule resources, because the container group is only used to isolate resources and run business code, and the exposed port or address needs to be realized by the service and entry rule resource objects, and the above-mentioned resources should be included in the actual business, so the dependent resource objects need to be determined to ensure the smooth running of the inference service.
[0053] Specifically, the embodiments of the present application can generate resource configuration according to the deployment parameters, and can create resource objects of the first container group and dependent resource objects thereof through an API of the container orchestration system, at least including mounting the model file and the entry script of the work component to the corresponding path of the first container group.
[0054] For example, the user issues a deployment request, and the request parameters of the deployment request include:
[0055] a model name (such as "qwen3-7b");
[0056] an inference engine type (such as "vllm", "sglang");
[0057] Inference Engine Image Name (e.g., vllm / vllm-openai:v0.9.2);
[0058] Model File Path (the path of the model file mounted in the first container group);
[0059] Resource Requirements (the number of accelerator cards, the size of memory, the number of CPUs, PVC / hostPath mounting, etc.);
[0060] Service Configuration (port, number of replicas, etc.).
[0061] The management component generates corresponding deployment instructions according to these request parameters, and the deployment instructions include resource configuration and creation of the first container group, such as generating the configuration of the first container group according to the resource requirements and the service configuration, and creating the first container group.
[0062] Further, in some embodiments of the present application, the management component is configured to extract at least one of the resource requirements, the service configuration, the model name, the engine type, the image name, and the target path in the deployment parameters, determine the resource configuration of the first container group according to the resource requirements and the service configuration, and determine the resource object of the first container group according to at least one of the model name, the engine type, the image name, and the target path.
[0063] The deployment parameters can include the model name, the inference engine type, the inference engine image name, the model file path (i.e., the target path), the resource requirements, and the service configuration.
[0064] It can be understood that the embodiments of the present application can determine the resource configuration of the first container group according to the resource requirements and the service configuration, determine the resource object of the first container group according to at least one of the model name, the engine type, the image name, and the target path, explicitly define the role of different parameters, and ensure the accurate generation of the resource configuration.
[0065] Further, in some embodiments of the present application, the inference module 12 is configured to parse the resource configuration and the resource object in the deployment instructions, deploy at least one first container group inside the inference module 12 according to the resource configuration, and deploy at least one image container inside the first container group according to the resource object.
[0066] It can be understood that the inference module 12 of the embodiments of the present application can parse the resource configuration and the resource object in the deployment instructions, deploy at least one first container group inside the inference module 12 according to the resource configuration, and deploy at least one image container inside the first container group according to the resource object, and convert the abstract configuration generated by the management component into actual container deployment operations.
[0067] Further, in some embodiments of the present application, the inference module 12 parses the model name, engine type, image name and target path in the resource object, determines the model file of the pre-trained language model according to the model name, mounts the entry script and the model file under the target path, and deploys at least one image container in the first container group according to the engine type and the image name.
[0068] Wherein, mounting is to associate the shared path of the storage component with the target path of the first container group, so that the first container group can access the model file in the storage module.
[0069] It can be understood that the inference module 12 of the embodiments of the present application can parse the resource configuration and the resource object in the deployment instruction, determine the model file of the pre-trained language model according to the model name, mount the entry script and the model file under the target path, ensure that the entry script and the model file of the work component can be accessed by the first container group, and deploy at least one image container in the first container group according to the engine type and the image name, so as to realize accurate deployment of the image container.
[0070] For example, the embodiments of the present application can use hostPath to mount the entry script under the target path.
[0071] Further, in some embodiments of the present application, the inference module 12 is configured to set the entry script as the entry of the first container group, so as to start the work component through the entry script.
[0072] It can be understood that the embodiments of the present application can set the entry script of the work component as the entry of the first container group, so as to start the work component through the entry script, override the default command of the image container, without modifying the original content of the image container, and be compatible with the official standard image.
[0073] Further, in some embodiments of the present application, the work component is configured to detect the environment variable of the image container, generate a start instruction according to the environment variable, and start the inference engine in the image container by using the start instruction.
[0074] Wherein, the environment variable is the environment variable of the image container, which can be defined as a variable starting with CHAT_, and is used to pass parameters required by the work component to start the inference engine.
[0075] It can be understood that the work component of the embodiments of the present application can detect the environment variable of the image container, generate a start instruction according to the environment variable, and then start the inference engine in the image container by using the start instruction, pass parameters through the environment variable, without hard coding the start instruction, and at the same time, can generate a start instruction of the corresponding inference engine according to the engine type parameter in the environment variable, so as to adapt to different inference engines.
[0076] The working component of the embodiment of the application can read the environment variable to obtain unknown and inference engine parameters of the pre-trained language model to generate the starting instruction of the inference engine. It should be noted that the inference engine of the embodiment of the application can be used as a command line to start a relatively independent process as a service, or can be imported as a module in the Python code of the working component after the working component, and then used as a sub-process / thread inside the working component. In engineering practice, as a sub-process / thread, it can realize mechanisms such as exception management, signal processing, and loading and unloading of models.
[0077] In other words, the script of the working component of the embodiment of the application can be used independently, or can be integrated with other applications. As described above, the working component can obtain necessary information for starting the model by obtaining the environment variable starting with ICHAT_, or can start the service according to the returned parameters after sending a verification request to the management component in the environment / engine verification stage. It should be noted that the verification can be controlled to be closed through the environment variable, that is, the model loading and starting are completely realized through the parameters in the environment variable.
[0078] Further, in some embodiments of the application, the working component is configured to send the environment variable to the management component before generating the starting instruction according to the environment variable; the management component identifies an engine list in the environment variable, determines a current version of the inference engine according to the engine list, and updates the current version of the inference engine in the image container according to a target version if the current version is inconsistent with the target version.
[0079] The engine list is the name and version information of the inference engine installed in the current image container included in the environment variable; the current version is the version of the inference engine actually installed in the image container; and the target version is the version of the inference engine instructed in the user deployment request.
[0080] It can be understood that the working component of the embodiment of the application sends the detected environment variable to the management component before generating the starting instruction, the management component parses the environment variable, extracts the engine list therein, determines the current version of the inference engine, and updates the current version of the inference engine in the image container according to the target version if the current version is inconsistent with the target version. By adding the version checking and updating steps before starting the inference engine, the known version compatibility problems are filtered, and the inference service starting failure or functional abnormalities caused by version mismatch are avoided.
[0081] It should be noted that the version checking process before starting the inference engine of the application may produce a false negative for a new image, at which time the checking can be turned off.
[0082] Further, in some embodiments of the present application, the work component sends an environment verification request to the management component before generating the start instruction according to the environment variable, the management component returns the environment variable required for running, and if the environment variable required for running returned by the management component is different from the environment variable of the image container, the environment variable of the image container is updated using the environment variable required for running returned by the management component.
[0083] It can be understood that the work component of the embodiments of the present application can send an environment verification request of the image container to the management component before generating the start instruction, the management component returns the environment variable required for running, and if the environment variable required for running returned by the management component is different from the environment variable of the image container, the environment variable of the image container is updated using the environment variable required for running returned by the management component, that is, through the rule that the parameter priority of the management component is higher than that of the environment variable in the image container, the conflict between the manually configured environment variable and the parameter issued by the management component is avoided, and the parameter uniformity is ensured.
[0084] Further, in some embodiments of the present application, the control module 13 internally deploys a second communication interface of the container orchestration system, the management component calls the second communication interface of the container orchestration system to communicate with the reasoning module 12, queries the first state information of the first container group through the second communication interface of the container orchestration system and the probe, and manages the first container group in the reasoning module 12 according to the first state information.
[0085] The container orchestration system is used to manage the deployment, expansion, operation and maintenance, etc. of the container; the second communication interface is an interface for the management component to interact with the container orchestration system, which can be an interface realized through a database of the container orchestration system; the probe can be a health check tool originally of the container orchestration system, and specifically can be a liveness probe and a readiness probe; and the first state information is the running state information of the first container group, such as whether the container process is alive, whether the port is available, etc.
[0086] It can be understood that the management component of the embodiments of the present application can call the second communication interface of the container orchestration system to communicate with the reasoning module 12, and query the first state information of the first container group through the second communication interface of the container orchestration system and the probe, fully utilize the fault detection capability of the container orchestration system, manage the first container group in the reasoning module 12 according to the first state information, realize real-time monitoring of the state of the first container group, and can timely grasp the running state of the first container group and discover and handle the exception.
[0087] Further, in some embodiments of the present application, the work component detects second state information of the pre-trained language model and third state information of the inference service, and sends the second state information and the third state information to the management component; and the management component generates a management instruction according to at least one of the second state information and the third state information.
[0088] The second state information is running state information of the pre-trained language model, such as whether the model is loaded; and the third state information is state information of the inference service, such as a first word element response time, a number of running requests, and the like.
[0089] It can be understood that the working component of the embodiment of the application sends the detected second state information of the pre-trained language model and the third state information of the inference service to the management component, and generates a management instruction according to at least one of the second state information and the third state information, so as to not only monitor the container level state, but also cover the model and inference service level, improve the comprehensiveness of monitoring, avoid misoperation caused by only according to the container state, and improve the accuracy of operation.
[0090] Further, in some embodiments of the application, the management component queries the first state information of the first container group through the second communication interface of the container orchestration system and the probe, queries the second state information of the pre-trained language model in the mirror container through the heartbeat signal of the working component, determines a port readiness state of the container orchestration system according to the first state information, determines a model semantic readiness state according to the second state information, and determines an available state of the inference service according to the port readiness state and the model semantic readiness state.
[0091] The heartbeat signal is state synchronization information sent by the working component to the management component periodically, including a UUID (Universally Unique Identifier) of the current working component, model details, and a state; the model semantic readiness state is a state of whether the pre-trained language model has service capability, such as completion of weight loading, completion of preheating, availability of a completion interface, and the like; and the available state of the inference service is a state for judging whether the inference service can normally process a user request, including cold start, preheating, readiness, and degraded operation.
[0092] It can be understood that the embodiment of the application can determine the port readiness state of the container orchestration system and the model semantic readiness state of the pre-trained language model through the first state information of the first container group and the second state information of the pre-trained language model respectively, ensure that the inference service can be normally used through joint readiness judgment, avoid misjudgment of the availability of the service only according to the port state, and improve the availability of the inference service.
[0093] Further, in some embodiments of the application, when the management component identifies a completion identifier of at least one of problem repair and function expansion of the first container group, a restart instruction of the first container group is generated, and the restart instruction is sent to the working component; and the working component restarts the entry script under the target path corresponding to the first container group in response to the restart instruction.
[0094] It can be understood that when the management component of the embodiment of the present application identifies the completion of the problem repair and function expansion of the first container group, the management component sends a restart instruction of the first container group to the working component, and the working component restarts the entry script under the target path corresponding to the first container group, so that the repaired configuration or the newly added function takes effect in the container group.
[0095] Further, in some embodiments of the present application, the management component obtains the business indicators, dimension indicators and resource indicators of the first container group through the working component, generates a scale-out instruction according to at least one of the business indicators, dimension indicators and resource indicators, sends the scale-out instruction to the working component, and the working component responds to the scale-out instruction to perform a scale-out action on the first container group.
[0096] The business indicators are indicators related to the inference service business, such as request queue length, first word element response time, and word element output time; the dimension indicators are indicators of inference service running dimensions, such as the number of running requests and the number of waiting requests; and the resource indicators are indicators of resources occupied by the inference service, such as GPU (Graphics Processing Unit) memory usage, CPU / memory occupancy, etc.
[0097] It can be understood that the management component of the embodiment of the present application obtains at least one of the business indicators, dimension indicators and resource indicators of the first container group through the working component, generates a scale-out instruction according to the obtained indicators, and sends the scale-out instruction to the working component. The working component responds to the scale-out instruction to perform a scale-out action of increasing or decreasing replicas on the first container group, thereby improving resource utilization by performing scale-out based on actual service running indicators.
[0098] Further, in some embodiments of the present application, the storage module 11 is provided with an adapter, and the working component calls the adapter in the storage module 11 to convert the required parameters of the inference engine into identifiable parameters of the inference engine through the adapter.
[0099] It can be understood that the embodiment of the present application converts the required parameters of the inference engine into identifiable parameters of the inference request through the adapter, without the need for manual adaptation by the user.
[0100] Further, in some embodiments of the present application, the storage module 11 supports at least one distributed storage, and the entry script and the model file are stored in the shared storage path of the distributed storage; after the inference module 12 deploys the first container group, the shared storage path is mounted to the target path of the first container group.
[0101] The distributed storage can be NFS (Network File System), Ceph (Ceph distributed storage system), GlusterFS (Gluster File System), etc. The shared storage path is a public path for storing an entry script and a model file, and can be accessed by all nodes on the container orchestration system.
[0102] It can be understood that the storage module 11 of the embodiment of the application supports at least one distributed storage, stores an entry script and a model file through a shared storage path, and after the inference module 12 completes the deployment of the first container group, mounts the shared storage path of the distributed storage to a target path of the first container group, and mounts the shared module to the same path of all the first container groups, so as to realize file sharing in the container orchestration system, without the need to copy files for each container group, but directly mounting the shared path can use core files, and at the same time supports multiple distributed storage schemes, so that a suitable storage scheme can be selected according to the scale and demand, and the scene applicability is improved.
[0103] The architecture of the working component of the embodiment of the application is shown in Figure 2 The working component layer includes an environment detector, a state synchronizer, a backend adapter, and a configuration manager. These modules deliver environment information, state data, adaptation parameters, and configuration instructions to the SGLang backend and the vLLM backend, which are two engine adaptation layers. The adaptation layers interface and drive the SGLang and vLLM inference engines in the inference engine layer. The environment detector is responsible for scanning the environment in the container. The state synchronizer reports the service state to the management component. The backend adapter converts general configurations into engine-specific parameters. The configuration manager parses environment variables to generate startup configurations. The engine adaptation layer serves as an intermediate layer between the working component and the inference engine, ultimately supporting the loading of models by the inference engine and providing services. This architecture allows the working component to non-invasively manage multiple inference engines. Like the shared storage module of the system supporting multiple distributed storage schemes (NFS, Ceph, etc.) that can be flexibly selected according to the cluster scale and demand, the design of the backend adapter and the multi-engine adaptation layer enables flexible adaptation to different inference engines, meeting the inference service needs in different scenarios.
[0104] Specifically, the inference service management system of the embodiment of the application is described below through an embodiment. The structure of the inference service management system is shown in Figure 3 The inference service management system includes a control module, an inference module, and a storage module. The inference module includes a working component.
[0105] The management component is deployed in the container orchestration system in the form of an independent second container group as a control center of the deployment system, and has the following functions: (1) request processing: providing a RESTful API (i.e., a first communication interface) interface to receive model deployment and management requests of a user transmitted from an interactive interface; (2) container orchestration system API client: interacting with an API interface service of the container orchestration system through an official client code library to implement CRUD (Create, Read, Update, and Delete) operations of the first container group; and (3) service management: maintaining state information of all inference services in the cluster, including the state of the first container group, the state of the working component, the model loading state, and the like, and transmitting model startup configurations specified by the interactive interface to the working component.
[0106] The working component is not a traditional resident process, but an intelligent script set, and has the following functions: (1) environment detection: detecting and collecting Python and inference engine version information; (2) configuration management: reading variables starting with ICHAT_ in the environment variable and converting them into parameters required by the inference engine; (3) communication: periodically synchronizing the state to the management component; and (4) backend adapter: converting the parameters required by the inference engine into parameters recognizable by the corresponding inference engine through mapping and the like.
[0107] The working component can be used for: (1) non-intrusive: mounted in the official inference engine image in the form of an external script without modifying the original image content; (2) adaptive: capable of automatically detecting the Python environment and the installed inference engine in the container, and converting the general configuration transmitted by the interactive interface into special configuration of the corresponding inference engine; and (3) state management: real-time monitoring of the inference service state and reporting to the management component.
[0108] The shared storage adopts a distributed file system architecture, and the specific implementation includes: (1) storage backend: supporting multiple distributed storage schemes such as NFS, Ceph, and GlusterFS; (2) mounting strategy: mounting the shared storage to the same path of all nodes in the cluster to realize file sharing in the cluster; in the present application, the entry script of the working component is located in the shared storage path, and is mounted to the specified path of the first container group in the hostPath mode when the model service is deployed.
[0109] It should be noted that the container group is a container orchestration system native resource object, and in engineering practice, resource configuration is usually used to cooperate with SVC, ISVC or KSVC and other resource objects to realize business deployment, traffic management and automatic expansion and contraction, and the execution entity of the final business code is located in the container group. Regardless of the deployment of any of the above resource objects, the container group will ultimately be used as the underlying implementation, so for convenience, the container group is described as the main body in the following. The container group usually undertakes the following responsibilities in this embodiment: (1) deploying inference engine images: using any official inference engine image, such as 'vllm / vllm-openai','sglang / sglang', etc., to deploy inference services using accelerator card resources; (2) resource configuration: allocating computing resources such as accelerator cards, memory and CPU according to the specified parameters of the management component; (3) network configuration: configuring service or entry rules to realize network exposure of the service; (4) health check: integrating the native live probe and ready probe of the container orchestration system.
[0110] Based on the inference service system as shown in Figure 3 , the implementation process of the inference service management system of the embodiment of the present application is as shown in Figure 4 and Figure 5 , the time sequence diagram of the inference service management is as shown in Figure 4 , and the flowchart of the inference service management is as shown in Figure 5 .
[0111] As shown in Figure 4 and Figure 5 , the overall implementation process of the inference service management includes:
[0112] Step 1: service deployment request processing, the user initiates a deployment request to the management component through an interactive interface or a communication interface, wherein the request parameters include:
[0113] (1) model name (such as "qwen3-7b");
[0114] (2) inference engine type (such as "vllm", "sglang");
[0115] (3) inference engine image name (such as vllm / vllm-openai:v0.9.2);
[0116] (4) model file path (the path of the model file mounted in the first container group);
[0117] (5) resource requirements (number of accelerator cards, memory size, number of CPUs, PVC / hostPath mounting, etc.);
[0118] (6) service configuration (port, number of replicas, etc.).
[0119] Step 2: Generate the first container group and related resource configuration, and set the entry to the entry script of the work component.
[0120] Generate the configuration of the first container group according to the resource requirements and service configuration, wherein the required parameters of the work component startup model are passed through the ICHAT_ environment variable at the beginning, and the environment variable at least includes the UUID uniquely allocated by the management component for the work component; wherein the default entry point and command parameters in the image should be overwritten by the above configuration as the startup instruction of the work component script, so as to realize the verification and management of the inference engine through the entry script of the work component.
[0121] Step 3: Call the interface of the container orchestration system to create the first container group resource object and its dependent resource object.
[0122] Step 4: Mount the model file and the entry script of the work component to the path corresponding to the first container group.
[0123] Steps 5-8 are the initialization and model loading of the work component. After the first container group is started, the work component script executes the following initialization process.
[0124] Step 5: Environment scanning: detect the Python environment, CUDA (Compute Unified Device Architecture) version, installed inference engine and its version in the image container.
[0125] Among them, the Python environment refers to the main business container, that is, vLLM, SGLang, Text Embedding Inf (Text Embedding Inference) and the like.
[0126] Step 6: Whether the environment variable needs to be verified, if yes, go to step 7, otherwise go to step 9.
[0127] Step 7: Compatibility verification: send a request to the management component to verify whether the current environment meets the running requirements of the specified inference engine, and report an error and exit if it does not meet the requirements. The verification performed in this step can filter known abnormal scenarios, but for some new images, it may produce a false negative, in which case the verification can be turned off.
[0128] The specific verification process includes: in a layer-by-layer progressive manner, taking vLLM as an example, first trying to import vLLM, failing to infer that vLLM does not exist, and succeeding to proceed to the next step to try to obtain vllm._version_ to determine the version information. The verification of the management component is performed as needed: first, the engine comparison is performed, such as the user requiring to use SGLang, but only vLLM exists at present, which is judged to fail; second, the version matching is performed, in the known version, different versions support different models, and optionally if the user selects the HuggingFace model name (the standard name of the model distribution platform), the matching can be performed, whether the current version of the inference engine meets the requirements, and if there is a problem, an error is prompted in time. If the user does not specify the HuggingFace model name, or the current vLLM version is newer and is not in the management component database, it is defaulted to pass and considered to be able to be deployed.
[0129] Step 8: Determine whether to continue execution, yes to step 9, otherwise end the process.
[0130] Step 9: Try to load the model according to the environment variable and the return body parameter.
[0131] The environment variable is read to obtain the model position and inference engine parameter, the inference engine start command is constructed, the model is tried to be loaded and the service is started.
[0132] It should be noted that vLLM itself can be started as a relatively independent process as a service by the above-mentioned instruction as a command line, or can be used as a module in the Python code of the work component after being imported by import, and is used as a sub-process / thread inside the work component. In engineering practice, as a sub-process / thread, it can realize the mechanisms of exception management, signal processing and model loading and unloading. The above-mentioned two schemes can complete the function, and the command line scheme is adopted for description in the present application for the convenience of understanding.
[0133] Step 10: Determine whether the server is started successfully, yes to step 11, otherwise end the process.
[0134] Step 11: State reporting, after the service is started, the current management component UUID, model details and state are sent to the management component through the heartbeat mechanism, and the ready state is synchronized.
[0135] Among them, the management component performs joint readiness and continuous monitoring mechanism.
[0136] Step 12: The container orchestration system layer is ready, the state of the first container group is queried through the readiness, inventory probe and second communication interface, it is confirmed that the container process and service port are available, that is, it is judged whether the state of the first container group is ready, ready to step 13, otherwise rejudge.
[0137] Step 13: Model semantics is ready, parse workgroup component event / low-frequency heartbeat reporting, confirm that weight loading is completed, model preheating is completed, completion interface is available, determine that the service is available, and end the process.
[0138] The joint condition of port readiness and model semantics readiness is used as the service availability threshold, and the cold start, preheating, readiness, and degraded running states are presented in the interactive interface.
[0139] In addition, the embodiment of the present application can periodically collect, subscribe to the first word element response time, the number of running requests, the number of waiting requests, and the word element output time, which are used for observability and scaling coordination (such as HPA (Horizontal Pod Autoscaler, horizontal automatic expander) / KEDA (Kubernetes Event-driven Autoscaling, event-based automatic expander)).
[0140] Compared with the first container group information obtained by the container orchestration system, the management component can obtain the environment information, inference engine information, model and service state information inside the container through the work component, and can realize more detailed information collection and display. The / metrics interface exposed by the work component can provide indicators including the first word element response time, the number of running requests, the number of waiting requests, and the word element output time for scaling.
[0141] Based on the above introduction, the main improvements of the embodiment of the present application include:
[0142] (1) Code and engine loose coupling: the work component starts the inference engine through a relatively stable and unchanged command / module method, which separates the business code and the inference engine, and supports detecting the environment in the container, which facilitates the connection and update of the latest inference engine image of the official, and also can update the entry script of the work component located in the hostPath by restarting the service, and quickly complete problem repair and function expansion.
[0143] (2) Dynamic dependency detection: the script of the work component has a dependency detection function. After the first container group is started, the script will first detect the installed inference engine (such as vLLM, SGLang) and its version in the container environment, and report the available engine list to the management component.
[0144] (3) Unified work component entry: the startup script of the work component is the constant entry point of the service. The script is responsible for environment detection, heartbeat communication with the management component, receiving instructions, and finally calling the actual inference engine startup command.
[0145] (4) Container orchestration system and work component double-layer management: The management component manages services through two dimensions: container orchestration system layer: through the service interaction with the container orchestration system communication interface, the life cycle of the first container group is managed, the deployment, scaling and fault self-healing of the service are realized; work component layer: after the successful operation of the first container group, the work component script in the first container group is called through the second communication interface to realize the environment detection, parameter configuration and state synchronization of the model to complete the fine control.
[0146] (5) Joint readiness determination method: the joint condition of the port readiness of the container orchestration system and the model semantic readiness of the work component is used as the service availability threshold to avoid misjudgment of business availability only by the port state.
[0147] (6) Non-load balancing orchestration and scaling coordination: the management component does not do request routing and load balancing, and the queue length, delay and GPU memory usage (etc. Indexes drive HPA / KEDA scaling, forming a technical boundary with the controller module.
[0148] Therefore, the inference service system of the present application can realize:
[0149] (1) Compatible standard image: No longer need to build customized images for each backend. By mounting the entry script of the work component as hostPath into the container, and setting it as the entry of the container (command parameters or parameter list correspond to the entry point and command parameters of the container), non-intrusive management of any official inference image is realized.
[0150] (2) The current LLM management framework can integrate and fully utilize the life cycle management, traffic management and automatic scaling of cloud native, etc. It is compatible with other cloud native frameworks such as kserve, knative, istio, etc. As long as it does not involve service forced start command or cannot mount hostPath framework / resource object, it can be used.
[0151] (3) Dynamic capability declaration and adaptation: the work component detects the available inference engine and version in the container at runtime to form a capability declaration (such as whether to support tensor parallel, maximum available model length, etc.), and the management component generates a resource and parameter set accordingly, reducing the misfit and image reconstruction cost.
[0152] (4) Joint readiness and index-driven scaling: improve business availability through joint determination of model semantic readiness and port readiness, and drive HPA / KEDA scaling with queue length, delay and GPU memory usage, etc. The management component does not bear the responsibility of request routing or load balancing.
[0153] The inference service management system provided by the embodiment of the present application comprises a storage module, an inference module and a control module. The storage module stores an entry script of a work component and a model file of a pre-training language model. The inference module deploys a first container group and deploys an image container inside the first container group. The storage module and the first container group are mounted. The work component is started through the entry script in the storage module. The decoupling of the entry script and the image container is realized. The specific entry script refers to the business code of the image container. The image container provides an inference environment of the inference service. The decoupling of the business code and the inference environment is realized through the mounting mode. When the image container is expanded, the influence of the entry script on the expansion does not need to be considered. The flexible expansion of the container image is realized. The expandability of the container image is improved. Meanwhile, the work component and the image container can be managed in a non-invasive manner through the mounting of the work component in the image container. The work component can start an inference engine in the image container. The inference engine loads a pre-training language model corresponding to the model file. The inference service is realized. The work component can also manage the inference service. The control module deploys a second container group. The management component is deployed inside the second container group. The management component can send a deployment instruction to the inference module and send a management instruction to the work component. Through the double management between the management component and the inference module and the work component, the flexible management of the image container and the inference service is realized. Therefore, the non-invasive and expandable management of the inference service is provided. The use experience of the user is improved.
[0154] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, can also be realized by hardware, but in many cases, the former is a better embodiment.
[0155] The embodiment of the present application also provides an electronic device comprising the inference service management system as above.
[0156] The embodiment of the present application also provides an inference service management method applied to the management component in the inference service management system as above.
[0157] Figure 6 The inference service management method provided by the embodiment of the present application is shown in a flowchart.
[0158] As shown in Figure 6 The inference service management method comprises the following steps S101-S104.
[0159] In step S101, the user interaction action of the interactive interface is identified.
[0160] The management component can identify the interactive action of the user through the first communication interface, and the interactive action is a creating, querying, modifying or the like operation performed by the user on the interactive interface.
[0161] In step S102, at least one of a deployment request and a management request of the user is determined according to the user interactive action, a deployment instruction is generated according to the deployment request, and a management instruction is generated according to the management request.
[0162] The deployment or management request includes a request parameter, such as a model name, an inference engine type, an inference engine image name, and the like.
[0163] In step S103, the deployment instruction is sent to the inference module, the inference module is started in response to the deployment instruction, at least one first container group is deployed in the inference module, at least one image container is deployed in the first container group, the storage module is mounted between the first container group, the work component is started through the entry script, the work component is mounted in the image container, the inference engine in the image container is started through the work component, the pre-trained language model corresponding to the model file is loaded through the inference engine, and the inference service is provided through the pre-trained language model.
[0164] The first container group can be a Pod on a container orchestration system; the image container is a container instance created based on an official inference engine image, is deployed in the first container group, and is a running carrier of the inference engine and the work component, and is essentially a standardized running environment loaded with an inference engine environment. The first container group provides network, storage, resource and other basic configurations for the image container.
[0165] It can be understood that the storage module can be mounted between the first container group, the work component is started through the entry script in the storage module, the entry script and the image container are decoupled, the specific entry script is the business code of the image container, the image container provides the inference environment of the inference service, the business code and the inference environment are decoupled through the mounting mode, when the image container is expanded, the influence of the entry script on the expansion is not considered, the container image is flexibly expanded, and the expandability of the container image is improved; at the same time, the work component is mounted in the image container, the non-intrusive management component between the work component and the image container is realized, the work component can start the inference engine in the image container, the inference engine loads the pre-trained language model corresponding to the model file, the inference service is realized, and the work component can also manage the inference service.
[0166] In some embodiments of the present application, before starting the inference engine in the mirror container through the work component, further comprising: obtaining a verification request of an environment variable of the mirror container; sending the verification request of the environment variable to the management component, and the management component verifies the environment variable in response to the verification request of the environment variable, and if the environment variable meets the preset running condition of the inference engine, a verification success result is generated and fed back to the work component; if the verification success result is received, the generation action of the start instruction is executed.
[0167] The environment variable is a system environment variable of the inference engine mirror container in the first container group, and the variable starting with ICHAT_ is injected by the management component when generating the container group configuration and is a core parameter carrier for starting the inference engine by the work component. The preset running condition is a condition that meets the configuration parameters of the mirror container for starting the inference engine.
[0168] It can be understood that, before starting the inference engine in the mirror container by the work component, the embodiments of the present application obtain the environment variable of the mirror container and generate a verification request, which is sent to the management component. The management component verifies the integrity, format correctness and matching with the inference engine of the environment variable in response to the request. If the preset running condition is met, a verification success result is fed back. After the work component receives the result, the generation action of the start instruction is executed. This avoids the inference engine from failing to start due to missing environment variables, format errors or mismatch with the engine, and improves the deployment success rate.
[0169] In step S104, the management instruction is sent to the work component, and the inference service is managed by the work component in response to the management instruction of the inference service.
[0170] It can be understood that, in the embodiments of the present application, the management instruction can be sent to the work component, and the inference service is managed by the work component in response to the management instruction of the inference service. Through the double management between the management component and the inference module and the work component, flexible management of the mirror container and the inference service is realized, and non-intrusive and extensible management of the inference service is provided, thereby improving the user experience.
[0171] It should be noted that the features of the embodiments corresponding to the inference service management method can be referred to the related description of the embodiments corresponding to the inference service management system, which will not be repeated here.
[0172] According to the inference service management method provided by the embodiment of the present application, the storage module is mounted between the first container group, the working component is started through the entry script in the storage module, the decoupling of the entry script and the image container is achieved, the specific entry script refers to the business code of the image container, the image container provides the inference environment of the inference service, the decoupling of the business code and the inference environment is achieved through the mounting mode, when the image container is expanded, the influence of the entry script on the expansion does not need to be considered, the flexible expansion of the container image is achieved, and the scalability of the container image is improved; meanwhile, by mounting the working component in the image container, the non-intrusive management component between the working component and the image container can be achieved, the working component can start the inference engine in the image container, the inference engine loads the pre-training language model corresponding to the model file, the inference service is achieved, the working component can also manage the inference service, sends the management instruction to the working component, and the management instruction of the inference service is responded through the working component to manage the inference service, so that through the double management between the management component and the inference module and the working component, the flexible management of the image container and the inference service is achieved, and then the non-intrusive and scalable management of the inference service is provided, and the use experience of the user is improved.
[0173] The implementation process of the inference service management method of the embodiment of the present application is described below through a specific embodiment, including:
[0174] Step 1: The user sends a model deployment request to the management component through an interactive interface, and the request carries (such as "qwen3-7b"), the inference engine type (such as "vLLM", "SGLang"), the inference engine image name (such as vllm / vllm-openai:v0.9.2), the model file path (the path of the model file mounted in the first container group), the resource requirement (the number of accelerator cards, the size of memory, the number of CPUs, PVC / hostPath mounting, etc.), and the service configuration (port, number of replicas, etc.) information.
[0175] Step 2: After receiving the request, the management component generates the container orchestration system configuration of the first container group according to the resource requirement and the service configuration, and also needs to generate the container orchestration system configuration of the PVC and the storage class resource object when using the PVC, and so on. The parameters required for starting the model by the working component are transmitted through the ICHAT_ starting environment variable, and the environment variable at least includes the UUID uniquely allocated by the management component for the working component, which is used for the management component to uniquely determine the working component service. The first container group execution entry is configured as the path where the working component script is located, the vLLM and SGLang inference engines are managed through the working component entry script; finally, the resource and other dependent resource objects of the first container group are created through the communication interface of the container orchestration system using the above configuration.
[0176] Step 3: The container orchestration system schedules the current first container group to the computing node according to the resource demand and usage, and when the node does not have the inference engine image specified in the request, the image will be pulled through the network. In this application, the inference image only requires to contain the inference engine supported by the backend adaptation layer in the work component script, and the official vLLM and SGLang images in the Docker image repository meet the requirements and can be easily obtained and updated to the latest version as needed. Then the model file and the work component entry script are mounted into the first container group according to the configuration, and the work component entry script is started.
[0177] Step 4: The work component script performs an initialization process. First, the Python environment, CUDA version, installed inference engine and its version in the container are detected; a request is sent to the management component to verify whether the current environment meets the running requirements of the specified inference engine, and if not, an error is reported and the process is exited (this step is an optional verification step). For example, when the interactive interface specifies to use vLLM as the inference engine, but uses the official SGLang image or other images that do not have the vLLM inference engine, an error message should be prompted and the process is exited. Further, when the image does not have the Python environment required for the work component entry script to execute, the work component will not send a request to the management component, and the unready state of the first container group will be maintained in the management component, prompting the user that there is a more basic environment problem in the image. The model location and inference engine parameters are obtained by reading the environment variables / management component return parameters (the return parameters have higher priority than the environment variables, and the work component will take the union of the two and follow the management component to avoid conflicts), the inference engine start command is constructed, the model is loaded, and the service is started. After the service is started, the UUID, model details and state of the current work component are sent to the management component through the heartbeat mechanism.
[0178] Step 5: After the model is loaded and ready, the work component periodically sends a heartbeat request to update the work component state recorded in the management component. When the management component does not receive a heartbeat request within a specified event interval, it considers that the work component is abnormal. The management component periodically detects the state of the first container group by calling the second communication interface of the container orchestration system, and considers that the first container group is ready when all containers of the first container group are ready. When the first container group is ready and the work component heartbeat request is ready, the service state is updated to ready in the interactive interface, and the service is ensured to be normally available through double verification. Then the first container group and the work component heartbeat state are continuously monitored, and the state changes can be viewed in the interactive interface.
[0179] The embodiments of the present application also provide a non-volatile computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above inference service management method embodiments when running.
[0180] In an example embodiment, the non-volatile computer-readable storage medium described above can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0181] Embodiments of the present application also provide a computer program product, which comprises a computer program. When the computer program is executed by a processor, the steps in any of the inference service management method embodiments described above are implemented.
[0182] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0183] The above provides a detailed description of the inference service management system, device, management method, medium and program product provided by the present application. The principles and implementation modes of the present application are described by applying specific examples in this paper. The above description of the examples is only to help understand the method of the present application and its core idea. It should be noted that for ordinary skilled persons in the technical field, without departing from the principles of the present application, the present application can be improved and modified in several ways. These improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A reasoning service management system, characterized in that, include: The storage module is used to store the entry scripts of the working components and the model files of the pre-trained language models; The inference module is used to respond to the deployment instructions of the pre-trained language model, deploy at least one first container group within the inference module, deploy at least one image container within the first container group, mount the storage module to the first container group, and start the working component through the entry script; the inference module is used to parse the resource configuration and resource objects in the deployment instructions, deploy at least one first container group within the inference module according to the resource configuration, and deploy at least one image container within the first container group according to the resource objects; The inference module parses the model name, engine type, image name and target path in the resource object, determines the model file of the pre-trained language model based on the model name, mounts the entry script and the model file to the target path, and deploys at least one image container in the first container group based on the engine type and image name. The working component is mounted in the image container. The working component starts the inference engine in the image container. The inference engine loads the pre-trained language model corresponding to the model file to provide inference services through the pre-trained language model. The working component responds to the management instructions of the inference service and manages the inference service. The control module has a second container group deployed within it, and a management component deployed within the second container group. The management component sends the deployment instructions to the inference module and the management instructions to the working component.
2. The reasoning service management system according to claim 1, characterized in that, The control module is equipped with a first communication interface for an interactive interface. The management component identifies user interaction actions on the interactive interface through the first communication interface, determines at least one of the user's deployment request and management request based on the user interaction actions, generates the deployment instruction based on the deployment request, and generates the management instruction based on the management request.
3. The reasoning service management system according to claim 2, characterized in that, The management component is used to extract deployment parameters from the deployment request, determine the resource configuration and resource objects of the first container group based on the deployment parameters, and generate the deployment instructions based on the resource configuration and the resource objects.
4. The reasoning service management system according to claim 3, characterized in that, The management component is used to extract at least one of the resource requirements, service configuration, model name, engine type, image name, and target path from the deployment parameters, determine the resource configuration of the first container group based on the resource requirements and the service configuration, and determine the resource object of the first container group based on at least one of the model name, the engine type, the image name, and the target path.
5. The reasoning service management system according to any one of claims 1-4, characterized in that, The inference module is used to set the entry script as the entry point of the first container group, so as to start the working component through the entry script.
6. The reasoning service management system according to claim 1, characterized in that, The working component is used to detect the environment variables of the image container, generate a startup command based on the environment variables, and use the startup command to start the inference engine within the image container.
7. The reasoning service management system according to claim 6, characterized in that, The working component sends the environment variables to the management component before generating the startup command based on the environment variables; the management component identifies the engine list in the environment variables, determines the current version of the inference engine based on the engine list, and updates the current version of the inference engine in the image container according to the target version if the current version is inconsistent with the target version.
8. The reasoning service management system according to claim 1, characterized in that, The control module is equipped with a second communication interface of the container orchestration system. The management component calls the second communication interface of the container orchestration system to communicate with the inference module, queries the first status information of the first container group through the second communication interface of the container orchestration system and the probe, and manages the first container group in the inference module according to the first status information.
9. The reasoning service management system according to claim 1 or 8, characterized in that, The working component detects the second state information of the pre-trained language model and the third state information of the inference service, and the working component sends the second state information and the third state information to the management component. The management component generates the management instruction based on at least one of the second status information and the third status information.
10. The reasoning service management system according to claim 1, characterized in that, The management component queries the first status information of the first container group through the second communication interface and probe of the container orchestration system, queries the second status information of the pre-trained language model in the image container through the heartbeat signal of the working component, determines the port readiness status of the container orchestration system based on the first status information, determines the model semantic readiness status based on the second status information, and determines the availability status of the inference service based on the port readiness status and the model semantic readiness status.
11. The reasoning service management system according to claim 1, characterized in that, When the management component recognizes a completion flag for at least one of the problem fixes and feature extensions of the first container group, it generates a restart command for the first container group and sends the restart command to the working component; the working component responds to the restart command to restart the entry script under the target path corresponding to the first container group.
12. The reasoning service management system according to claim 1, characterized in that, The management component obtains the business metrics, dimension metrics, and resource metrics of the first container group through the working component, generates a scaling instruction based on at least one of the business metrics, dimension metrics, and resource metrics, and sends the scaling instruction to the working component. The working component responds to the scaling instruction to perform scaling operations on the first container group.
13. The reasoning service management system according to claim 1, characterized in that, The storage module is equipped with an adapter. The working component calls the adapter in the storage module to convert the required parameters of the inference engine into recognizable parameters of the inference engine.
14. The reasoning service management system according to claim 1 or 13, characterized in that, The storage module supports at least one distributed storage system, and stores the entry script and the model file through the shared storage path of the distributed storage system; after the inference module completes the deployment of the first container group, it mounts the shared storage path to the target path of the first container group.
15. An electronic device, characterized in that, Including the reasoning service management system as described in any one of claims 1-14.
16. A reasoning service management method, characterized in that, The method is applied to a management component in a reasoning service management system as described in any one of claims 1-14, wherein the method includes the following steps: Identify user interaction actions on the interactive interface; Based on the user interaction action, determine at least one of the user's deployment request and management request, generate a deployment instruction based on the deployment request, and generate a management instruction based on the management request; The deployment command is sent to the inference module, which responds to the deployment command by deploying at least one first container group within the inference module, deploying at least one image container within the first container group, mounting the storage module to the first container group, starting the working component through the entry script, mounting the working component in the image container, starting the inference engine within the image container through the working component, and loading the pre-trained language model corresponding to the model file through the inference engine to provide inference services through the pre-trained language model. The management command is sent to the working component, which then responds to the management command of the inference service to manage the inference service.
17. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the inference service management method as described in claim 16.
18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the inference service management method as described in claim 16.
Citation Information
Patent Citations
Online model reasoning system
CN111414233A
Method and system for constructing machine learning model automated production line
WO2023071075A1