Service starting and control method, electronic equipment, storage medium and program product

By recording the processor process and creating a startup scheduling file, the startup process of the inference service is optimized, and the problem of long startup time of the inference service is solved, the startup efficiency and performance are improved, and the waste of GPU resources is reduced.

CN120353516AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510858005.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In the prior art, the startup time of the inference service is long, which affects its performance, especially when the model file is large, the model loading time is long, resulting in waste of GPU resources and the inability to quickly alleviate the pressure of service access.

Method used

By recording the processor process, forming a running process file, creating an adaptive startup scheduling file, omitting the steps to deploy the processor, using snapshot mirroring and startup scheduling files to quickly start the inference service, and optimizing the startup process.

Benefits of technology

It improves the startup efficiency and performance of inference services, reduces startup time, ensures sufficient utilization of GPU resources, and optimizes the performance of computing clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353516A_ABST
    Figure CN120353516A_ABST
Patent Text Reader

Abstract

The invention discloses a service starting and control method, electronic equipment, a storage medium and a program product, and relates to the technical field of cloud services. The service starting method comprises the following steps: after a current inference service runs, recording a processor process related to the deployment of the current inference service in a computing cluster to form a running process file, and taking the current inference service as a target inference service capable of improving the starting efficiency. If yes, when the starting request of the first reasoning service belonging to the target reasoning service is obtained, the running process file can be written in so as to omit the steps of reasoning a processor needing to be deployed and the like. Therefore, the technical problem that the performance of the inference service is influenced by long inference service starting time is solved, and the technical effects that the inference service starting process is optimized, the inference service restarting efficiency can be improved, and the inference service performance can be improved are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cloud services, and particularly to a service startup method, a service control method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] An important component of cloud services is the inference service, which refers to a cloud service that deploys a trained inference model online to perform inference tasks such as real-time prediction. However, due to the complexity of the inference model architecture of the inference service and the large size of the model file, it takes a long time to start or reload the inference service each time, thereby affecting the performance of the inference service. Summary of the Invention

[0003] This application provides a service startup method, a service control method, an electronic device, a computer-readable storage medium, and a computer program product to at least solve the problem in the related art that the long startup time of the inference service affects its performance.

[0004] This application provides a service startup method, which includes: in response to the current inference service running, obtaining the processor process when the current inference service runs in the computing cluster, and recording the processor process to form a running process file; creating an executable file adapted to the running process file as a startup scheduling file; wherein, the startup scheduling file is used to start the current inference service; taking the current inference service that has completed the creation of the startup scheduling file as the target inference service; in response to obtaining a startup request for the first inference service, executing its startup scheduling file and writing it into the running process file to start the first inference service; wherein, the first inference service belongs to the target inference service.

[0005] This application also provides a service control method, which includes: in response to the first run of the current inference service, determining whether the current inference service runs successfully; in response to the current inference service running successfully, executing the service startup method as described above; scheduling the current inference service to execute an inference task.

[0006] This application also provides an electronic device, which includes: a memory and a processor; the memory is used to store a computer program; the processor is used to implement the steps of the service startup method as described above when executing the computer program; or, implement the steps of the service control method as described above.

[0007] This application also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the service startup method as described above; or, implements the steps of the service control method as described above.

[0008] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the service startup method described above; or, implements the steps of the service control method described above.

[0009] Through the present application, since the processor process of the target inference service is identified in advance and a running process file is formed. When obtaining a startup request for the first inference service belonging to the target inference service, the running process file can be written to omit steps such as the processor to be deployed for inference. Therefore, the technical problem that the startup duration of the inference service is too long and affects its performance can be solved, and the startup process of the inference service can be optimized, which is beneficial to improving the restart efficiency of the inference service, and further can improve the technical effect of the performance of the inference service. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a schematic diagram of a scenario of an embodiment of the application scenario of the service startup method of the present application; Figure 2 It is a schematic flowchart of an embodiment of the service startup method of the present application; Figure 3 It is a schematic flowchart of an embodiment of running the current inference service of the present application; Figure 4 It is a schematic flowchart of an embodiment of expanding the inference service of the present application; Figure 5 It is a comparison schematic diagram of an embodiment of the startup of the conventional inference service and the startup of the target inference service of the present application; Figure 6 It is a schematic flowchart of an embodiment of the service control method of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0013] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0015] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the service startup method depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0016] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the scenario of an embodiment of Application Scenario 1 of the service startup method of this application.

[0017] In this embodiment, the application scenario of the service startup method may include a computing cluster, a service management device, a first storage unit, and a second storage unit.

[0018] Among them, the computing cluster can run the inference service.

[0019] The computing cluster can be considered as a system composed of multiple interconnected computers, and each computer can be considered as a computing node. For example, the computing node can be a server, etc. The computing node can include a processor capable of supporting operations. For example, the processor type can include at least one of CPU (Central Processing Unit), GPU (Graphics Processing Unit), and acceleration card (i.e., acceleration processor). That is to say, when the processor type includes CPU, GPU, and acceleration processor, the service startup method of this application can be compatible with the deployment details of the inference service on the processors of the foregoing processor types, thereby improving the functionality of service startup management and further improving the startup efficiency of the inference service.

[0020] The inference service can be understood as a process of deploying a trained inference model to a computing cluster to perform inference tasks such as real-time prediction. It enables the application of machine learning or deep learning models to practical problems. The inference model obtains the newly input data and makes inferences based on this new data to output prediction results or decision-making suggestions. The inference service can be natural language processing, image recognition, recommendation systems, autonomous driving, etc., and is not specifically limited here.

[0021] The service management device can control the startup of the inference service in the computing cluster and further control the inference service to perform inference tasks. That is to say, the service management device can implement the steps of the service startup method or the steps of the service control method. The service startup method and the service control method will be elaborated in detail later and will not be repeated here. In other words, the service management device can have the ability to manage the lifecycle of the inference service and can be used to implement the creation, expansion, and running of the inference service from an inference file, etc.

[0022] The service management device can control the deployment of the current inference service to the computing cluster. Here, the current inference service refers to the inference service that is currently being deployed or started. After the current inference service is successfully deployed and reliably started, a running process file and a startup scheduling file of the current inference service can be created, and the running process file can be stored in the first storage unit, and the startup scheduling file can be stored in the second storage unit.

[0023] For example, when the running process file is in the format of an image file, the first storage unit can be an image file library, etc., which is not limited here. The second storage unit can be a database, etc., which is not limited here.

[0024] It can be seen that the embodiments of the present application provide a service startup method. The working principle of the service startup method will be described in detail below in combination with the execution process of the service startup method.

[0025] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the service startup method of the present application.

[0026] S101: In response to the running of the current inference service, obtain the processor process when the current inference service is running in the computing cluster, and record the processor process to form a running process file.

[0027] In this embodiment, as described above, the current inference service refers to the inference service that is currently being started or managed for startup.

[0028] It is possible to monitor whether the current inference service is running. Herein, the running of the current inference service means that the current inference service has been started and can be used to execute inference tasks. Optionally, this can be performed when the current inference service runs for the first time in the computing cluster or has been started and run multiple times. This is not limited herein.

[0029] When it is monitored that the current inference service is running, the processor process when the current inference service runs in the computing cluster can be obtained, and the obtained processor process can be recorded to form the running process file of the current inference service, so as to record the processor process on which the current inference service can run, so that when the current inference service is started again later, it is not necessary to immediately perform inference on the processor on which it can be deployed, which is beneficial to improving the startup efficiency of the inference service.

[0030] S102: Create an executable file adapted to the running process file as the startup scheduling file; wherein, the startup scheduling file is used to start the current inference service.

[0031] In this embodiment, in response to the completion of the creation of the running process file, it can be considered that the efficiency of starting the inference service using the original file for starting the current inference service, i.e., the initial startup file, is relatively low. Therefore, an executable file adapted to the running process file can be created as the startup scheduling file, so that when the current inference service is started later, the startup scheduling file can be used to start the inference service, thereby improving its startup efficiency.

[0032] S103: Take the current inference service for which the startup scheduling file creation is completed as the target inference service.

[0033] In this embodiment, in response to the completion of the creation of the startup scheduling file of the current inference service, it can be considered that the efficiency of redeploying and starting the current inference service can be improved. Therefore, the current inference service is taken as the target inference service.

[0034] It should be noted that the target inference service can refer to a type of inference service defined in this article in this article, and no more special meanings are given. The target inference service and the conventional inference service elaborated in detail later are relative concepts. The target inference service is an inference service that can use the startup scheduling file to improve its startup efficiency, and the conventional inference service is an inference service that is started using its initial startup file.

[0035] S104: In response to obtaining the startup request of the first inference service, execute its startup scheduling file and write it into the running process file to start the first inference service; wherein, the first inference service belongs to the target inference service.

[0036] In this embodiment, when a startup request for a first inference service belonging to a target inference service is obtained, it can be considered that the startup efficiency can be improved by using the running process file and the startup scheduling file. Therefore, the scheduling file of the first inference service can be executed and written into the running process file to achieve the startup of the first inference service.

[0037] For example, the startup request may include one or more of a request for restarting the first inference service, a request for expanding the first inference service (the original first inference request is running and a new first inference service is created), and a request for deploying the first inference service elsewhere, which is not limited herein.

[0038] That is to say, in this embodiment, after the current inference service runs, the processor processes involved in deploying the current inference service to the computing cluster are recorded to form a running process file, and the current inference service is used as the target inference service that can improve the startup efficiency. In other words, in this embodiment, the processor processes of the target inference service can be identified in advance and a running process file can be formed. In this way, when a startup request for a first inference service belonging to the target inference service is obtained, it can be written into the running process file to omit steps such as inferring the processors to be deployed. Therefore, the startup process of the inference service can be optimized to facilitate improving the startup efficiency when expanding the inference service, and thus the performance of the inference service can be improved.

[0039] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an embodiment of running the current inference service in this application.

[0040] S201: Obtain the initial information of the current inference service.

[0041] In this embodiment, the model file and the inference framework image of the current inference service can be obtained as the initial information.

[0042] Generally speaking, taking executable files such as the initial startup file and the startup scheduling file as an example. When the current inference service is first created, the user can upload the model file and the inference framework image of the current inference service to the computing cluster and the service management device. In this way, the model file and the inference framework image can be used as the initial information.

[0043] S202: Build an initial startup file based on the initial information and create the current inference service.

[0044] In this embodiment, the processor type, the number of each type of processor, the scaling index, the container image, and the startup instruction indicated by the initial information can be parsed as the resource requirement information of the current inference service; wherein, the processor type includes at least one of a central processing unit, a graphics processing unit, and an acceleration processing unit.

[0045] An initial startup file can be generated by combining the initial information and the resource requirement information to obtain the initial startup file of the current inference service.

[0046] Generally speaking, the resource requirements of the inference service are determined according to the initial information, and the deployment script of the current inference service, that is, the initial startup file deployscripts_defalut, is adaptively generated. The initial startup file can include setting the number of CPUs, the number of GPUs, scaling metrics, container image inferenceImage_default, startup command startCommand_default, etc.

[0047] Furthermore, the initial startup file can be run to load the inference model on which the current inference service depends, so as to create the current inference service in the computing cluster.

[0048] Specifically, the processor type and the number of each type of processor indicated by the initial startup file can be read as the target processor and its target number. The target number of target processors in the computing cluster is selected as the associated processors. The current inference service is deployed on the computing nodes carrying the associated processors. A running process file is formed from the processor processes of the current inference service and the associated processors. Generally speaking, in this way, a computing node with an appropriate number of GPUs and CPUs can be selected for the current inference service to support the running of the current inference service.

[0049] S203: Monitor the deployment status of the current inference service.

[0050] In this embodiment, during the process of creating the current inference service, the deployment status of the current inference service can be monitored to evaluate whether the current inference service is running.

[0051] Optionally, the loading status of the inference model can be monitored to characterize the deployment status of the current inference service by using the loading status of the inference model. Or, the deployment status monitoring can be achieved by means of feedback information after deployment, identifying configurations based on configuration files, etc., which will not be elaborated here.

[0052] S204: Determine whether the current inference service is running.

[0053] In this embodiment, when it is determined that the current inference service is running, step S205 can be executed. When it is determined that the current inference service is running, step S204 can be continued to be executed.

[0054] Optionally, it can be determined whether the current inference service is in a ready state based on the loading status of the inference model.

[0055] S205: Obtain the processor processes of the current inference service running in the computing cluster.

[0056] In this embodiment, the computing cluster may include one or more computing nodes. In this embodiment, the number of computing nodes in the computing cluster is multiple as an example.

[0057] At the same time, in this embodiment, in response to the current reasoning service being in a ready state, it can be determined that the current reasoning service is running, and the current reasoning service is not loaded to the load balancer temporarily.

[0058] That is to say, in this embodiment, when the current reasoning service is considered to be in a ready state, the current reasoning service is controlled not to provide reasoning services to the outside world temporarily, and a running process file of the processor process is created for the current reasoning service. In this way, the acquisition of additional processor processes that may exist when the current reasoning service performs reasoning tasks can be reduced, which is conducive to ensuring the simplicity of the running process file, so as to reduce the formation of redundant processor processes when the current reasoning service is started using the running process file and the startup scheduling file, and further improve the startup efficiency of the current reasoning service. At the same time, it is also possible to reduce the impact of the reasoning performance of the current reasoning service during the formation of the running process file, so as to improve the startup efficiency of the current reasoning service while taking into account the reasoning performance when it performs reasoning tasks, so as to reliably improve the service reasoning and computing cluster performance. For example, by calling the computing cluster API (Application Programming Interface, application programming interface) and other methods, the processor process snapshot of the container (i.e., the current reasoning service) can be obtained as a running process file, and when the running process file is created, the current reasoning service is added to the load balancer to be scheduled by the load balancer to provide reasoning services to the outside world, that is, the reasoning task can be executed.

[0059] S206: Determine whether obtaining the processor process is successful.

[0060] In this embodiment, when the processor process is successfully acquired, step S207 may be executed. When the processor process is unsuccessful, step S210 may be executed.

[0061] S207: Generate an operation process file and store it in the first storage unit.

[0062] In this embodiment, the processor process may be recorded to form an operating process file.

[0063] Specifically, the computing nodes included in the computing cluster can be obtained in advance, and a process inspection component can be deployed to the computing nodes to control the container management component and the node communication component of the computing nodes to enable the container snapshot function. For example, the process inspection component can be designed independently; it can also be a pre-existing process inspection component. For example, the process inspection component can be CRIU (Checkpoint and Restore in Userspace), etc. CRIU is an open-source tool that can implement the checkpoint and restore functions of user-space processes. In this embodiment, it can be used to identify and obtain the current inference service processor process.

[0064] Schedule the process inspection component to perform a process inspection on the computing node to obtain the processor process of the current inference service. A process snapshot of the processor process is created through the container snapshot function.

[0065] Convert the format of the process snapshot into an image format as the running process file.

[0066] Furthermore, the first storage unit and the second storage unit can be connected.

[0067] The running process file can be stored in the first storage unit. The startup scheduling file is stored in the second storage unit, and both the running process file and the startup scheduling file carry the service label of the current inference service. So that when the current inference service is created, its running process file and startup scheduling file can be obtained from the first storage unit and the second storage unit.

[0068] With such a design, the first storage unit and the second storage unit can be used to store the running process files and startup scheduling files of multiple inference services, that is, the number of inference services included in the target inference service is multiple. When multiple inference services belonging to the target inference service are started, the startup efficiency can be improved, and the service startup and the performance of the target inference service can be further improved.

[0069] S208: Create an executable file that adapts to the running process file as the startup scheduling file.

[0070] In this embodiment, the initial startup file for deploying and running the current inference service to the computing cluster can be obtained.

[0071] Control the hardware information and computing resources in the startup scheduling file to match the initial startup file; among them, the computing resources and hardware information are associated with the associated processor, and the associated processor is a processor with a processor process. In this way, by extracting the hardware information and computing resource information involved in the initial startup file and controlling the content in the startup scheduling file that matches the extracted information, it is possible to ensure that the startup scheduling file supports the reliable startup of its affiliated inference service, ensure that the inference service started using the startup scheduling file can have corresponding inference capabilities, so as to be able to improve the startup efficiency of the inference service while taking into account the reliability of the inference service, and further optimize the performance of the inference service and its affiliated computing cluster.

[0072] Optionally, it is possible to locate the inference image for deploying the current inference service inside the initial startup file; among them, the inference image is used to identify the associated processor when deploying and starting the current inference service, and form a processor process with the associated processor.

[0073] In this way, the inference image can be replaced with a running process file, locate and remove the startup instruction of the inference image, and form a startup scheduling file. In this way, it is possible to delete and adjust the code related to the inference image on the basis of the initial startup file to form a startup scheduling file, which can reduce the difficulty of generating the startup scheduling file and also reduce the risk of the startup scheduling file being unreliable, that is, it can improve the generation efficiency and reliability of the startup scheduling file, so as to be conducive to the target inference service to start efficiently and reliably based on the startup scheduling file, further improve the reliability of service startup, and improve the reliability of the inference service and the computing cluster.

[0074] Specific examples of the initial startup file and the startup scheduling file will be given in combination with the file code later, and will not be elaborated here for the time being.

[0075] In this embodiment, it is possible to determine whether the startup scheduling file of the current inference service has been created. When it is determined that the startup scheduling file of the current inference service has not been created, it is possible to wait for the startup scheduling file to be created. When it is determined that the startup scheduling file of the current inference service has been created, step S209 can be executed.

[0076] Furthermore, when it is sensed that the startup scheduling file has been created, it is also possible to identify whether the startup scheduling file has failed to be created, that is, considering that being created does not equal being created successfully, so as to further enhance the robustness of service startup. When it is determined that the startup scheduling file has failed to be created, step S210 can be executed. When it is determined that the startup scheduling file has been created successfully, step S209 can be executed.

[0077] S209: Load the current inference service as the target inference service into the load balancer.

[0078] In this embodiment, in response to the completion of the creation of the scheduling file, the current inference service is used as the target inference service. The target inference service can be loaded into the load balancer of the computing cluster, so that the load balancer schedules the target inference service to execute the inference task.

[0079] S210: Load the current inference service as a regular inference service into the load balancer.

[0080] In this embodiment, in response to the failure of creating the startup scheduling file of the current inference service, it is used as a regular inference service.

[0081] In this way, in this embodiment, it is also possible to take into account the inference services for which it is difficult to obtain the processor process and create the startup scheduling file. Such inference services are used as regular inference services, and when there is a startup requirement for them, the initial startup file can still be used for startup to ensure the reliable startup of the regular inference service, thereby improving the compatibility of service startup.

[0082] S211: Determine that the startup of the current inference service is completed.

[0083] Taking the foregoing running process file as an example of an image format file, generally speaking, after the successful execution of creating the processor process snapshot in this embodiment, the snapshot file can be converted into a snapshot image inferenceImage_snash, and the snapshot image can be uploaded to the image repository, that is, the first storage unit. In order to be able to create the startup scheduling file of the current inference service based on the snapshot image of the current inference service, that is, the startup scheduling script file deployscripts_snash. deployscripts_snash can include the settings of the number of CPUs, the number of GPUs, the scaling metrics, the container image, the GPU model, the CPU model, inferenceImage_snash, the startup command startCommand_snash, etc. required by the current inference service.

[0084] At the same time, according to the hardware requirements for the operation of the GPU process, it is possible to control the newly started first inference service to use the same hardware as that used when forming its startup scheduling file. For example, it is possible to query to obtain the resource information of the computing node running the first inference service, such as the GPU card model, the CPU model, etc. The relevant details of controlling the newly started first inference service to use the same hardware as that used when forming its startup scheduling file will be elaborated in detail later.

[0085] The following gives examples of the specific initial startup file and startup scheduling file in combination with the file code.

[0086] The initial startup file may include the following code: apiVersion: inference.inais / v1alpha1 kind: DistributedInference metadata: name: vllm namespace: ds-dist spec: hpa: maxReplicas: 1 metrics: - resource: name: cpu target: averageUtilization: 2 type: Utilization type: Resource minReplicas: 1 scaleTargetRef: apiVersion: inference.inais / v1alpha1 kind: DistributedInference name: vllm mpiReplicaSpecs: Launcher: replicas: 1 template: spec: nodeSelector: accelerator: "NVIDIA-H20-3e" containers: - command: - / bin / sh - -c - / vllm-workspace / start.sh --tensor-parallel-size 4 --pipeline-parallel-size 4 --model-path / huggingface / --model-id model image: 192.168.16.0.151:5000 / mine / vllm / vllm:v0.7.2 imagePullPolicy: IfNotPresent env: - name: "h_parameter" value: "--max-model-len 65536 --gpu-memory-utilization 0.98 --max-num-seqs 256" name: launcher ports: - containerPort: 8080 protocol: TCP resources: limits: cpu: 10 memory: 100G nvidia.com / gpu: 4 requests: cpu: 10 memory: 100G nvidia.com / gpu: 4 volumeMounts: - mountPath: / huggingface name: ww - mountPath: / dev / shm name: shared-mem volumes: - hostPath: path: / data0 / Models / DeepSeek-R1-671B type: Directory name: ww - emptyDir: medium: Memory sizeLimit: 10.24G name: shared-mem terminationGracePeriodSeconds: 0 Worker: replicas: 3 template: spec: nodeSelector: accelerator: "NVIDIA-H20-3e" containers: - image: 172.16.0.151:5000 / mine / vllm / vllm:v0.7.2-dis imagePullPolicy: IfNotPresent name: worker resources: limits: cpu: 10 memory: 100G nvidia.com / gpu: 4 requests: cpu: 10 memory: 100G nvidia.com / gpu: 4 volumeMounts: - mountPath: / huggingface name: ww - mountPath: / dev / shm name: shared-mem command: - / bin / sh - -c - / vllm-workspace / start_worker.sh volumes: - hostPath: path: / data0 / Models / DeepSeek-R1-671B type: Directory name: ww - emptyDir: medium: Memory sizeLimit: 10.24G name: shared-mem terminationGracePeriodSeconds: 0 replicas: 1 slots: 1 sshAuthMountPath: / root / .ssh The startup scheduling file can contain the following code: apiVersion: inference.inais / v1alpha1 kind: DistributedInference metadata: name: vllm namespace: ds-dist spec: hpa: maxReplicas: 1 metrics: - resource: name: cpu target: averageUtilization: 2 type: Utilization type: Resource minReplicas: 1 scaleTargetRef: apiVersion: inference.inais / v1alpha1 kind: DistributedInference name: vllm mpiReplicaSpecs: Launcher: replicas: 1 template: spec: nodeSelector: accelerator: "NVIDIA-H20-3e" containers: image: 192.168.16.0.151:5000 / mine / vllm / vllm:v0.7.2-snashpot imagePullPolicy: IfNotPresent env: - name: "h_parameter" value: "--max-model-len 65536 --gpu-memory-utilization 0.98 --max-num-seqs 256" name: launcher ports: - containerPort: 8080 protocol: TCP resources: limits: cpu: 10 memory: 100G nvidia.com / gpu: 4 requests: cpu: 10 memory: 100G nvidia.com / gpu: 4 volumeMounts: - mountPath: / huggingface name: ww - mountPath: / dev / shm name: shared-mem volumes: - hostPath: path: / data0 / Models / DeepSeek-R1-671B type: Directory name: ww - emptyDir: medium: Memory sizeLimit: 10.24G name: shared-mem terminationGracePeriodSeconds: 0 Worker: replicas: 3 template: spec: nodeSelector: accelerator: "NVIDIA-H20-3e" containers: - image: 172.16.0.151:5000 / mine / vllm / vllm:v0.7.2-dis imagePullPolicy: IfNotPresent name: worker resources: limits: cpu: 10 memory: 100G nvidia.com / gpu: 4 requests: cpu: 10 memory: 100G nvidia.com / gpu: 4 volumeMounts: - mountPath: / huggingface name: ww - mountPath: / dev / shm name: shared-mem command: - / bin / sh - -c - / vllm-workspace / start_worker.sh volumes: - hostPath: path: / data0 / Models / DeepSeek-R1-671B type: Directory name: ww - emptyDir: medium: Memory sizeLimit: 10.24G name: shared-mem terminationGracePeriodSeconds: 0 replicas: 1 slots: 1 sshAuthMountPath: / root / .ssh It can be seen that the difference between the startup scheduling file and the initial startup file is that the code part of the inference image in the initial startup file code is deleted. The specific code content is as follows: “ - command: - / bin / sh - -c - / vllm-workspace / start.sh --tensor-parallel-size 4 --pipeline-parallel-size 4 --model-path / huggingface / --model-id model” At the same time, the startup scheduling file also supplements the relevant information about the running process file. The specific code is as follows: The code field of “-snashpot” is added to “image: 192.168.16.0.151:5000 / mine / vllm / vllm:v0.7.2-snashpot”. Among them, snashpot is an example of the running process file in the image format. That is to say, the inference image and its startup instructions can no longer be configured in the startup scheduling file.

[0087] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a specific embodiment for expanding the inference service of this application.

[0088] S301: Obtain the current expansion instruction and parse the indicated inference service as the service to be verified.

[0089] In this embodiment, the current expansion instruction represents the instruction obtained currently and used for expanding the inference service accordingly.

[0090] In this embodiment, the current expansion instruction can be parsed to obtain the inference service to be expanded, and the parsed inference service is used as the service to be verified. The content to be verified is to evaluate the startup files (initial startup file and startup scheduling file) relied on for its expansion.

[0091] S302: Identify whether the service to be verified belongs to the target inference service.

[0092] In this embodiment, when it is identified that the service to be verified belongs to the target inference service, step S303 is executed. When it is identified that the service to be verified does not belong to the target inference service, step S313 is executed.

[0093] Optionally, it is possible to identify whether there is a running process file and / or a startup scheduling file of the service to be verified to determine whether it belongs to the target inference service. That is, when there is a running process file and / or a startup scheduling file of the service to be verified, it is considered that the service to be verified belongs to the target inference service; when there is no running process file and / or startup scheduling file of the service to be verified, it is considered that the service to be verified does not belong to the target inference service, that is, the service to be verified belongs to the conventional inference service.

[0094] S303: Take the service to be verified as the first inference service and retrieve its startup scheduling file.

[0095] In this embodiment, in response to the service to be verified belonging to the target inference service, it can be considered that the startup scheduling file can be used to expand its capacity to improve the creation and startup efficiency during its expansion. Therefore, the service to be verified can be taken as the first inference service, and the startup scheduling file of the service to be verified can be obtained.

[0096] S304: Evaluate the scheduling factor of the processor specification associated with the first inference service.

[0097] In this embodiment, considering that there may be insufficient remaining resources of the computing node and / or the associated processor when the first inference service is expanded, the scheduling factor indicating whether the remaining resources are sufficient can be evaluated.

[0098] S305: Verify whether the associated processor has sufficient resources to expand the first inference service.

[0099] In this embodiment, when the associated processor has sufficient resources to expand the first inference service, step S306 is executed. When the associated processor does not have sufficient resources to expand the first inference service, step S308 is executed.

[0100] The scheduling factor can be used to verify whether the associated processor has sufficient resources to expand the first reasoning service. Optionally, in this embodiment, it can be determined whether the associated processor of the first reasoning service has sufficient resources; or, it can be determined whether the associated processor of the first reasoning service has sufficient resources to characterize whether the associated processor has sufficient resources; or, it can be determined whether the associated processor of the first reasoning service and the computing nodes deployed thereon have sufficient resources. When both have sufficient resources, it is considered that the associated processor has sufficient resources, otherwise it is considered that the associated processor has insufficient resources, thereby taking into account the associated processor and its computing nodes, and improving the reliability of the expansion of the first reasoning service, so as to further improve the performance of the reasoning service and the computing cluster.

[0101] S306: The resource scheduling associated processor executes its startup scheduling file and writes the running process file to create and run a new first reasoning service as the second reasoning service.

[0102] In this embodiment, in response to the associated processor having sufficient resources to expand the first reasoning service, resources can be scheduled for the associated processor. The startup scheduling file of the first reasoning service is executed and the running process file is written to create and run a new first reasoning service as a second reasoning service. That is, the second reasoning service is a new first reasoning service formed by the expansion.

[0103] S307: Load the second inference service to the load balancer.

[0104] In this embodiment, in response to the second reasoning service being created, the second reasoning service can be loaded to the load balancer, so that the load balancer can be informed that the second reasoning service is in a ready state and can schedule reasoning tasks to it.

[0105] S308: Delete the first reasoning service from the associated processor.

[0106] In this embodiment, in response to the associated processor not having sufficient resources to expand the first reasoning service, that is, it can be considered that the associated processor has insufficient resources, the first reasoning service can be deleted from the associated processor so that the first reasoning service and its second reasoning service can be reliably deployed.

[0107] S309: Inferring the inference image in the initial startup file to select another processor with sufficient resources as the update processor.

[0108] In this embodiment, in response to the associated processor not having sufficient resources to expand the first inference service, the processor for deploying the first inference service and the expanded second inference service can be selected. Therefore, the initial startup file of the first inference service can be obtained to extract the inference image in the initial startup file and perform re-inference at the current moment, so as to select another processor with sufficient resources as the update processor.

[0109] S310: Deploy the first inference service to the updated processor using the initial startup file, and use the updated processor as the new associated processor.

[0110] In this embodiment, in response to completing the inference and selection of the updated processor, the first inference service can be deployed to the updated processor using the initial startup file, and the updated processor can be used as the new associated processor.

[0111] S311: Deploy the second inference service to the new associated processor.

[0112] In this embodiment, the second inference service can be similarly deployed to the newly inferred associated processor. Optionally, the second inference service can be deployed synchronously or asynchronously when the first inference service is deployed; or, when the first inference service is deployed and running, create its new run process file and startup scheduling file, and deploy the second inference service based on the newly created run process file and startup scheduling file. There is no strict limitation here.

[0113] S312: Load the first inference service and the second inference service into the load balancer.

[0114] In this embodiment, in response to the first inference service and the second inference service completing the deployment on the new associated processor, the first inference service and the second inference service can be loaded into the load balancer. So that the load balancer knows that the first inference service and the second inference service are both in a ready state and can schedule inference tasks to both of them.

[0115] S313: Use the service to be verified as the third inference service, and create a new third inference service as the fourth inference service using its initial startup file.

[0116] In this embodiment, in response to the service to be verified not belonging to the target inference service, it can be considered that the service to be verified belongs to the regular inference service. The service to be verified can be used as the third inference service, and a new third inference service can be created as the fourth inference service using its initial startup file.

[0117] S314: Load the fourth inference service into the load balancer.

[0118] In this embodiment, in response to completing the creation of the fourth inference service, the fourth inference service can be loaded into the load balancer. So that the load balancer knows that the fourth inference service is in a ready state and can schedule inference tasks to the fourth inference service.

[0119] S315: Determine that the inference service expansion and startup indicated by the current expansion instruction are completed.

[0120] That is to say, in this embodiment, in response to the startup request being an expansion request, the startup scheduling file of the first inference service can be obtained as the target startup file, and the running process file of the first inference service can be obtained as the target process file. The target startup file is executed and written into the target process file to create a new first inference service as the second inference service. The resource usage information when the first inference service is scheduled by the computing cluster is obtained. The computing nodes corresponding to the physical resources in the resource usage information are identified as associated nodes. The physical resources of the associated nodes are used for inference tasks when the second inference service is scheduled. With such a design, when creating a new first inference service, that is, its second inference service, it is possible to achieve a strong match between the resource usage of the second inference service, such as node resources and processor resources, and the resource usage of the first inference service, which is beneficial to ensuring the reliable operation of the second inference service and reliably executing inference tasks to ensure the performance of the first inference service and the second inference service.

[0121] In response to obtaining the startup request of the third inference service, the third inference service is restarted using the initial startup file of the third inference service. Among them, the third inference service belongs to a conventional inference service. As described above, the conventional inference service is formed by the current inference service for which the startup scheduling file creation fails.

[0122] As Figure 5 exemplified Figure 5 in the figure is a comparison schematic diagram of an embodiment of the startup of the conventional inference service and the startup of the target inference service in this application. Combining the implementation manners described above, it can be seen that when starting the target inference service, the process of immediately performing inference on the deployed processor during the startup of the conventional inference service can be omitted, which is beneficial to improving the startup efficiency of the inference service.

[0123] Generally speaking, taking the running process file as a snapshot image as an example, when the service management device receives an inference service expansion request or an administrator's instruction to create an inference service with the same configuration, etc., it can be considered that the startup request of the inference service is obtained. The service management device can determine whether there is a snapshot image of the inference service.

[0124] When there is a snapshot image of the inference service, the inference service can be used as the first inference service. The first inference service can be started based on the new inference service deployment script deployscripts_snash (i.e., the startup scheduling file) and the snapshot image inferenceImage_snash (i.e., the running process file) to adapt to the startup request.

[0125] Taking the start-up process as an example of the expansion process, the GPU model and CPU model required by the first inference service can be used as the scheduling factors for this expanded inference service. The scheduler is used to select a computing node with a strictly matching GPU and / or CPU specification to start the first inference service. When there are no resources in the computing cluster that meet the number of GPU cards, GPU card models, and CPU models, the inference service fallback is triggered, the first inference service is deleted, and the default inference service deployment script deployscripts_defalut (i.e., the initial start-up file) is used to create the first inference service again.

[0126] When there is no snapshot image inferenceImage_snash of the inference service, the default inference service deployment script deployscripts_defalut is directly used to create the inference service. The scheduler selects a node that meets the number of GPU cards and CPU number according to the GPU resource usage of the cluster to start the inference service. Further, the inference service can also be used as the current inference service to create a snapshot image for it.

[0127] After the inference service requested by the start-up request runs successfully, the access information of the inference service replica can be loaded into the complex balancer for scheduling by the load balancer.

[0128] The embodiments of the present application provide a service control method. Combining with the execution process of the service control method, the working principle of the service control method is described in detail.

[0129] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of an embodiment of the service control method of the present application: S401: In response to the initial operation of the current inference service, determine whether the current inference service runs successfully.

[0130] S402: In response to the successful operation of the current inference service, execute the service start method.

[0131] S403: Schedule the current inference service to execute an inference task.

[0132] Generally speaking, in large-scale AI (Artificial Intelligence) inference services, the training framework usually needs to load the model file stored in the shared storage and read it into the GPU video memory. As a result, when the model file is large, it usually takes a long time to complete the loading of the model file. For example, it may take about 30 minutes or even longer to load the model, which will cause the inference service to start very slowly and easily lead to waste of GPU resources.

[0133] Especially in the online environment of the inference service, when there is a large access pressure on the inference service, the automatic scaling of the inference service can usually be triggered. The newly added inference services formed by scaling also take a long time to run normally, and during this period, the service access pressure cannot be quickly relieved.

[0134] In response to this, the service startup method provided in this application proposes a unified method for generating and restoring processor process snapshots based on running process files such as process images, which can be used in processing such as CPUs, GPUs, and acceleration cards.

[0135] Taking CPUs and GPUs as examples, the CPU and GPU process snapshots can be encapsulated as container images, so as to support node migration and fast startup across computing nodes. Specifically, after the inference service runs successfully for the first time, operations for taking CPU and GPU process snapshots can be added to save the states of the GPU process and CPU process in the initialization running state as container images.

[0136] In this way, when the inference service needs to perform replica scaling or the platform creates inference services of the same specification, a computing node matching the snapshot image can be selected based on a resource scheduling mechanism that meets the GPU and CPU specifications, and the snapshot container image can be used to quickly start the inference service, thereby reducing the service startup time and ensuring that the inference service can run when the number of GPU resources is sufficient. And when the cluster resources do not meet the GPU and CPU specifications, the default inference service deployment script can be used to deploy the inference service to ensure the reliable deployment and startup of the inference service and optimize the performance of the inference service and the computing cluster.

[0137] In other words, this application is an optimization solution for the startup mechanism of large model inference services. When managing inference services on the AI platform, a mechanism for generating and using GPU process snapshots can be added to optimize and remove the process of loading model files into the GPU video memory, thereby reducing the inference service startup time. When the inference service is first deployed and runs successfully, the platform first takes snapshots of the CPU process and GPU process of the inference service. After the snapshot is completed, the inference service is exposed to participate in the execution of inference tasks, and at the same time, an asynchronous process is started to convert the snapshot file into a container image and upload it to the image repository, and a deployment script for the inference service based on the snapshot image can be defined. When the administrator creates an inference service of the same specification or the inference service is scaled, it is first judged whether there is a snapshot image, and a computing node with the same GPU specification and the same CPU specification is selected based on the snapshot image during resource scheduling to run the inference service. When the inference service starts, it does not need to read the model file, thus accelerating the startup of the inference service.

[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, and no strict limitation is made here.

[0139] An embodiment of the present application also provides a service management device.

[0140] In this embodiment, the service management device may include a connection module and a control module.

[0141] The connection module can be connected to the computing cluster.

[0142] The control module can be connected to the connection module to implement the service startup method or service control method as described in any of the previous embodiments.

[0143] Specifically, the service startup method may at least include: in response to the current inference service running, obtaining the processor process when the current inference service runs on the computing cluster, and recording the processor process to form a running process file; creating an executable file adapted to the running process file as a startup scheduling file; wherein the startup scheduling file is used to start the current inference service; taking the current inference service that has completed the creation of the startup scheduling file as the target inference service; in response to obtaining a startup request for the first inference service, executing its startup scheduling file and writing it into the running process file to start the first inference service; wherein the first inference service belongs to the target inference service.

[0144] The service control method may at least include: in response to the first run of the current inference service, determining whether the current inference service runs successfully; in response to the current inference service running successfully, executing the service startup method as described in any of the above embodiments; scheduling the current inference service to execute an inference task.

[0145] Furthermore, for the description of the features in the corresponding embodiment of the service management device, reference can be made to the relevant descriptions in the corresponding embodiments of the service management device method and the service control method, which will not be elaborated here one by one.

[0146] An embodiment of the present application also provides an electronic device.

[0147] The electronic device may include a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above service startup method or service control method embodiments. That is, the memory is used to store the computer program. The processor is used to implement the steps of the service startup method as described above when executing the computer program; or, implement the steps of the service control method as described above.

[0148] Embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any of the above service startup method or service control method embodiments when running. That is, when the computer program is executed by a processor, it implements the steps of the service startup method as described above; or, implements the steps of the service control method as described above.

[0149] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0150] Embodiments of the present application also provide a computer program product. The computer program product may include a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above service startup method or service control method embodiments. That is, when the computer program is executed by a processor, it implements the steps of the service startup method as described above; or, implements the steps of the service control method as described above.

[0151] Embodiments of the present application also provide another computer program product. The computer program product may include a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above service startup method or service control method embodiments.

[0152] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0153] The above has introduced in detail a service startup method, a service control method, an electronic device, a computer-readable storage medium, and a computer program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A service startup method, characterized in that, The described service startup method includes: In response to the current inference service running, obtain the processor process when the current inference service runs in the computing cluster, and record the processor process to form a running process file; Create an executable file adapted to the running process file as a startup scheduling file; wherein, the startup scheduling file is used to start the current inference service; Take the current inference service that has completed the creation of the startup scheduling file as the target inference service; In response to obtaining a startup request for the first inference service, execute its startup scheduling file and write it into the running process file to start the first inference service; wherein, the first inference service belongs to the target inference service.

2. The service startup method according to claim 1, wherein The creating an executable file adapted to the running process file as a startup scheduling file includes: Obtain the initial startup file for deploying and running the current inference service to the computing cluster; Control the hardware information and computing resources in the startup scheduling file that match the initial startup file; wherein, the computing resources and the hardware information are associated with the associated processor, and the associated processor is the processor with the processor process.

3. The service startup method according to claim 1, wherein The taking the current inference service that has completed the creation of the startup scheduling file as the target inference service includes: Judge whether the startup scheduling file of the current inference service is created; In response to the completion of the creation of the scheduling file, take the current inference service as the target inference service; Load the target inference service to the load balancer of the computing cluster to schedule the target inference service to execute inference tasks through the load balancer.

4. The service startup method according to claim 3, wherein The responding to the current inference service running includes: Obtain the initial startup file of the current inference service; Run the initial startup file to load the inference model relied on by the current inference service; Judge whether the current inference service is in a ready state based on the loading state of the inference model; In response to the current inference service being in the ready state, determine that the current inference service is running, and do not load the current inference service to the load balancer for the time being.

5. The service startup method according to claim 4, wherein The obtaining the initial startup file of the current inference service includes: Obtain the model file and the inference framework image of the current inference service as initial information; Parse the processor type, the number of each type of processor, the scaling index, the container image, and the startup instruction indicated by the initial information as the resource requirement information of the current inference service; wherein, the processor type includes at least one of a central processing unit, a graphics processing unit, and an acceleration processing unit; Generate the initial startup file by combining the initial information and the resource requirement information.

6. The service startup method according to claim 1, characterized in that Before the responding to the current inference service running includes: Obtain the initial startup file of the current inference service; Read the processor type and the number of each type of processor indicated by the initial startup file as the target processor and its target number; Select the target number of the target processors in the computing cluster as the associated processors; Deploy the current inference service to a computing node carrying the associated processor; and obtain a processor process of the current inference service on the associated processor to form the running process file.

7. The service startup method according to claim 1, wherein After creating the executable file adapted to the running process file as the startup scheduling file, the following steps are further included: Connect the first storage unit and the second storage unit; Store the running process file in the first storage unit, store the startup scheduling file in the second storage unit, and make both the running process file and the startup scheduling file carry the service label of the current inference service; so that when creating the current inference service, obtain its running process file and startup scheduling file from the first storage unit and the second storage unit.

8. The service startup method according to claim 1, wherein When executing the startup scheduling file and writing into the running process file to start the first inference service, the following steps are further included: In response to the startup request being an expansion request, obtain the startup scheduling file of the first inference service as the target startup file, and obtain the running process file of the first inference service as the target process file; Execute the target startup file and write into the target process file to create a new first inference service as the second inference service; Obtain the resource usage information when scheduling the first inference service in the computing cluster; wherein, the computing cluster includes multiple computing nodes; Identify the computing nodes corresponding to the physical resources in the resource usage information as associated nodes; Use the physical resources of the associated nodes to perform inference tasks when scheduling the second inference service.

9. The service startup method according to claim 1 or 8, characterized in that When obtaining the processor process of the current inference service running in the computing cluster and recording the processor process to form the running process file, the following steps are included: Pre-obtain the computing nodes included in the computing cluster, deploy a process inspection component to the computing nodes, and control the container management component and node communication component of the computing nodes to enable the container snapshot function; Schedule the process inspection component to perform a process inspection process on the computing nodes to obtain the processor process of the current inference service; Make a process snapshot of the processor process through the container snapshot function; Convert the format of the process snapshot into an image format as the running process file.

10. The service startup method according to claim 1, wherein After creating the executable file adapted to the running process file as the startup scheduling file, the following steps are further included: In response to the failure of creating the startup scheduling file of the current inference service, regard it as a regular inference service; In response to obtaining the startup request of the third inference service, restart the third inference service using the initial startup file of the third inference service; wherein, the third inference service belongs to the regular inference service.

11. The service startup method according to claim 1 or 10, characterized in that, The service startup method further includes: Obtain the current expansion instruction; Parse the inference service indicated to be expanded by the current expansion instruction as the service to be verified; Identify whether the service to be verified belongs to the target inference service; In response to the service to be verified belonging to the target inference service, determine it as the first inference service; otherwise, determine it as the third inference service belonging to the regular inference service.

12. A service control method, characterized in that, The service control method includes: In response to the initial operation of the current inference service, determine whether the current inference service operates successfully; In response to the successful operation of the current inference service, execute the service startup method according to any one of claims 1 to 11; Schedule the current inference service to execute an inference task.

13. An electronic device, characterized in that, The electronic device includes: a memory for storing a computer program; a processor for implementing the steps of the service startup method according to any one of claims 1 to 11 when executing the computer program; or, implementing the steps of the service control method according to claim 12.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the service startup method according to any one of claims 1 to 11 when executed by a processor; or, implements the steps of the service control method according to claim 12.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the service startup method according to any one of claims 1 to 11 when executed by a processor; or, implements the steps of the service control method according to claim 12.

Citation Information

Patent Citations

  • Method, apparatus, device, and storage medium for scheduling jobs in cluster

    CN109117265A

  • Method and device for publishing reasoning service in GPU cluster, equipment and medium

    CN111324457A

  • Container starting method and device

    CN114064190A

  • Inference service deployment method and device, equipment and storage medium

    CN116301912A

  • Inference service management method, equipment, medium and computer program product

    CN119415273A

Cited By

  • Inference service copy pool management method, electronic equipment and storage medium

    CN122287909A

  • Inference service replica pool management method, electronic device, and storage medium

    CN122287909B