Service updating method, electronic equipment, storage medium and program product
By receiving client requests, determining storage path and version information, performing update operations and restarting the inference framework container, the problem of untimely update of inference framework services in cloud computing is solved, and timely updates are achieved without additional acceleration cards.
Patent Information
- Application Number
- CN202510840286.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In cloud computing, the update of the inference framework service is not timely due to the use of accelerator card resource, and the update of the container group cannot be completed in time.
By receiving the client's service update request, determine the storage path of the original inference framework code, read the original version information, and perform update operations based on the updated version information, restart the inference framework container to complete the update.
Without an additional acceleration card, the updated reasoning framework code can be applied in time, which improves the speed and efficiency of service updates.
Smart Images

Figure CN120353487A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to a service update method, an electronic device, a storage medium, and a program product. Background Art
[0002] In the field of cloud computing technology, in order to improve resource utilization, various services can generally be deployed using container groups. For example, question-and-answer services, inference services, etc. Generally, multiple acceleration cards can be provided on each computing node in a computer cluster, and at least one container group can be deployed on each acceleration card.
[0003] In an inference service, as user requirements change continuously, the inference framework code will change. Furthermore, the container group needs to be updated. When performing the update, in order to avoid affecting the inference service, a rolling update mechanism is generally adopted, that is, first obtain the mirror file of the updated inference service, and create and start a new container group based on the mirror file of the updated inference service. When it is detected that the new container group has been successfully started, first switch the service flow of the original container group to the new container group, and then delete the old container group. However, in order to make full use of the acceleration card resources, when users deploy services on computing nodes, they usually occupy all the acceleration cards. In this way, when updating the container group, there are no additional acceleration cards available for deploying the new container group, resulting in the inference framework service being unable to be updated in a timely manner. Summary of the Invention
[0004] This application provides a service update method, device, electronic device, storage medium, and program product to solve the problem of untimely update of the inference framework service.
[0005] This application provides a service update method, including: Receiving a service update request sent by a client, where the service update request includes service indication information, location indication information, and update version information; Determining the original version information of the original inference framework code according to the service indication information and the location indication information; When it is determined to perform an update operation according to the update version information and the original version information, performing an update operation on the original inference framework code according to a preset update method and the update version information; When it is detected that the update operation is completed, restarting the inference framework container based on the updated inference framework code to complete the update operation of the inference framework service.
[0006] This application also provides a service update device, including: A receiving module, configured to receive a service update request sent by a client, where the service update request includes service indication information, location indication information, and update version information; A determination module, configured to determine the original version information of the original inference framework code according to service indication information and location indication information; An update module, configured to, when it is determined to perform an update operation according to the update version information and the original version information, perform an update operation on the original inference framework code according to a preset update method and the update version information; A restart module, configured to, when it is detected that the update operation is completed, restart the inference framework container based on the updated inference framework code to complete the update operation of the inference framework service.
[0007] This application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above service update methods when executing the computer program.
[0008] This application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program, when executed by a processor, implements the steps of any of the above service update methods.
[0009] This application further provides a computer program product, including a computer program, and the computer program, when executed by a processor, implements the steps of any of the above service update methods.
[0010] Through this application, when receiving a service update request sent by a client, the first storage path of the original inference framework code can be determined according to the service indication information and location indication information included in the service update request. Furthermore, based on the first storage path, the original version information of the original inference framework code is read. When it is determined to perform an update operation according to the update version information and the original version information, the original inference framework code stored under the first storage path is updated according to a preset pulling method. When it is detected that the update operation is completed, the updated inference framework code is read from the storage path. Based on the updated inference framework code, the inference framework container is restarted. In this way, only the original inference framework code needs to be updated, and only the inference framework container needs to be restarted after the update. Even without an additional acceleration card, the updated inference framework code can be applied in a timely manner, that is, the update of the inference framework service can be completed in a timely manner. In addition, in the related art, a container group needs to be restarted. Since the inference engine is included in the container group and the startup time of the inference engine is relatively long, the speed of restarting the container group is slow. However, this solution only needs to restart the inference framework container, which can improve the update speed of the overall service and apply the new inference framework code more timely. Description of the Drawings
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0012] Figure 1 It is a schematic diagram of the architecture of a container group orchestration platform provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a service update method provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of a service deployment method provided by an embodiment of the present application; Figure 4 It is a schematic flowchart of a data record storage method provided by an embodiment of the present application; Figure 5 It is a schematic diagram of the architecture of a container group deployment provided by an embodiment of the present application; Figure 6 It is a schematic flowchart of another service deployment method provided by an embodiment of the present application; Figure 7 It is a schematic flowchart of another service update method provided by an embodiment of the present application; Figure 8 It is a schematic flowchart of a service update device provided by an embodiment of the present application; Figure 9 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0015] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific implementation manners.
[0016] The following will give a detailed explanation of the professional terms involved in this application.
[0017] Container: It is a lightweight, portable, and self - contained software packaging and runtime environment. It packages an application and all its dependencies into a standardized unit.
[0018] Pod (Container Group): It is the smallest scheduling and management basic unit in the container group orchestration platform. It is a collection of one or more closely related containers, and these containers can share storage and network resources.
[0019] Inference service: It refers to the task process of predicting (Inference) a deep - learning model in an actual application scenario, which is used to provide actual services, such as face recognition, speech translation, recommendation systems, etc. For example, the deep - learning model can be a Large Language Model (LLM), which will be abbreviated as the model hereinafter.
[0020] Inference framework: It refers to the software that provides full - process support for model inference. It is used to provide an application programming interface (API) for the inference engine, interact with the inference engine using pre - configured chat templates (ChatTemplate) for the application (such as chat templates in jinja2 syntax), manage the life cycle of the inference engine, and monitor the inference engine, such as model loading, unloading, version control, resource allocation, health check, etc.
[0021] Inference engine: It refers to the runtime core responsible for hardware - level acceleration and efficient execution of calculations. It is used to deploy the model file to the acceleration card and use the acceleration card for model inference to complete the inference service.
[0022] Persistent Volume: It refers to a storage resource in a computer cluster, which can be pre - configured by an administrator or dynamically configured using a storage class. The storage class can be a Network File System (NFS) class.
[0023] Persistent Volume Claim: It refers to a claim for storage resource requirements (such as resource size, access mode, etc.). The container group orchestration platform generally binds or creates the corresponding PV based on the PVC.
[0024] The service update method provided by this application can be applied to a container group orchestration platform. As Figure 1 shown, the container group orchestration platform may include at least one computing node. The computing node can be a server, for example, an NFS server. When there is only one computing node in the computer cluster, this computing node can serve as both a management node and a storage node at the same time. Or, when the computer cluster includes multiple computing nodes, one of the multiple computing nodes can be set as the management node, and each computing node can serve as a storage node. The above-mentioned computing node can only execute computing services (for example, deploying container groups). Correspondingly, a storage node can also be set in the computer cluster. Each computing node may include a Central Processing Unit (CPU), memory, multiple acceleration cards, etc. For example, the acceleration card can be a Graphics Processing Unit (GPU).
[0025] Among them, the container group can be a business container group, a code management container group, etc.
[0026] The business container group can be used to provide actual services for external clients or web page calls. The business container group can be an inference business container group. The inference business container group can be used to provide actual inference services for external clients or web page calls. For example, the inference business container group can be an LLM inference business container group. The inference business container group may include an initialization full-pull container, an inference engine container, an inference framework container, etc. Among them, the inference engine container and the inference framework container are both container instances. The initialization full-pull container is only executed once when the inference business container group is started, and is used to fully pull the code and start the inference engine container and the inference framework container. The inference engine container is used to deploy the model file to the acceleration card and use the acceleration card for model inference to complete the inference service. The inference framework container is used to provide an application programming interface for the inference engine container to interact with the inference engine container. In addition, the inference framework container can also manage the life cycle of the inference engine container and monitor various performance indicators of the inference engine container. The business container group can be mounted on at least one PV to use the PV to store relevant code files, for example, to store the inference framework code. The inference framework code can be code written in the Python language.
[0027] The code management container group can be used to provide management of the code files stored in the PVs mounted by different business container groups, as well as independent update and centralized management of the code files stored in the PVs. The code management container group may include a code management container.
[0028] During the process of the inference service container group performing the inference service, if the inference framework code is updated, it is necessary to update the inference framework code in a timely manner and execute the inference service based on the updated inference framework code.
[0029] The following takes the container group orchestration platform including one computing node as an example for detailed description.
[0030] An embodiment of the present application provides a service update method, which can be executed by the above-mentioned computing node to complete the update and application of the inference framework code in a timely manner. As Figure 2 shown, the specific processing steps of the service update method may include: Step S201, receive a service update request sent by the client.
[0031] Among them, the service update request may include service indication information, location indication information, and update version information. The service indication information may include one or more of scenario identification information, service identification information, etc. The scenario identification information can be used to indicate the specific environment on which the service depends, and it can be a scenario name, for example, "inference", "llm", "chat". The service identification information can be used to distinguish different services, such as service names, for example, "qwen3-instruce", "deepseek-r1", "inf-svc", "ocr", etc. The location indication information can be used to indicate the storage location of the relevant code involved in the service, and it can be a PVC name, for example, "qwen3infcode", "dsr1", etc. The update version information can be a version number.
[0032] Specifically, the user can perform an update operation on the client (for example, enter the update version information in the service interface of any business and click the update button). After the client detects the user's update operation, it generates a service update request (carrying the update version information) and sends it to the corresponding computing node. When the computing node receives the service update request sent by the client, it can parse the service update request to obtain the service indication information, location indication information, and update version information, etc.
[0033] Step S202, determine the original version information of the original inference framework code according to the service indication information and the location indication information.
[0034] Specifically, since the version specified by the user may be the same as the version of the original inference framework code, if the service update operation is directly triggered when a service update request is received, it will cause resource waste (restarting the container wastes resources). Therefore, to avoid this problem, the computing node can first determine the original version information of the original inference framework code according to the service indication information and the location indication information for subsequent version comparison operations. Accordingly, the computing node can find the original version information according to the following specific steps: Step 1, determine the first storage path of the original inference framework code according to the service indication information and the location indication information.
[0035] Step 1, obtain the second storage path mounted by the pre-built code management container.
[0036] Among them, the lower-level path of the second storage path may include at least one sub-path. The second storage path may be the storage path corresponding to the PV mounted by the code management container group to which the code management container belongs. The sub-path may be the storage path corresponding to the PV mounted by the business container group. For example, the storage path corresponding to the PV mounted by the LLM inference business container group. The format of the sub-path may be "scenario identification information - location indication information - unique identifier", where the unique identifier may be a Universally Unique Identifier (UUID). Since the PV can be created through a storage class, when creating, the host path needs to be specified, that is, the path corresponding to the PV mounted by each business container group is a sub-path under the storage class path under the host path. Therefore, the code management container group only needs to mount the path corresponding to the storage class to manage the code files stored in all PVs. For example, the host path may be " / mnt / inaisfs / ", and the sub-path may be " / mnt / inaisfs / llm-codeV1-21b8e9ec-187c-11f0-b184-0bfa87fb8069", where "llm" may be the scenario identification information, "codeV1" may be the name of the PV storing the inference framework code, and "codeV1-21b8e9ec-187c-11f0-b184-0bfa87fb8069" may be the UUID. The code management container group can be mounted to " / mnt / inaisfs / " (that is, the second storage path) to manage the code of all business container groups.
[0037] Step 2, determine the sub-path that matches both the service indication information and the location indication information in at least one sub-path as the first storage path.
[0038] Specifically, since the second storage path mounted by the code management container group to which the code management container belongs is located above the storage path (i.e., the above-mentioned sub-path) mounted by the business container group, the computing node can first obtain the second storage path mounted by the code management container group, and then determine a sub-path that matches both the service indication information and the location indication information as the first storage path in the lower-level path of the second storage path. In the LLM scenario, the inference frameworks used by different services can be the same. When matching the first storage path, the matching operation can be performed only by combining the scenario identification information in the service indication information with the location indication information. Or, when the inference frameworks used by different services in the LLM scenario are different, the matching operation can be performed by combining the scenario identification information, service identification information, and location indication information in the service indication information.
[0039] For example, in the case where the same inference framework is applicable to all services in the LLM scenario, the structure of the storage path can be: “ / data / code-root ├──chat-qwen7b-chat-123abc ├──llm-qwen3infcode-456def └──...” Among them, “ / data / code-root” can be the above-mentioned second storage path, and both “chat-qwen7b-chat -123abc” and “llm-qwen3infcode-456def” are sub-paths. When the service indication information is “llm” and the location indication information is “qwen3infcode”, the determined first storage path is “llm-qwen3infcode-456def”. All LLM inference business container groups can be all mounted to the same storage path. For example, all are mounted to “llm-qwen3infcode-456def”.
[0040] By specifically setting up a code management container group, mounting the code management container group to a specified path, and mounting each business container group to the lower-level path of the specified path where the code management container group is mounted, centralized management of code files can be achieved. And during the process of updating the code, the corresponding first storage path can be determined through simple matching of the service indication information and the location indication information, and the location where the inference framework code to be updated is located can be accurately determined.
[0041] Under the above storage path structure, the user can access any sub-path under the second storage path mounted by the code management container through the client, and perform operations such as viewing, modifying, and deleting the code files under any sub-path, so as to achieve centralized management of the code.
[0042] Step 2: Read the original version information based on the first storage path.
[0043] Step 1: Read the original inference framework code from the storage location corresponding to the first storage path based on the first storage path.
[0044] Step 2: Determine the target field corresponding to the preset field identification information in the original inference framework code according to the preset field identification information.
[0045] Step 3: Extract the original version information from the target field.
[0046] Specifically, the computing node can read the original inference framework code from the storage location corresponding to the first storage path based on the first storage path. Furthermore, according to the preset field identification information, determine the target field corresponding to the preset field identification information in the original inference framework code, and extract the original version information from the target field.
[0047] Alternatively, the computing node can directly read the original version information from the storage location corresponding to the first storage path based on the first storage path. In this way, through the first storage path, the original version information of the inference framework code currently used can be determined in a timely manner, which is convenient for accurately determining whether to perform an update operation subsequently and avoiding the problem of resource waste.
[0048] In some alternative embodiments, the computing node can also determine the first storage path in the following manner: Determine the target unique identifier that matches both the service indication information and the location indication information in the pre-constructed correspondence table according to the service indication information and the location indication information; furthermore, generate the first storage path according to the second storage path, the service indication information, the location indication information, and the second storage path.
[0049] In this way, the first storage path can be directly generated, and there is no need to perform a storage path search operation to read the original inference framework code, which is more convenient.
[0050] Step S203: When it is determined to perform an update operation according to the updated version information and the original version information, perform an update operation on the original inference framework code according to the preset update method and the updated version information.
[0051] Among them, the preset update method can be an incremental update method or a full-scale update method.
[0052] Specifically, the computing node can determine whether the updated version information is consistent with the original version information. If so, it determines to execute the update operation. If not, it determines not to execute the update operation, which can save resources. When determining to execute the update operation, the computing node can perform an update operation on the original inference framework code according to the preset update method and the updated version information. Accordingly, the computing node can perform the following specific steps: Step 1: Obtain the update code of the inference framework from the preset storage location according to the preset update method and the updated version information.
[0053] Specifically, to improve the flexibility of the update, this solution provides multiple methods. The computing node can perform corresponding update operations according to the preset update method, that is, first obtain the storage location information corresponding to the preset update method according to the preset update method. Then, according to the preset update method and the updated version information, obtain the update code of the inference framework from the storage location corresponding to the storage location information. Specifically, it can include the following two methods: Method 1: When the preset update method is the incremental update method, the storage location corresponding to the storage location information is the target code repository provided by the target code platform, and the update code is the incremental update code. Accordingly, the computing node can use the pre-built code pulling tool and the updated version information to incrementally pull the incremental update code from the target code repository.
[0054] The first version information of the inference framework code that the computing node pulled last time is recorded in the code pulling tool. The computing node can input the updated version information into the code pulling tool, and the code pulling tool can automatically pull the incremental update code from the target code repository according to the first version information and the updated version information, that is, pull the difference code between the two versions of the first version information and the updated version information. For example, the code pulling tool can be the Go-git tool, and the target code repository can be the GitLab code repository.
[0055] Since Method 1 is applicable to services with frequent updates and the update cycle of the inference framework code is short, this method can be used to update the inference framework code more efficiently.
[0056] Method 2: When the preset update method is the full update method, the storage location corresponding to the storage location information is the target storage node, and the update code is the full update code corresponding to the updated version information. Accordingly, the computing node can obtain the full update code according to the following steps: Step 1: Obtain the second storage path corresponding to the target storage node.
[0057] Step 2: Generate a code update instruction according to the preset instruction template, the second storage path, and the updated version information.
[0058] Among them, the preset instruction template can be an "scp instruction", and the "scp instruction" is a Linux instruction.
[0059] Step 3: Send the code update instruction to the target storage node to notify the target storage node to obtain the full update code corresponding to the update version information.
[0060] Step 4: Receive the full update code corresponding to the update version information sent by the target storage node.
[0061] Specifically, the target storage node can be a node in the container group orchestration platform, which is specifically used to store inference framework codes of various versions. Therefore, when performing an update, the computing node can first obtain the second storage path corresponding to the target storage node. Add the second storage path to the first specified position of the preset instruction template, and add the update version information to the second specified position of the preset instruction target to generate a code update instruction. The computing node can send the generated code update instruction to the target storage node. In this way, when the target storage node receives the code update instruction, it can parse out the update version information from it, and based on the update version information, determine the target inference framework code corresponding to the update version information among the inference framework codes of various versions stored in the target storage node as the full update code. The target storage node can send the full update code to the computing node, and the computing node can receive the full update code, that is, obtain the update code of the inference framework.
[0062] In any way, through the update version information, version control can be achieved, and rollback or update of any version can be realized, rather than only being able to passively update to the latest version.
[0063] Step 2: Based on the update code, perform an update operation on the original inference framework code under the first storage path.
[0064] Method 1: When the update code is an incremental update code, the computing node can directly add the incremental update code to the original inference framework code under the first storage path to complete the update operation of the original inference framework code.
[0065] Method 2: When the update code is a full update code, the computing node can replace the original inference framework code stored under the first storage path with the full update code.
[0066] Step S204: When it is detected that the update operation is completed, restart the inference framework container based on the updated inference framework code to complete the update operation of the inference framework service.
[0067] Specifically, when the computing node detects that the update operation is completed, it can first close the original inference framework container, read the updated inference framework code from the first storage path, and create and start a new inference framework container based on the updated inference framework code to complete the update operation of the inference framework service.
[0068] In some alternative embodiments, after step S204, the computing node can manage the inference engine container by using the restarted inference framework container. Accordingly, the computing node can also perform the following steps: Step 1, obtain the port information corresponding to each of the pre-built at least one inference engine container.
[0069] Step 2, re-take over the inference engine container corresponding to the port information based on the port information.
[0070] Specifically, the inference engine container periodically sends a registration request to the inference framework container in the inference service container group to which it belongs. After restarting the inference framework container, the port information corresponding to each of the at least one inference engine container can be obtained in real time through the registration request, and the inference engine container can be continuously monitored based on the port information. In this way, the inference engine container can be monitored in a timely manner, and the fault problems of the inference engine container can be processed in a timely manner.
[0071] In the service update method according to the embodiment of the present application, when the computing node receives a service update request sent by a client, it can first determine the first storage path of the original inference framework code according to the service indication information and the location indication information included in the service update request. Further, based on the first storage path, the original version information of the original inference framework code is read. When it is determined to perform an update operation according to the update version information and the original version information, the original inference framework code stored in the first storage path is updated according to a preset pulling method. When it is detected that the update operation is completed, the updated inference framework code is read from the storage path. Based on the updated inference framework code, the inference framework container is restarted. In this way, only the original inference framework code needs to be updated, and only the inference framework container needs to be restarted after the update. Even in the case of no additional acceleration card, the updated inference framework code can be applied in a timely manner, that is, the update of the inference framework service can be completed in a timely manner. In addition, in the related art, the container group needs to be restarted. Since the inference engine is included in the container group and the startup time of the inference engine is relatively long, the speed of restarting the container group is slow. However, in this solution, only the inference framework container needs to be restarted, which can improve the update speed of the overall service and apply the new inference framework code more timely.
[0072] Before executing the above service update method, the computing node can first deploy the service. Accordingly, the embodiment of the present application provides a service deployment method, which can be executed by any one of the above computing nodes. As Figure 3As shown in the figure, the specific processing steps of the service update method may include: Step S301: Receive a service deployment request sent by the client.
[0073] Among them, the service deployment request includes annotation information.
[0074] Specifically, the user can write a configuration file on the client. For example, the configuration file can be a file in the "yaml" format, and the name of the configuration file can be "service.yaml". Correspondingly, the user can input the first instruction and the name of the configuration file on the client (for example, kubectl apply –f your-service.yaml). The client can generate a service deployment request according to the first instruction and the name of the configuration file and send it to the client. For example, the first instruction can be a "kubectl instruction". Or, the user can click a target button in the relevant service interface of the client, and the client can automatically call the first interface (for example, it can be the kubernetes api-server) to generate a service deployment request and send it to the computing node. After receiving the service deployment request sent by the client, the computing node can parse the service deployment request to obtain the annotation information, and then perform different service deployment operations according to the different annotation information.
[0075] Step S302: Determine whether to perform the deployment operation of the inference framework service according to the annotation information.
[0076] Specifically, the computing node can determine whether to perform the deployment operation of the inference framework service according to the content included in the annotation information. Correspondingly, the computing node can perform the following specific steps: Step 1: Determine whether the annotation information includes a preset field.
[0077] Among them, the preset field can be the "LLM-service field".
[0078] Step 2: When determining whether the annotation information includes a preset field, determine to perform the deployment operation of the inference framework service.
[0079] Or, Step 3: When determining that the annotation information does not include a preset field, determine not to perform the deployment operation of the inference framework service.
[0080] Specifically, the computing node can first determine whether the annotation information includes the "LLM-service field". If so, it is determined that the current LLM inference service requires the prior deployment of the inference framework service. If not, it is determined that the current is performing other services, and it can be processed according to the deployment method of other services without special processing of this solution.
[0081] Step S303, when it is determined to perform the deployment operation of the inference framework service, obtain the latest version of the inference framework code from the preset storage location according to the preset update method.
[0082] Method 1: Use a pre-built code pulling tool to pull the latest version of the inference framework code from the target code repository.
[0083] Method 2: Obtain the second storage path corresponding to the preset storage location. Based on the preset instruction template and the second storage path, generate a code reading instruction. Send the code reading instruction to the target storage node to notify the target storage node to search for and send the latest version of the inference framework code. Receive the latest version of the inference framework code sent by the target storage node.
[0084] Method 3: Extract the latest version of the inference framework code from a pre-built initialization full-pull container.
[0085] Specifically, in Method 1, the computing node only needs to input the service deployment instruction into the code pulling tool, which can trigger the code pulling tool to pull the latest version of the inference framework code from the target code repository. In this way, it is relatively convenient.
[0086] In Method 2, the computing node can add the second storage path to the first specified position of the preset instruction template to generate a code acquisition instruction and send it to the target storage node. In this way, when the target storage node receives the code acquisition instruction, it can determine the latest version of the inference framework code from various versions of the inference framework code stored in itself and send it to the computing node. The computing node can then obtain the latest version of the inference framework code sent by the target storage node. In this way, obtaining code from the local target storage node has high stability, fast deployment speed, and multiple computing nodes can obtain code from the same storage node simultaneously, which is suitable for large-scale distributed training and inference scenarios.
[0087] In Method 3, the initialization full-pull container can be pre-configured with the latest version of the inference framework code. Correspondingly, when it is determined to deploy the LLM inference service, the latest version of the inference framework code can be extracted from the code position corresponding to the preset position information in itself according to the preset position information. Since the code has been pre-packaged in the container and only needs to be extracted according to the preset position for use, there is almost no network delay, and the efficiency of deploying the inference framework service is relatively high.
[0088] Step S304, store the latest version of the inference framework code under the first storage path.
[0089] Among them, the latest version of the inference framework code is the original inference framework code.
[0090] Specifically, the computing node can directly store the latest version of the inference framework code under the first storage path.
[0091] Step S305: Create and start an inference framework container based on the latest version of the inference framework code read from the first storage path.
[0092] Specifically, the computing node can read the latest version of the inference framework code from the first storage path, create and start an inference framework container. The inference framework container started here can be the original inference framework container mentioned in the above step S204.
[0093] The service deployment method of the embodiments of the present application controls the deployment logic through simple annotations without the need to modify the deployment script or manual intervention, which is relatively simple. Moreover, obtaining the latest version of the inference framework code from the preset storage location ensures that the deployed version is always the latest and verified code version, avoiding functional defects or security issues caused by local caching of old version code. In addition, saving the latest version of the code to the first storage path as the original inference framework code used during runtime enables all inference service container groups that mount the first storage path to use the same version of the original inference framework code, ensuring the consistency of the inference service deployment.
[0094] To facilitate users to view the service deployment information of the inference framework at any time, when executing the above service update method, the computing node can also record the relevant information of the inference framework in the database. Correspondingly, the embodiments of the present application provide a data record storage method, which can be executed by any one of the above computing nodes. As Figure 4 shown, the specific processing steps of the data record storage method may include: Step S401: Obtain the storage resource configuration information corresponding to the first storage path.
[0095] Among them, the storage resource configuration information may be the above-mentioned PVC.
[0096] Specifically, the computing node can determine the PVC that matches the first storage path from at least one pre-constructed PVC as the above-mentioned storage resource configuration information according to the first storage path.
[0097] Step S402: Generate a data record corresponding to the inference framework code according to the storage resource configuration information, service indication information, first storage path, and original version information.
[0098] Specifically, the computing node can combine the storage resource configuration information, service indication information, first storage path, and original version information to obtain a data record corresponding to the inference framework code.
[0099] Step S403: Store the data record in a pre-built target database for querying the service deployment information of the inference framework.
[0100] Specifically, the computing node can store the data record in the pre-built target database, so that the user can view the data record in the target database subsequently.
[0101] In the data record storage method of the embodiment of the present application, by generating a data record and storing it in the target database during the service deployment process, it is convenient for the user to understand the relevant situation of the currently deployed inference framework code at any time. Further, when the user learns the original code version information through the data record, it can be determined whether to perform an update operation, so as to avoid the problem of wasting computing node resources caused by inputting the update version information consistent with the original version information without any understanding by the user.
[0102] The above service deployment method, data record storage method, and service update method will be described in detail below with a specific example.
[0103] As Figure 5 shown, the LLM inference service container group 1 and the LLM inference service container group 2 can be deployed on the computing node. Both LLM inference service container groups can include an initialization full-pull container, an inference engine container, and an inference framework container. The inference framework code used by the inference framework container can be managed by the code management container in the code management container group, that is, the user can manage the inference framework code stored in the storage locations corresponding to the lower-level paths of the storage paths mounted by the code management container group through the user interface (UI). The following takes any one of the LLM inference service container groups as an example for description.
[0104] In the architecture as Figure 5 shown, the above operations as Figures 2 to 4 shown can be performed by the computing node using different program components (for example, the intermediate components mentioned below) and containers.
[0105] Step S201, Step S301, and Step S302 can be executed by the computing node using an intermediate component (for example, it can be a Webhook), Step S202 and Step S203 are executed by the computing node using the code management container, and Step S204 is executed by the computing node using the inference framework container. Steps S303 to S305 can be executed by the computing node using the initialization full-pull container. Steps S401 to S403 can be executed by the code management container.
[0106] In the service deployment stage, the process of the service deployment method can be as Figure 6 shown.
[0107] Users can create an LLM inference service through "kubectl commands" or the "Kubernetes API". Correspondingly, based on the user's operation, the client sends a service deployment request to the computing node.
[0108] The intermediate component executes the above steps S301 and S302. When the intermediate component of the computing node receives the service deployment request, it can determine whether the "annotations" included in the service deployment request (i.e., the above-mentioned annotation information) include the "LLM-service field". When the intermediate component determines that the "annotations" include the "LLM-service field", it determines that an LLM inference service needs to be deployed currently. Or, when the intermediate component determines that the "annotations" do not include the "LLM-service field", it determines that other services need to be deployed currently (this solution does not interfere in this case). When it is determined that an LLM inference service needs to be deployed currently, the intermediate component can extract the service scenario name, service name, PVC name, etc. from the service deployment request and send them to the code management container. And the intermediate component can send an inference service deployment request to the initialization full pull container in the LLM inference service container group.
[0109] The initialization full pull container executes the above steps S303 to S305. When the initialization full pull container receives the inference service deployment request, it can perform the operation of pulling the inference framework code, store the pulled inference framework code in the storage path it mounts (i.e., the storage path mounted by the inference service container group it belongs to, for example, the above-mentioned first storage path), and then, read the inference framework code from this storage path and create and start an inference framework container based on this inference framework code.
[0110] The code management container executes the above steps S401 to S403. The code management container can determine the storage path corresponding to the inference framework code according to the received service scenario name and PVC name, and determine the version information of the pulled inference framework code based on this storage path. In addition, the code management container can also obtain the PVC corresponding to this LLM inference service container group. Finally, the code management container can generate a data record according to the service scenario name, service name, PVC name, the version information of the pulled inference framework code, and the PVC, and store it in the database.
[0111] In the service update stage, the process of the service update method can be as Figure 7 shown.
[0112] The intermediate component executes step S201. After receiving the service update request sent by the client, the intermediate component can forward the service update request to the code management container.
[0113] The code management container executes step S202 and step S203. When receiving the service update request forwarded by the intermediate component, the code management container can update the inference framework code stored under the storage path mounted by the LLM inference service container group through a preset update method.
[0114] The inference framework container executes step S204. Unicorn WSGI (WebServer Gateway Interface) code can be set in the inference framework container. Unicorn WSGI code is a kind of Web Server. The “–reload parameter” can be set in this code. When this parameter is started, Unicorn can monitor whether the inference framework code is updated. When it is monitored that the inference framework code is updated, it triggers the Unicorn WSGI reload, and then triggers the restart of the inference framework container.
[0115] After the inference framework container restarts, it can receive the registration request periodically sent by the inference engine container. The inference framework container can extract the port information from the registration request. Furthermore, based on the port information, it monitors the relevant information of the inference engine container. For example, it obtains the model name, version loaded by the inference engine container, as well as the statistical information of relevant performance metrics, health check information, etc.
[0116] A proxy mode can be adopted between the inference framework container and the inference engine container, that is, an inference engine service object is abstracted from the inference framework service to implement the above interaction function. Among them, abstracting an inference engine service object can be a code class, and this code class can include methods corresponding to the operations that the inference engine code can implement, so that the inference framework container can operate the inference engine container just like operating its own internal code. This code class can be written using the Remote Procedure Call (gRPC) protocol or the HyperText Transfer Protocol (HTTP).
[0117] When this code class is written using the gRPC protocol, the code class can be as follows: syntax = "proto3"; package vllm_proxy; service VLLMService { / / Method for creating an inference engine container rpc CreateEngine(EngineArgs) returns (EngineResponse); / / Method for generating text, a streaming RPC, i.e., the computing node can return multiple responses (streaming output) rpc Generate(GenerateRequest) returns (stream GenerateResponse); / / Method for health checking of the inference engine container rpc CheckHealth(HealthCheckRequest) returns (HealthCheckResponse); / / Method for shutting down the inference engine container rpc ShutdownEngine(ShutdownRequest) returns (ShutdownResponse); } / / Configuration parameters corresponding to AsyncEngineArgs, parameters that need to be configured when creating the inference engine container message EngineArgs { string model = 1; bool enable_lora = 2; int32 max_loras = 3; string tokenizer_mode = 4; bool trust_remote_code = 5; int32 tensor_parallel_size = 6; int32 block_size = 7; int32 swap_space = 8; float gpu_memory_utilization = 9; int32 max_num_seqs = 10; string quantization = 11; int32 max_model_len = 12; map<string, int32>limit_mm_per_prompt = 13; string guided_decoding_backend = 14; string scheduling_policy = 15; bool enable_prefix_caching = 16; } / / Response method after creating the inference engine container message EngineResponse { string engine_id = 1; bool success = 2; string error_message = 3; } / / Low-Rank Adaptation (LoRA) request, used to specify LoRA in the generation request message LoRARequest { string lora_name = 1; int32 lora_int_id = 2; string lora_local_path = 3; } / / Method for sampling parameters message SamplingParams { string lora_name = 1; int32 n = 2; int32 best_of = 3; float presence_penalty = 4; float frequency_penalty = 5; float temperature = 6; float top_p = 7; int32 top_k = 8; int32 max_tokens = 9; repeated string stop = 10; repeated int32 stop_token_ids = 11; bool skip_special_tokens = 12; / / Guiding the decoding to generate specific parameters GuidedDecodingParams guided_decoding = 13; } message GuidedDecodingParams { string json = 1; string regex = 2; repeated string choice = 3; string grammar = 4; bool json_object = 5; string backend = 6; string whitespace_pattern = 7; } / / Multimodal data message MultiModalData { repeated string image = 1; } / / Generation request message GenerateRequest { string engine_id = 1; oneof input { string text_prompt = 2; ComplexPrompt complex_prompt = 3; } SamplingParams sampling_params = 4; string request_id = 5; LoRARequest lora_request = 6; } message ComplexPrompt { string prompt = 1; MultiModalData multi_modal_data = 2; } / / Generated Token Output message OutputToken { int32 index = 1; string text = 2; repeated float logprobs = 3; string finish_reason = 4; repeated int32 token_ids = 5; } / / Generate Response message GenerateResponse { string request_id = 1; repeated OutputToken outputs = 2; repeated int32 prompt_token_ids = 3; bool is_final = 4; string error_message = 5; } / / Health Check Request message HealthCheckRequest { string engine_id = 1; } message HealthCheckResponse { bool is_healthy = 1; string status_message = 2; } / / Request to close the inference engine container message ShutdownRequest { string engine_id = 1; } message ShutdownResponse { bool success = 1;} The computing node can set up a monitoring component (e.g., it can be a liveness Probe), set the monitoring period duration of this monitoring component to a value greater than 15 s, and set the failure threshold (Failure Thresholds) of this monitoring component to a value greater than 2. In this way, this monitoring component can monitor whether the LLM inference service container group is running normally according to the monitoring period duration. When the LLM inference service container group is detected to be running abnormally for multiple consecutive periods (greater than 2), then restart the LLM inference service container group. Since it takes a certain amount of time to restart the inference framework container, during this period, the LLM inference service container group is running abnormally. If the LLM inference service container group is directly restarted, it will lead to a longer restart time and introduce unnecessary time consumption. At the same time, in order to be able to complete the necessary health checks, periodic checks can be carried out according to the above-set monitoring period duration.
[0118] Multiple inference engine containers can be deployed in the inference service container group. Therefore, in order to prevent conflict problems, random ports can be assigned to each inference engine container. In addition, for the convenience of management, the inference framework container can specify the port using environment variables, and the port in this environment variable can be consistent with the target Port parameter during service deployment. In this way, each inference engine container can determine the port of the inference framework container through the target Port parameter, and based on this port, send a registration request to the inference framework container. The inference framework container can extract the port information from the registration request and, based on the port information, take over each inference engine container included in its affiliated inference service container group again.
[0119] In the related technology, in order to facilitate the control of the resources occupied by the inference engine, generally, the inference framework is used as the parent process and the inference engine is used as the sub-process. Therefore, when updating the inference framework code, it is necessary to end both the process where the inference framework is located and the process where the inference engine is located. However, restarting the inference engine takes a long time, resulting in a long interruption time of the inference service. By isolating the inference framework and the inference engine, that is, separating the inference framework container and the inference engine container into two completely independent containers, and using the proxy mode to let the inference framework container operate the inference engine container, during the update, there is no need to restart or shut down the inference engine container, and the restart speed is relatively fast, and the new inference framework code can be applied to the inference service in a timely manner. The code management container is specifically used to execute the related functions of code management, which can realize the centralized management of the code of all service container groups, facilitate user management, and can also realize flexible control of the code version.
[0120] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0121] An embodiment of the present application further provides a service update device, as Figure 8 shown, including: A receiving module 810, configured to receive a service update request sent by a client, where the service update request includes service indication information, location indication information, and update version information; A determining module 820, configured to determine the original version information of the original inference framework code according to the service indication information and the location indication information; An updating module 830, configured to, when it is determined to perform an update operation according to the update version information and the original version information, perform an update operation on the original inference framework code according to a preset update method and the update version information; A restart module 840, configured to, when it is detected that the update operation is completed, restart the inference framework container based on the updated inference framework code to complete the update operation of the inference framework service.
[0122] In some optional implementation manners, the determining module 820 is specifically configured to: Determine the first storage path of the original inference framework code according to the service indication information and the location indication information; Read the original version information based on the first storage path.
[0123] In some optional implementation manners, the determining module 820 is specifically configured to: Obtain the second storage path mounted by the pre-built code management container, where the lower-level path of the second storage path includes at least one sub-path; Determine, according to the service indication information and the location indication information, a sub-path that matches both the service indication information and the location indication information in the at least one sub-path as the first storage path.
[0124] In some optional implementation manners, the determining module 820 is specifically configured to: Read the original inference framework code from the storage location corresponding to the first storage path based on the first storage path; Determine a target field corresponding to the preset field identification information in the original inference framework code according to the preset field identification information; Extract the original version information from the target field.
[0125] In some optional implementation manners, the updating module 830 is specifically configured to: Obtain the update code of the inference framework from the preset storage location according to the preset update method and update version information; Perform an update operation on the original inference framework code under the first storage path based on the update code.
[0126] In some alternative embodiments, when the preset update method is an incremental update method, the preset storage location is the target code repository provided by the target code platform, and the update code is an incremental update code; the update module 830 is specifically configured to: Incrementally pull the incremental update code from the target code repository by using a pre-built code pulling tool and the update version information.
[0127] In some alternative embodiments, when the preset update method is a full update method, the preset storage location is the target storage node, and the update code is the full update code corresponding to the update version information; the update module 830 is specifically configured to: Obtain a second storage path corresponding to the target storage node; Generate a code update instruction according to the preset instruction template, the second storage path, and the update version information; Send the code update instruction to the target storage node to notify the target storage node to obtain the full update code corresponding to the update version information; Receive the full update code corresponding to the update version information sent by the target storage node.
[0128] In some alternative embodiments, the update module 830 is specifically configured to: Replace the original inference framework code stored under the first storage path with the full update code.
[0129] In some alternative embodiments, the apparatus further includes a deployment module 850, which is configured to: Receive a service deployment request sent by a client, where the service deployment request includes annotation information; Determine whether to perform a deployment operation of the inference framework service according to the annotation information; When it is determined to perform a deployment operation of the inference framework service, obtain the inference framework code of the latest version from the preset storage location according to the preset update method; Store the inference framework code of the latest version under the first storage path, where the inference framework of the latest version is the original inference framework code; Create and start an inference framework container based on the inference framework code of the latest version read from the first storage path.
[0130] In some alternative embodiments, the deployment module 850 is specifically configured to: Determine whether the annotation information includes a preset field; When it is determined whether the annotation information includes a preset field, it is determined to perform an inference framework service deployment operation; Or, When it is determined that the annotation information does not include a preset field, it is determined not to perform an inference framework service deployment operation.
[0131] In some alternative embodiments, the apparatus further includes a storage module 860, configured to: Obtain storage resource configuration information corresponding to a first storage path; Generate a data record corresponding to the inference framework code according to the storage resource configuration information, service indication information, the first storage path, and the original version information; Store the data record in a pre-constructed target database for querying service deployment information of the inference framework.
[0132] In some alternative embodiments, the apparatus further includes a monitoring module 870, configured to: Obtain port information corresponding to each of at least one pre-constructed inference engine container; Re-take over the inference engine container corresponding to the port information based on the port information.
[0133] For the description of the features in the embodiments corresponding to the service update apparatus, reference may be made to the relevant descriptions in the embodiments corresponding to the service update method, which will not be elaborated here one by one.
[0134] An embodiment of the present application further provides an electronic device, as Figure 9 shown, including a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any one of the above service update method embodiments. The electronic device may be the above computing node.
[0135] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above service update method embodiments when running.
[0136] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc, etc., various media that can store a computer program.
[0137] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described service update method embodiments are implemented.
[0138] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described service update method embodiments are implemented.
[0139] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0140] The above has introduced in detail a service update method, apparatus, electronic device, storage medium, and program product provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A service update method, characterized in that, Including: Receiving a service update request sent by a client, where the service update request includes service indication information, location indication information, and update version information; Determining the original version information of the original inference framework code according to the service indication information and the location indication information; When it is determined to perform an update operation according to the update version information and the original version information, performing an update operation on the original inference framework code according to a preset update method and the update version information; When it is detected that the update operation is completed, restarting the inference framework container based on the updated inference framework code to complete the update operation of the inference framework service.
2. The service update method according to claim 1, wherein The determining the original version information of the original inference framework code according to the service indication information and the location indication information includes: Determining a first storage path of the original inference framework code according to the service indication information and the location indication information; Reading the original version information based on the first storage path.
3. The service update method according to claim 2, wherein The determining the first storage path of the original inference framework code according to the service indication information and the location indication information includes: Obtaining a second storage path mounted by a pre-built code management container, where a lower-level path of the second storage path includes at least one sub-path; Determining, according to the service indication information and the location indication information, a sub-path that matches both the service indication information and the location indication information from at least one of the sub-paths as the first storage path.
4. The service update method according to claim 2, characterized in that The reading the original version information based on the first storage path includes: Reading the original inference framework code from a storage location corresponding to the first storage path based on the first storage path; Determining a target field corresponding to the preset field identification information in the original inference framework code according to the preset field identification information; Extracting the original version information from the target field.
5. The service update method according to any one of claims 2 to 4, characterized in that, The performing an update operation on the original inference framework code according to a preset update method and the update version information when it is determined to perform an update operation according to the update version information and the original version information includes: Obtaining an update code of the inference framework from a preset storage location according to the preset update method and the update version information; Performing an update operation on the original inference framework code under the first storage path based on the update code.
6. The service update method according to claim 5, wherein When the preset update method is an incremental update method, the preset storage location is a target code repository provided by a target code platform, and the update code is an incremental update code; The obtaining an update code of the inference framework from the preset storage location according to the preset update method and the update version information includes: Incrementally pulling the incremental update code from the target code repository by using a pre-built code pulling tool and the update version information.
7. The service update method according to claim 5, wherein When the preset update method is a full-volume update method, the preset storage location is a target storage node, and the update code is a full-volume update code corresponding to the update version information; Obtaining an update code of an inference framework from the preset storage location according to the preset update method and the update version information includes: Obtaining a second storage path corresponding to the target storage node; Generating a code update instruction according to a preset instruction template, the second storage path, and the update version information; Sending the code update instruction to the target storage node to notify the target storage node to obtain the full update code corresponding to the update version information; Receiving the full update code corresponding to the update version information sent by the target storage node.
8. The service update method according to claim 7, wherein Performing an update operation on the original inference framework code under the first storage path based on the update code, including: Replacing the original inference framework code stored under the first storage path with the full update code.
9. The service update method according to claim 5, wherein Before receiving a service update request sent by a client, the method further includes: Receiving a service deployment request sent by the client, where the service deployment request includes annotation information; Determining whether to perform a deployment operation of the inference framework service according to the annotation information; When it is determined to perform the deployment operation of the inference framework service, obtaining the inference framework code of the latest version from the preset storage location according to the preset update method; Storing the inference framework code of the latest version under the first storage path, where the inference framework of the latest version is the original inference framework code; Creating and starting the inference framework container based on the inference framework code of the latest version read from the first storage path.
10. The service update method according to claim 9, characterized in that, Determining whether to perform a deployment operation of the inference framework service according to the annotation information includes: Determining whether the annotation information includes a preset field; When it is determined that the annotation information includes a preset field, determining to perform the deployment operation of the inference framework service; Or, When it is determined that the annotation information does not include a preset field, determining not to perform the deployment operation of the inference framework service.
11. The service update method according to claim 9, wherein The method further includes: Obtaining storage resource configuration information corresponding to the first storage path; Generating a data record corresponding to the inference framework code according to the storage resource configuration information, the service indication information, the first storage path, and the original version information; Storing the data record in a pre-constructed target database for querying service deployment information of the inference framework.
12. The service update method according to any one of claims 1 to 4, characterized in that, After detecting that the update operation is completed, restarting the inference framework container based on the updated inference framework code to complete the update operation of the inference framework service, the method further includes: Obtaining port information corresponding to at least one pre-constructed inference engine container respectively; Re-taking over the inference engine container corresponding to the port information based on the port information.
13. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the service update method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the service update method according to any one of claims 1 to 12 are implemented.
15. A computer program product, characterized in that, The computer program product includes a computer program, wherein when the computer program is executed by a processor, the steps of the service update method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Configuration information updating method and device under micro-service architecture and electronic equipment
CN112882738A
Multi-model hot updating method and device
CN115129426A
Microservice code version monitoring method and device, equipment, medium and product
CN118939308A
Method for updating of service module in extension service framework and the server using the same
KR102204581B1
Method and system for service rolling-updating in a container orchestrator system
US20210072966A1