Fine-tuning model dynamic loading method and related device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-08-07
AI Technical Summary
然而,此方式需要对每个业务场景均进行独立部署,需要很高的资源成本,且推理服务的运维效率较低
[0008]上述方案,通过控制器响应于在容器云平台增改目标自定义资源,调用推理服务单元中的辅助容器触发微调模型的加载流程,辅助容器利用微调模型的模型相关信息从预设模型库下载微调模型,并将微调模型存储在共享存储卷,从而调用基础容器从共享存储卷动态加载微调模型并将微调模型并入基础容器的基础模型中,使得推理服务单元具备目标推理服务,通过在容器云平台的推理服务单元内设置基础容器与辅助容器的双容器架构,结合目标自定义资源,实现了微调模型的自动化、解耦式的动态加载,不仅能够借助辅助容器独立完成微调模型的下载与共享存储卷的写入操作,避免对基础容器内基础模型运行状态的干扰,保障推理服务的连续性与稳定性,还能通过共享存储卷让基础容器快速获取并融合微调模型,高效完成基础模型权重参数的优化与目标推理服务的部署,大幅提升了容器云平台中模型迭代与推理服务更新的灵活性、便捷性和资源利用率,提高推理服务的运维效率以及减少资源消耗。
Smart Images

Figure CN122526641A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large model technology, and in particular to a dynamic loading method, computer device and storage medium for fine-tuning models. Background Technology
[0002] With the rapid development of large models, in order to adapt large models to the inference needs of different business scenarios, it is usually necessary to fine-tune the training based on the general base model to obtain a fine-tuned model adapted to specific tasks, and then deploy it to the online inference platform to provide services.
[0003] Currently, for various business scenarios, the usual approach is to first merge the adapted, fine-tuned model with the base model to generate a new, larger model, and then deploy the larger model. However, this method requires independent deployment for each business scenario, incurring high resource costs, and the operational efficiency of the inference service is relatively low. Summary of the Invention
[0004] The main technical problem addressed in this application is to provide a dynamic loading method and related apparatus for fine-tuning models, which can automate the dynamic loading process of fine-tuning models, improve the operational efficiency of inference services, and reduce resource consumption.
[0005] The first aspect of this application provides a dynamic loading method for a fine-tuned model. The method includes: a controller, in response to adding or modifying a target custom resource on a container cloud platform, calling an auxiliary container in an inference service unit to trigger a loading process for the fine-tuned model; wherein the target custom resource contains model-related information of the fine-tuned model; the auxiliary container uses the model-related information of the fine-tuned model to download the fine-tuned model from a preset model library and stores the fine-tuned model in a shared storage volume; and a base container is called to dynamically load the fine-tuned model from the shared storage volume and incorporate the fine-tuned model into the base model of the base container, enabling the inference service unit to have a target inference service, wherein the fine-tuned model is used to fine-tune the weight parameters of the base model to achieve the target inference service.
[0006] A second aspect of this application provides a computer device including a memory and a processor coupled to each other, the memory storing program data and the processor executing the program data to implement the steps performed by at least one of the controller, auxiliary container, and base container in the dynamic loading method of the above-described fine-tuning model.
[0007] A third aspect of this application provides a computer-readable storage medium storing program data, which, when executed by a processor, implements the steps performed by at least one of the controller, auxiliary container, and base container in the dynamic loading method of the above-described fine-tuning model.
[0008] The above solution, through the controller's response to adding or modifying target custom resources on the container cloud platform, calls the auxiliary container in the inference service unit to trigger the loading process of the fine-tuned model. The auxiliary container uses the model-related information of the fine-tuned model to download the fine-tuned model from the preset model library and stores the fine-tuned model in the shared storage volume. This then calls the basic container to dynamically load the fine-tuned model from the shared storage volume and merge the fine-tuned model into the basic model of the basic container, enabling the inference service unit to have the target inference service. By setting up a dual-container architecture of basic container and auxiliary container in the inference service unit of the container cloud platform, combined with the target custom resources, the automated and decoupled dynamic loading of the fine-tuned model is realized. It can not only independently complete the download of the fine-tuned model and the writing operation of the shared storage volume with the help of the auxiliary container, avoiding interference with the running state of the basic model in the basic container and ensuring the continuity and stability of the inference service, but also allow the basic container to quickly obtain and merge the fine-tuned model through the shared storage volume, efficiently complete the optimization of the basic model weight parameters and the deployment of the target inference service. This greatly improves the flexibility, convenience and resource utilization of model iteration and inference service updates in the container cloud platform, improves the operation and maintenance efficiency of the inference service and reduces resource consumption.
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them: Figure 1 This is a flowchart illustrating the first embodiment of the dynamic loading method for fine-tuning the model in this application; Figure 2 This is an example schematic diagram of an embodiment of the container cloud platform of this application; Figure 3 This is an example schematic diagram of another embodiment of the container cloud platform of this application; Figure 4 This is an example schematic diagram of an embodiment of the dynamic loading method for fine-tuning the model in this application; Figure 5 This is an example schematic diagram of another embodiment of the dynamic loading method for fine-tuning the model in this application; Figure 6 This is a flowchart illustrating the second embodiment of the dynamic loading method for fine-tuning the model in this application; Figure 7 This is a flowchart illustrating the third embodiment of the dynamic loading method for fine-tuning the model in this application; Figure 8 This is a schematic diagram of the structure of an embodiment of the computer device of this application; Figure 9 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0012] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0013] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0014] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0015] Through long-term research, the inventors of this application have discovered that in large-scale business scenarios, it is often necessary to fine-tune the model to cope with the explosive growth of the number of parameters. The large-scale LoRA (Low-Rank Adaptation) model can be compressed into the product of two small matrices through low-rank decomposition, as follows: For various business scenarios, the large model can be split into a LoRA model and a base model. During inference deployment, the LoRA model and the base model need to be merged to generate a new model, and then the large model is deployed. This significantly reduces computing costs and storage requirements while maintaining model performance.
[0016] However, deploying multiple business scenarios independently requires high resources. If the models of multiple business scenarios are fine-tuned based on the same base model and share the weight parameters of the base model, this will lead to a waste of storage resources.
[0017] To address the aforementioned technical issues, this application provides a dynamic loading method and related apparatus for fine-tuning models, which can automate the dynamic loading process of fine-tuning models, improve the operational efficiency of inference services, and reduce resource consumption.
[0018] Please see Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the dynamic loading method for fine-tuning the model in this application. The method may include the following steps: S11: In response to adding or modifying the target custom resource on the container cloud platform, the controller calls the auxiliary container in the inference service unit to trigger the loading process of the fine-tuning model; wherein, the target custom resource contains model-related information of the fine-tuning model.
[0019] In this embodiment, the fine-tuning model is obtained by fine-tuning the target inference service based on the base model and the fine-tuning model. The fine-tuning model may include, but is not limited to, the LoRA model. The following embodiments of this application use the LoRA model as an example for illustration. The base model may include, but is not limited to, a large model or a large language model. For each target inference service (such as each business scenario), fine-tuning training of the fine-tuning model and the base model is required to generate a fine-tuning model corresponding to each target inference service. During the fine-tuning training process, the weight parameters of the base model remain unchanged, and only the two low-rank matrices (such as A and B) of the fine-tuning model need to be trained to achieve fine-tuning. Multiple target inference services correspond to the same base model or multiple base models. Multiple target inference services may be fine-tuned based on the same base model, sharing the weight parameters of the base model, with only the fine-tuning model being different.
[0020] The fine-tuning model is used to fine-tune the weight parameters of the base model to achieve the target inference service. For example, the weight parameters of the base model are... The fine-tuned model is as follows: The parameter update amount can be determined using two low-rank matrices (A and B) of the fine-tuning model. Then use parameter update amount Weight parameters of the base model Fine-tuning is performed to obtain a new base model, which serves as the target large model. The weight parameters of the target large model are as follows: The target large model is used to implement the corresponding target inference service.
[0021] Please see Figure 2 The dynamic loading method for fine-tuning models described in this application is applied to container cloud platforms, such as, but not limited to, Kubernetes (K8s) platforms. A container cloud platform includes a controller (Kubernetes operator) and at least one inference service unit (Pod). Each inference service unit contains a base container (model-base container) and auxiliary containers (LoRA Runtime Sidecar containers). A Pod is the smallest deployable computing unit created and managed in Kubernetes. A Pod contains a group (or more) of containers that can share network, storage, and other environmental resources within the Pod. The controller (Kubernetes operator) can be used to control the entire process of the dynamic loading method for fine-tuning models. Optionally, the base container can be, but is not limited to, the main application container, and the auxiliary container can be, but is not limited to, the sidecar container. In Kubernetes, the sidecar container is a container added next to the main application container to extend or modify the inference service capabilities of the main application container. The sidecar container shares the same Pod resources as the main application container, but they run in different containers and can independently manage their own lifecycles and configurations. The auxiliary container can interact with the base container to achieve dynamic loading and unloading of the LoRA model. The auxiliary container can also interact with the preset model library to download LoRA model artifact files from the preset model library to the local shared storage volume, making the LoRA model artifact files visible in the base container.
[0022] The container cloud platform also includes a pre-built model library, such as Harbor, which stores the artifact files of the models, such as the artifact files of the fine-tuned models and / or the artifact files of the base models.
[0023] Container cloud platforms also include callback service interfaces, which may include, but are not limited to, webhooks. These callback service interfaces are used to verify / intercept Kubernetes resource functionalities. For example, when a Kubernetes Pod is created, a webhook intercepts the basic service resource definitions used to create the Pod, enabling the automatic injection of auxiliary containers into the Pod.
[0024] The container cloud platform also includes a basic inference service interface (Service model-base) and a target inference service interface (Service model-LoRA). The basic inference service interface is used to connect the inference service unit to the basic inference service, and the target inference service interface is used to connect the inference service unit to the target inference service.
[0025] The LoRA operator monitors resource changes of the target custom resource, which is a custom resource (CR) adapted to the fine-tuning model, denoted as LoraAdapter CR. In response to additions or modifications to the target custom resource LoraAdapter CR on the container cloud platform, the LoRA operator invokes the LoRA Runtime Sidecar container within the inference service unit Pod to trigger the loading process of the fine-tuning model. The target custom resource LoraAdapter CR contains model-related information for the fine-tuning model.
[0026] In some implementations, the model-related information in the target custom resource LoraAdapter CR includes a first identifier for the fine-tuning model, a second identifier for the base model, address information, and the port of the container corresponding to the LoRA model (e.g., the port of the base container and / or auxiliary container corresponding to the LoRA model). The address information is the storage path of the fine-tuning model in a preset model library. The first identifier identifies the fine-tuning model, and the second identifier identifies the base model. The fine-tuning model is associated with the base model. For example, the first identifier of the fine-tuning model is the name of the fine-tuning model, and the second identifier of the base model is the name of the base model.
[0027] For example, the YAML file (configuration file) of the target custom resource LoraAdapter CR can contain the definition of the LoRA model, the name of the base model associated with the LoRA model, the Pod tag selector of the base model associated with the LoRA model, and the address information of the LoRA model in Harbor. The definition of the LoRA model includes the name of the LoRA model, the namespace (the logical isolation unit of the resource), and the resource tag (such as the loraadapter tag, which can customize the attributes of the LoRA model). The resource tag includes the name of the LoRA model and the port of the container corresponding to the LoRA model. This application does not limit the specific target custom resource.
[0028] Optionally, the main fields of the above YAML file may include the name of the LoRA model, the name of the base model that the LoRA model depends on or is associated with, and the address information of the LoRA model's artifact file in Harbor.
[0029] After the Pod is deployed, you only need to create a LoraAdapter CR resource in Kubernetes, and the LoRA operator can automatically complete the entire process of dynamically loading the LoRA model in the LoraAdapter CR resource.
[0030] In some implementations, the base model of the second identifier can be determined as the target base model based on the target custom resource LoraAdapter CR, and the loading process of the fine-tuning model can be triggered by calling the auxiliary container in the inference service unit Pod corresponding to the target base model. Alternatively, the target container can be determined based on the port of the container corresponding to the LoRA model, based on the target custom resource LoraAdapter CR, and the loading process of the fine-tuning model can be triggered by calling the auxiliary container in the inference service unit Pod corresponding to the target container.
[0031] In some implementations, the LoRA operator, in response to adding or modifying a target custom resource (LoraAdapter CR) on the container cloud platform, retrieves several loaded models with the same base model. These loaded models are the fine-tuned models loaded by auxiliary containers within the inference service unit (Pod). The same base model corresponds to the base model associated with the fine-tuned models in the LoraAdapter CR. Specifically, it retrieves the base model corresponding to the LoraAdapter CR, queries several loaded LoraAdapter CR resources, and filters several LoRA models associated with the same base model, thus obtaining several loaded models, which are also the LoRA models loaded by auxiliary containers within the inference service unit (Pod).
[0032] Using several loaded models, determine whether the fine-tuning model (such as the LoRA model) in the target custom resource LoraAdapter CR meets the conditions for successful loading. These conditions include name duplication with several loaded models and the loading status of the loaded models being "loading complete" (or "loading successful"). If the LoRA model has the same name as several loaded models and its loading status is "loading successful," it indicates that the LoRA model has been loaded and cannot be loaded again. In response to the fine-tuning model meeting the successful loading conditions, the process returns a "repeat loading" status for the LoRA model and ends. Otherwise, the LoRA model is not loaded. In response to the fine-tuning model not meeting the successful loading conditions, the corresponding auxiliary container (LoRA Runtime Sidecar container) in the inference service unit Pod is invoked to trigger the dynamic loading process of the LoRA model.
[0033] S12: The auxiliary container uses the model-related information of the fine-tuning model to download the fine-tuning model from the preset model library and stores the fine-tuning model on the shared storage volume.
[0034] The LoRA Runtime Sidecar container, an auxiliary container, uses model-related information from the fine-tuned model to download the artifact files of the fine-tuned model from the pre-defined model library and stores them on a shared storage volume. Within the inference service unit Pod, the base container (model-base container) and the LoRA Runtime Sidecar container share the same shared storage volume. Files stored on this shared storage volume are visible to both the base container (model-base container) and the LoRA Runtime Sidecar container, allowing both to read or write files to the shared storage volume.
[0035] In some implementations, the LoRA Runtime Sidecar container uses the address information of the fine-tuning model (the field parameter artifactURL) to download the artifact file of the fine-tuning model from the pre-defined model library Harbor, and stores the artifact file of the fine-tuning model in the shared storage volume of the inference service unit Pod.
[0036] S13: Call the base container to dynamically load the fine-tuned model from the shared storage volume and merge the fine-tuned model into the base model of the base container, so that the inference service unit has the target inference service. The fine-tuned model is used to fine-tune the weight parameters of the base model to achieve the target inference service.
[0037] The auxiliary container calls the base container in the Pod to dynamically load the LoRA model from the shared storage volume. For example, the auxiliary container sends a notification to the base container that the LoRA model has been downloaded. The base container reads the LoRA model from the shared storage volume, loads the LoRA model into the processor (such as GPU, Graphics Processing Unit) corresponding to the base container, and merges the LoRA model into the base model of the base container to obtain the target large model. This enables the inference service unit Pod to have the target inference service. The LoRA model is used to fine-tune the weight parameters of the base model to achieve the target inference service.
[0038] In some implementations, after the LoRA model is dynamically loaded, the LoRA Runtime Sidecarcontainer modifies the loading progress of the LoRA model.
[0039] In some implementations, after the auxiliary container modifies the loading progress of the fine-tuning model, the LoRA operator continuously polls the auxiliary container, the LoRA Runtime Sidecar container, for the loading progress of the fine-tuning model. For example, it can query the interface of the auxiliary container, the LoRA Runtime Sidecar container, to obtain the loading progress of the fine-tuning model. In response to the loading progress of the fine-tuning model being complete, the LoRA operator modifies the loading status of the fine-tuning model in the target custom resource, the LoraAdapter CR, to be complete, and / or creates a target inference service interface (Service) for the inference service unit (Pod), implementing LoRA model inference interface routing. The target inference service interface is used for the inference service unit to access the target inference service.
[0040] In some embodiments, the LoRA operator monitors resource changes of the target custom resource LoraAdapter CR. Resource changes include adding or modifying the target custom resource, such as creating a new target custom resource, updating the target custom resource, or changing the loading status of the fine-tuning model in the target custom resource to an abnormal state.
[0041] For creating or updating target custom resources, the fine-tuned model in the target custom resource can be dynamically loaded directly according to the above embodiments. For cases where the loading status of the fine-tuned model in the target custom resource is changed to an abnormal state, which corresponds to an abnormal restart of the inference service unit Pod, it is necessary to dynamically load the fine-tuned model in the target custom resource in the abnormal state.
[0042] Please see Figure 3After the LoRA model is dynamically loaded, the auxiliary container LoRA Runtime Sidecar container periodically polls the base container model-base container to detect the loading status of the LoRA model in the Pod or the base container model-base container. In response to changes in the loading status of the LoRA model, the loading status of the LoRA model in the LoraAdapter CR resource is modified accordingly.
[0043] In the event of an abnormal restart of the inference service unit Pod, self-recovery of the LoRA model can be implemented, and an abnormal restart operation can be performed in case of Pod failure or other abnormal situations. The auxiliary container LoRA Runtime Sidecar container in the Pod periodically polls the base container (model-base container) to check the loading status of the LoRA models in the base container. Specifically, in response to an abnormal restart of the inference service unit, the auxiliary container retrieves several target custom resources associated with the base model in the base container. It queries the base container to find that the loading status of the associated fine-tuning model is not loaded, and modifies the loading status of the fine-tuning model in the target custom resource to an abnormal state; where the associated fine-tuning model corresponds to the associated target custom resource. For example, the LoRA Runtime Sidecar container in the Pod queries several LoraAdapter CR resources, filters several LoRA models associated with the base model in the base container, iterates through the LoraAdapter CRs of the associated LoRA models, and calls the base container's interface to query the loading status of the corresponding dynamically loaded LoRA models. If the LoRA model is not loaded in the base container due to an abnormal Pod restart, the LoRA Runtime Sidecarcontainer modifies the corresponding LoraAdapter CR to set the loading status of the LoRA model to an abnormal state.
[0044] The controller listens for resource changes in the LoraAdapter CR. If it detects that the loading status of the LoRA model in the LoraAdapter CR has been changed to an abnormal state, the LoRA model in the abnormal state can be used as a fine-tuning model for reloading. The controller calls the LoRA Runtime Sidecar container in the corresponding Pod to trigger the loading process of the fine-tuning model.
[0045] Unlike the above embodiments, in step S11, after the controller responds to resource changes in the LoraAdapter CR (such as adding or modifying the LoraAdapter CR), it queries whether the loading status of the LoRA model in the LoraAdapter CR is abnormal. If the loading status of the LoRA model in the LoraAdapter CR is not abnormal (i.e., normal), other update processes can be processed. If the loading status of the LoRA model in the LoraAdapter CR is abnormal, the LoRA Runtime Sidecar container in the corresponding Pod is invoked to trigger the LoRA model loading process. This loading process can be called the reloading process, and the LoRA Runtime Sidecar container executes the reloading process of the corresponding LoRA model.
[0046] In some implementations, the auxiliary container determines whether the shared storage volume of the inference service unit stores a reloaded fine-tuned model. If a reloaded fine-tuned model is stored, the download process for the fine-tuned model is skipped, and the subsequent steps of calling the base container to dynamically load the fine-tuned model from the shared storage volume and merging the fine-tuned model into the base model of the base container are executed. If no reloaded fine-tuned model is stored, the auxiliary container uses the model-related information of the fine-tuned model to download the fine-tuned model from a preset model library and stores the fine-tuned model in the shared storage volume.
[0047] Then, the base container is invoked to dynamically load the fine-tuned model from the shared storage volume and incorporate the fine-tuned model into the base model of the base container, so that the inference service unit has the target inference service and completes the self-recovery process of the fine-tuned model.
[0048] The above solution, through the controller's response to adding or modifying target custom resources on the container cloud platform, calls the auxiliary container in the inference service unit to trigger the loading process of the fine-tuned model. The auxiliary container uses the model-related information of the fine-tuned model to download the fine-tuned model from the preset model library and stores the fine-tuned model in the shared storage volume. This then calls the basic container to dynamically load the fine-tuned model from the shared storage volume and merge the fine-tuned model into the basic model of the basic container, enabling the inference service unit to have the target inference service. By setting up a dual-container architecture of basic container and auxiliary container in the inference service unit of the container cloud platform, combined with the target custom resources, the automated and decoupled dynamic loading of the fine-tuned model is realized. It can not only independently complete the download of the fine-tuned model and the writing operation of the shared storage volume with the help of the auxiliary container, avoiding interference with the running state of the basic model in the basic container and ensuring the continuity and stability of the inference service, but also allow the basic container to quickly obtain and merge the fine-tuned model through the shared storage volume, efficiently complete the optimization of the basic model weight parameters and the deployment of the target inference service. This greatly improves the flexibility, convenience and resource utilization of model iteration and inference service updates in the container cloud platform, improves the operation and maintenance efficiency of the inference service and reduces resource consumption.
[0049] For ease of understanding, the dynamic loading method for the fine-tuning model described above in this application will be explained below with some specific embodiments. In the following examples, the controller is referred to as LoRA operator, the inference service unit is referred to as Pod, the auxiliary container is referred to as LoRA Runtime Sidecar container, and the target custom resource is referred to as LoraAdapter CR.
[0050] Example 1 Please see Figure 4 The LoRA operator automatically and dynamically loads the LoRA model based on the resource definition of the LoraAdapter CR. For example, in the case of adding, creating, or updating a LoraAdapter CR, this embodiment includes the following steps: 1. The LoRA operator monitors resource changes in the LoraAdapter CR.
[0051] The LoraAdapter CR contains the LoRA model, its associated base model, and the address information of the LoRA model in Harbor.
[0052] 2. Query all LoraAdapter CR lists, filter the LoRA model list associated with the base model in the LoraAdapter CR, and obtain several loaded models with the same base model.
[0053] 3. Check whether the LoRA model has been loaded in LoraAdapter CR and whether the loading status is successful.
[0054] Check if the LoRA model has the same name as a LoRA model in the LoRA model list, and whether the LoRA model was successfully loaded.
[0055] If a LoRA model has the same name and its loading status is successful, it indicates that the LoRA model has already been loaded and cannot be loaded again, thus ending the process.
[0056] 4. In response to the LoRA model not being loaded, call the interface of the LoRA RuntimeSidecar container in the Pod with the corresponding name of the basic model to trigger the dynamic loading process of the LoRA model.
[0057] 5. The LoRA Runtime Sidecar container performs dynamic loading of the corresponding LoRA model.
[0058] The LoRA Runtime Sidecar container downloads the LoRA model artifact file from Harbor based on the LoRA model's address information in Harbor (the field parameter artifactURL).
[0059] After the LoRA model is downloaded, the LoRA model artifact file is saved to the Pod's shared storage volume.
[0060] The interface of the base container is called to dynamically load the LoRA model from the shared storage volume of the Pod. The base container reads the LoRA model file from the shared storage volume and loads the LoRA model into the GPU of the base container, so that the LoRA model is incorporated into the base model in the base container.
[0061] After the LoRA model is dynamically loaded, the LoRA Runtime Sidecar container modifies the loading progress of the LoRA model.
[0062] 6. The LoRA operator continuously polls the LoRA Runtime Sidecar container to query the LoRARuntime Sidecar container interface and obtain the loading progress of the LoRA model.
[0063] In response to the LoRA model's loading progress indicating completion, a LoRA model Service (target inference service interface) is created, implementing the LoRA model's inference interface routing, i.e., the Pod implements the target inference service.
[0064] The above approach, through the Kubernetes operator framework and CRD (Custom Resource Definition), combined with the LoRA Runtime Sidecar container, automates the entire process of dynamic loading of LoRA models, simplifying the management and maintenance of multiple LoRA models.
[0065] Example 2 Please see Figure 5 In the event of an abnormal Pod restart, the LoRA model's self-recovery process is executed. This embodiment includes the following steps: 1. In response to an abnormal Pod restart, the LoRA Runtime Sidecar container in the Pod periodically polls the Pod's base container to check the loading status of the LoRA model.
[0066] The LoRA Runtime Sidecar container queries the LoraAdapter CR list and filters several LoRA models associated with the base model in the base container.
[0067] Iterate through the list of LoraAdapter CRs corresponding to the associated LoRA models, and call the base container to query the loading status of the corresponding LoRA model, that is, query the loading status of the LoRA model corresponding to the base model.
[0068] Modify the loading status of the LoRA model in the corresponding LoRAAdapter CR to an abnormal state. Due to an abnormal Pod restart, the LoRA model was not loaded, so the loading status of the LoRA model in the corresponding LoRA Runtime Sidecar container is modified to an abnormal state.
[0069] 2. The LoRA operator monitors resource changes in the LoraAdapter CR.
[0070] The LoraAdapter CR contains the LoRA model, its associated base model, and the address information of the LoRA model in Harbor.
[0071] 3. The LoRA operator queries whether the loading status of the LoRA model in the LoraAdapter CR is abnormal.
[0072] 4. If the LoRA model loading status in LoraAdapter CR is normal, then proceed with other update processes.
[0073] 5. In response to an abnormal loading status of the LoRA model in LoRaAdapter CR, call the interface of the LoRA Runtime Sidecar container in the Pod corresponding to the basic model name to trigger the LoRA model reloading process.
[0074] 6. The LoRA Runtime Sidecar container executes the reloading process for the corresponding LoRA model.
[0075] The LoRA Runtime Sidecar container determines whether the Pod's shared storage volume contains (or has) the corresponding LoRA model file. If the LoRA model is present in the shared storage volume, the download process is skipped. If the LoRA model is not present in the shared storage volume, the LoRA Runtime Sidecar container downloads the LoRA model artifact file from Harbor based on the model's address information in Harbor (the `artifactURL` field).
[0076] After the LoRA model is downloaded, the LoRA model artifact file is saved to the Pod's shared storage volume.
[0077] The interface of the base container is called to dynamically load the LoRA model from the shared storage volume of the Pod. The base container reads the LoRA model file from the shared storage volume and loads the LoRA model into the GPU of the base container, so that the LoRA model is incorporated into the base model in the base container.
[0078] After the LoRA model is dynamically loaded, the LoRA Runtime Sidecar container modifies the loading progress of the LoRA model.
[0079] 7. The LoRA operator continuously polls the LoRA Runtime Sidecar container to query the LoRARuntime Sidecar container interface and obtain the loading progress of the LoRA model.
[0080] 8. In response to the LoRA model's loading progress indicating that loading is complete, the LoRA model's self-recovery process ends.
[0081] The above method utilizes the LoRA Runtime Sidecar container to periodically detect the loading status of the LoRA model, and uses the Kubernetes operator to reload the LoRA model with an abnormal loading status, thereby achieving self-recovery of the LoRA model inference service.
[0082] In some embodiments, the controller LoRA operator listens for resource changes in the target custom resource LoraAdapter CR. The resource changes also include deleting the target custom resource LoraAdapter CR. In this case, the unloading process of the fine-tuning model can be triggered.
[0083] Please see Figure 6 , Figure 6 This is a flowchart illustrating a second embodiment of the dynamic loading method for fine-tuning the model in this application. The method may include the following steps: S21: In response to deleting the target custom resource on the container cloud platform, the controller calls the auxiliary container in the inference service unit to trigger the unloading process of the fine-tuning model; wherein, the target custom resource contains model-related information of the fine-tuning model.
[0084] When a target custom resource is deleted, the uninstallation process of the fine-tuned model in the target custom resource can be triggered. The target custom resource contains the fine-tuned model, the associated base model, and the address information of the fine-tuned model in the preset model library.
[0085] S22: The auxiliary container uses the model-related information of the fine-tuning model to remove the fine-tuning model from the shared storage volume.
[0086] The auxiliary container of the inference service unit is invoked to remove the fine-tuned model from the shared storage volume of the inference service unit using the model-related information of the fine-tuned model.
[0087] S23: Call the base container to separate the fine-tuning model from the base model, so that the inference service unit is updated to the base inference service.
[0088] The basic container of the inference service unit is called to separate the fine-tuning model from the basic model, that is, to separate the fine-tuning model from the target large model. As a result, the basic container only contains the basic model, so that the inference service unit is updated to the basic inference service, and the target inference service is unloaded from the inference service unit.
[0089] In some embodiments, the inference service unit needs to be deployed before the dynamic loading of the LoRA model described above. See [link to relevant documentation]. Figure 7 , Figure 7 This is a flowchart illustrating the third embodiment of the dynamic loading method for fine-tuning the model in this application. The method may include the following steps: S31: The controller receives the basic service resource definition; wherein, the basic service resource definition contains relevant information about the basic model and resource tags. The relevant information about the basic model is used to deploy the basic container, and the resource tags indicate whether the basic model needs to load the fine-tuning model.
[0090] At least one inference service unit can be deployed on a container cloud platform (such as Kubernetes). Each inference service unit contains auxiliary containers and a base container. One inference service unit can be deployed for each base model.
[0091] The LoRA operator receives the basic service resource definition. The basic service resource definition is used to deploy the inference service unit Pod. The basic service resource definition contains information about the Pod, information about the basic inference service interface, information about the basic model, and resource tags. The information about the basic model is used to deploy the basic container. The resource tags indicate whether the basic model needs to load the fine-tuning model. The information about the Pod is used to deploy the Pod. The information about the basic inference service interface is used to create the basic inference service interface.
[0092] For example, an inference service unit Pod can be deployed on Kubernetes through basic service resource definitions. The YAML file defining the basic service resources can contain definition fields such as the version of the Kubernetes application resources used, the Deployment resource type, and the Service resource type. Among them, the Deployment resource type is used to create a group of Pods, and the Service resource type is used to create the service interface of the Pods.
[0093] The Deployment resource type defines and manages a group of Pods, including resource tags (Deployment tags, which can be customized for Deployment attributes), Deployment name, namespace (the logical isolation unit of resources), number of Pods created and managed by the Deployment, selector for Pods associated with the Deployment, Pod tags (which can be customized for Pod attributes), and metadata definitions of the underlying containers in the Pod. Among these, resource tags include the name of the underlying model and whether dynamic loading of the LoRA model is enabled (i.e., whether the underlying model needs to load the LoRA model), etc. The metadata definitions of the underlying containers in the Pod include container image name, image pull strategy, container name, container port, container port communication protocol, etc.
[0094] Service resource types include Service tags (customizable Service attributes), Service port forwarding definitions, and tag selectors for the Pods associated with the Service. Among them, Service tags include information such as the base model name, Service name, and namespace corresponding to the Service, and Service port forwarding definitions include information such as the Service entry port, Service port communication protocol, and the port of the backend Pod.
[0095] The above configuration file defines two core resources: Deployment and Service. Deployment is used for application deployment of Pods, that is, to deploy Pods and their underlying container model-base container. Service is used to create the basic inference service interface Service model-base.
[0096] S32: The callback service interface responds to the resource tag in the basic service resource definition, which indicates that the basic model needs to load the fine-tuning model. It injects the resource definition of the auxiliary container into the basic service resource definition to obtain the modified basic service resource definition.
[0097] Continue reading Figure 2 The container cloud platform also includes callback service interfaces, which may include, but are not limited to, webhooks. The LoRA operator controller can create Pods in Kubernetes using the received basic service resource definitions. Due to the webhook mechanism, when creating a Pod, the webhook intercepts the basic service resource definitions, modifies them, and automatically injects the resource definitions of the LoRA Runtime Sidecar container. After injection, Kubernetes deploys the basic container and auxiliary containers to complete the deployment of the Pod for the basic inference service.
[0098] Specifically, in response to the resource tag in the basic service resource definition indicating that the basic model needs to load the fine-tuning model, the Webhook injects the resource definition of the auxiliary container LoRA Runtime Sidecar into the basic service resource definition to obtain the modified basic service resource definition.
[0099] S33: The controller creates an inference service unit on the container cloud platform based on the modified basic service resource definition, thereby obtaining an inference service unit with basic inference services, and / or creates a basic inference service interface for the inference service unit. The inference service unit deploys a basic container and an auxiliary container, and the basic inference service interface is used to access the basic inference service for the inference service unit.
[0100] The controller creates a Pod in Kubernetes based on the modified basic service resource definition, resulting in a Pod with basic inference services. The Pod contains basic containers and auxiliary containers.
[0101] In some implementations, the controller creates a base inference service interface (Service model-base) for the Pod, which is used to connect the inference service unit to the base inference service.
[0102] This application has the following beneficial effects: By introducing custom resources and controllers from Kubernetes for the dynamic loading of LoRA models, this application automates the management of LoRA model dynamic loading, improves operational efficiency, and enables self-recovery of LoRA model inference services, thereby enhancing inference service reliability. Based on the capabilities of the Kubernetes platform, this application manages inference services in a cloud-native manner. Through custom resources, it defines the description of LoRA model inference deployment resources. Using the Kubernetes operator framework, it automates the entire process of dynamic LoRA model loading by monitoring the deployment of LoraAdapter CR resources and combining Pod-side auxiliary containers. This improves the operational efficiency of inference services, enables self-recovery of LoRA model inference services, and enhances the reliability of multi-LoRA model inference services.
[0103] It is understood that in the above method of specific implementation, the order in which each step is written does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0104] It is understood that the dynamic loading method for the fine-tuning model in this application can be executed by a computer device, which can be any device with processing capabilities, such as a mobile device, computer, server, etc., and this application does not impose any restrictions on it. In some possible implementations, the dynamic loading method for the fine-tuning model can be implemented by the processor calling program data stored in memory.
[0105] Please see Figure 8 , Figure 8This is a schematic diagram of the structure of a computer device according to an embodiment of this application. The computer device 40 includes a memory 41 and a processor 42 coupled to each other. The memory 41 stores program data, and the processor 42 executes the program data to implement the steps performed by at least one of the controller, auxiliary container, and basic container in the dynamic loading method of the above-described fine-tuning model. In a specific implementation scenario, the computer device 40 may include, but is not limited to, a microcomputer or a server. In addition, the computer device 40 may also include mobile devices such as laptops and tablets, which are not limited here.
[0106] In this embodiment, processor 42 can also be referred to as a CPU (Central Processing Unit). Processor 42 may be an integrated circuit chip with signal processing capabilities. Processor 42 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor, or processor 42 can be any conventional processor. Furthermore, processor 42 can be implemented using integrated circuit chips.
[0107] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 50 stores program data 51 that can be executed by a processor. When the program data 51 is executed by the processor, it implements the steps performed by at least one of the controller, auxiliary container, and basic container in the dynamic loading method of the above-described fine-tuning model.
[0108] The computer-readable storage medium 50 in this embodiment includes: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program data 51. Alternatively, it can be a server storing the program data 51. The server can send the stored program data 51 to other devices for execution, or it can run the stored program data 51 itself.
[0109] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments. For the sake of brevity, this application will not repeat the details here.
[0110] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to. For the sake of brevity, the present application will not repeat them here.
[0111] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A dynamic loading method for fine-tuning a model, characterized in that, Applied to a container cloud platform, the container cloud platform comprising a controller and at least one inference service unit, each inference service unit containing a base container and auxiliary containers, the method includes: In response to adding or modifying target custom resources on the container cloud platform, the controller invokes the auxiliary container in the inference service unit to trigger the loading process of the fine-tuning model; wherein, the target custom resource contains model-related information of the fine-tuning model; The auxiliary container uses the model-related information of the fine-tuning model to download the fine-tuning model from a preset model library and stores the fine-tuning model in a shared storage volume; The base container dynamically loads the fine-tuned model from the shared storage volume and incorporates the fine-tuned model into the base model of the base container, so that the inference service unit has the target inference service. The fine-tuned model is used to fine-tune the weight parameters of the base model to achieve the target inference service.
2. The method according to claim 1, characterized in that, The model-related information includes a first identifier of the fine-tuning model, a second identifier of the base model, and address information, wherein the address information is the storage path of the fine-tuning model in a preset model library; the process of triggering the loading of the fine-tuning model by invoking the auxiliary container in the inference service unit includes: The base model of the second identifier is determined to be the target base model; The loading process of the fine-tuning model is triggered by calling the auxiliary container in the inference service unit corresponding to the target basic model; And / or, the auxiliary container uses the model-related information of the fine-tuning model to download the fine-tuning model from a preset model library, including: The auxiliary container uses the address information of the fine-tuning model to download the fine-tuning model from a preset model library.
3. The method according to claim 1, characterized in that, The controller, in response to adding or modifying target custom resources on the container cloud platform, invokes the auxiliary container in the inference service unit to trigger the loading process of the fine-tuning model, including: In response to adding or modifying target custom resources on the container cloud platform, the controller obtains several loaded models of the same basic model, wherein the loaded models are fine-tuned models loaded in the auxiliary containers of the inference service unit; Using the aforementioned loaded models, determine whether the fine-tuning model in the target custom resource meets the conditions for successful loading; In response to the fact that the fine-tuning model does not meet the conditions for successful loading, the auxiliary container in the inference service unit is invoked to trigger the loading process of the fine-tuning model.
4. The method according to claim 1, characterized in that, The step of dynamically loading the fine-tuning model from the shared storage volume by the base container and merging the fine-tuning model into the base model of the base container includes: The auxiliary container modifies the loading progress of the fine-tuning model; After modifying the loading progress of the fine-tuning model, the following is included: The controller polls the auxiliary container for the loading progress of the fine-tuning model; In response to the loading progress of the fine-tuning model being completed, the loading status of the fine-tuning model in the target custom resource is modified to be completed, and / or a target inference service interface is created for the inference service unit, wherein the target inference service interface is used to allow the inference service unit to access the target inference service.
5. The method according to claim 1, characterized in that, The target custom resource to be added or modified includes any one of the following: Create a new target custom resource; Update target custom resources; Change the loading status of the fine-tuning model in the target custom resource to an abnormal state.
6. The method according to claim 5, characterized in that, The controller responds to changes to the target custom resource on the container cloud platform, including: In response to an abnormal restart of the inference service unit, the auxiliary container obtains several target custom resources associated with the basic model in the basic container; If the loading status of the associated fine-tuning model in the base container is not loaded, the loading status of the fine-tuning model in the target custom resource is changed to an abnormal state; wherein, the associated fine-tuning model corresponds to the associated target custom resource; And / or, the fine-tuning model in an abnormal state is used as the fine-tuning model for reloading; the auxiliary container uses the model-related information of the fine-tuning model to download the fine-tuning model from a preset model library and stores the fine-tuning model on a shared storage volume, including: The auxiliary container determines whether the shared storage volume stores a reloaded fine-tuning model; In response to the absence of a reloaded fine-tuning model stored, the auxiliary container performs the steps of downloading the fine-tuning model from a preset model library using the model-related information of the fine-tuning model, and storing the fine-tuning model in a shared storage volume.
7. The method according to claim 1, characterized in that, The container cloud platform also includes a callback service interface; the method further includes: The controller receives a basic service resource definition; wherein the basic service resource definition contains relevant information about the basic model and resource tags, the relevant information about the basic model is used to deploy the basic container, and the resource tags indicate whether the basic model needs to load the fine-tuning model; The callback service interface responds to the resource tag in the basic service resource definition, which indicates that the basic model needs to load the fine-tuning model. It then injects the resource definition of the auxiliary container into the basic service resource definition to obtain the modified basic service resource definition. The controller creates an inference service unit on the container cloud platform based on the modified basic service resource definition, thereby obtaining an inference service unit with basic inference services, and / or creates a basic inference service interface for the inference service unit, wherein the inference service unit deploys a basic container and an auxiliary container, and the basic inference service interface is used to access the basic inference service for the inference service unit.
8. The method according to claim 1, characterized in that, The method further includes: In response to deleting the target custom resource on the container cloud platform, the controller invokes the auxiliary container in the inference service unit to trigger the unloading process of the fine-tuning model; wherein, the target custom resource contains model-related information of the fine-tuning model; The auxiliary container uses the model-related information of the fine-tuning model to delete the fine-tuning model from the shared storage volume; The fine-tuning model is separated from the base model by calling the base container, so that the inference service unit is updated to the base inference service; And / or, the fine-tuning model is a LoRA model, which is obtained by fine-tuning the target inference service based on the base model and the fine-tuning model.
9. A computer device, characterized in that, The method includes a memory and a processor coupled to each other, the memory storing program data, and the processor executing the program data to implement the steps performed by at least one of the controller, auxiliary container, and base container in the method of any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The system stores program data that, when executed by a processor, implements the steps performed by at least one of the controller, auxiliary container, and base container in the method described in any one of claims 1 to 8.