A method and device for incremental deployment of deep learning models

By defining a folder-based distributed storage structure and state marking mechanism in the deep learning model, the problem of deep learning models in the existing technology cannot be deployed incrementally and updated hotly is solved, efficient model updates and video memory management are achieved, and the adaptability and flexibility of the model are improved.

CN119377645BActive Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411911222.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-06-06
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing deep learning models cannot achieve incremental deployment and hot updates in actual applications, resulting in insufficient generalization capabilities of the model in a dynamically changing environment and occupies too much video memory during hot updates and version switching.

Method used

By defining a folder-based distributed storage structure, the network structure of the model is decoupled from the weight parameters, and independent files are used to store weights of each layer. Use configuration files to manage the mapping relationship between weights and operators, and ensure the atomicity of the update process through the status marking mechanism.

Benefits of technology

It realizes efficient incremental deployment and thermal updates of deep learning models, reduces video memory usage, and improves the flexibility and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377645B_ABST
    Figure CN119377645B_ABST
Patent Text Reader

Abstract

The present invention discloses an incremental deployment method and device for a deep learning model, including: defining a folder-based distributed storage structure for model storage, decoupling the model's network structure from weight parameters; reading a configuration file, establishing an operator-weight mapping relationship, loading each layer's weights according to the mapping relationship and performing forward reasoning; when a deployment end receives an update request, first downloading a new configuration file, parsing the change information therein; if it is an incremental deployment configuration file, forming a new configuration file with the original configuration file; if it is a configuration file that modifies the mapping relationship of the original configuration file, using the new configuration file; otherwise rejecting the update request; determining the operator set that needs to be updated through the new configuration file, and creating a copy for each operator to be updated; loading the new weight into the operator copy, and ensuring the atomicity of the update process through a state marking mechanism. The present invention solves the problem that deep learning models cannot be incrementally deployed and hot updated in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and in particular to a method and device for incrementally deploying a deep learning model. Background Art

[0002] With the rapid development and widespread application of deep learning technology, model deployment has become a key link in the implementation of artificial intelligence. In the current deep learning ecosystem, model training and deployment are usually regarded as two independent links. The training phase is generally carried out in a powerful data center, using deep learning frameworks such as PyTorch to continuously optimize the model weights through forward and back propagation algorithms. However, after being trained with preset data sets, these models often lack the ability to generalize in specific application scenarios, resulting in their performance not meeting expectations. In order to solve this problem, transfer learning and incremental learning have become important research directions.

[0003] For example, Chinese patent document with publication number CN113723518A discloses a task hierarchical deployment method, apparatus and computer equipment based on transfer learning; Chinese patent document with publication number CN115952873A discloses a model training and deployment method based on incremental learning and federated learning.

[0004] Incremental learning allows models to be retrained or fine-tuned with newly acquired data in practical applications, which not only improves the effectiveness of the model but also enables it to better adapt to dynamically changing environments. In this case, the weights of the model will inevitably change. Studies have found that most neural network layers are usually irrelevant to specific tasks, and only a few layers need to be fine-tuned based on new data. Therefore, if the model needs to be completely replaced for each update, it will cause excessive video memory to be occupied during hot updates and version switching, and even double the memory usage.

[0005] Some existing solutions, such as LoRA, generate a low-rank decomposition matrix (Δw) of weight increments, which can only update the weight increment matrix during the update process, making it more efficient. However, the applicability of this method is relatively limited, mainly targeting models using specific fine-tuning methods, and cannot be widely used in incremental learning scenarios where the original weights need to be updated.

[0006] Therefore, designing a system that can efficiently manage model weight updates and is not limited to specific fine-tuning methods has become the core motivation of the technical solution of the present invention. The goal of the present invention is to optimize the incremental update process of the model in a modular way, reduce the memory usage, and improve the flexibility and adaptability of the model. Summary of the invention

[0007] The present invention provides a method and device for incremental deployment of a deep learning model, which solves the problem in the prior art that deep learning models cannot be incrementally deployed and cannot be hot-updated.

[0008] A method for incremental deployment of a deep learning model, comprising the following steps:

[0009] (1) Define a folder-based distributed storage structure for model storage, decouple the model's network structure from weight parameters, and use independent files to store the weights of each layer;

[0010] (2) Read the configuration file, establish the mapping relationship between operators and weights, and then load the weights of each layer according to the mapping relationship and perform forward reasoning;

[0011] (3) When the deployment end receives an update request, it first downloads the new configuration file and parses the change information therein; if it is an incremental deployment configuration file, a new configuration file is formed with the original configuration file; if it is a configuration file that modifies the mapping relationship of the original configuration file, the modified new configuration file is used; otherwise, the update request will be rejected;

[0012] (4) Determine the set of operators that need to be updated through the new configuration file, and create a copy for each operator to be updated; load the new weights into the operator copies, and ensure the atomicity of the update process through the state marking mechanism.

[0013] In step (1), a folder-based distributed storage structure is defined for model storage, including:

[0014] The network structure is stored in a graph and saved in a file. The weight files are put into a folder and identified by name and id. The operators in the network structure specify the weights required to be calculated in the reasoning process by specifying the weight id, thereby decoupling the network structure and weight parameters of the model.

[0015] In step (2), when establishing the mapping relationship between the operator and the weight, the correctness of the mapping relationship is checked. If the mapping relationship is incorrect, an error is returned. If the mapping relationship is correct, a mapping relationship table is established and the weights required for reasoning are loaded into the memory.

[0016] In step (2), each operator has one or more weight files for performing forward reasoning. The weight files specifically used by each operator are loaded according to the mapping relationship and forward reasoning is performed.

[0017] In step (3), the configuration file monitoring module detects the update status of the configuration file in real time to ensure the timeliness of the update trigger.

[0018] In step (4), the atomicity of the update process is ensured through the state marking mechanism, specifically:

[0019] Set five states for the operator: initial state, updating, updating completed, in use, and waiting to be recycled;

[0020] When new weights are loaded into an operator copy, the copy is marked as being updated; after loading is complete, it is marked as updated; the new operator is used in the next inference and marked as in use; at the same time, the original operator is marked as pending recycling; the state marking mechanism is used to ensure that the inference weight update is completed without interrupting the inference, and finally the resources occupied by the operator in the pending recycling state are recycled.

[0021] An incremental deployment device for a deep learning model, comprising:

[0022] The file management module is used to define a folder-based distributed storage structure for model storage, decouple the model's network structure from weight parameters, use independent files to store the weights of each layer, and manage the mapping relationship between weights and network structure through configuration files;

[0023] The configuration file update module is used to detect the update status of the configuration file in real time, and parse the change information in the newly downloaded configuration file; if it is an incremental deployment configuration file, a new configuration file is formed with the original configuration file; if it is a configuration file that modifies the mapping relationship of the original configuration file, the new configuration file is used; otherwise, the update request will be rejected;

[0024] The hot update module is used to update the inference weights by reading the configuration file without interrupting the inference. First, the set of operators that need to be updated is determined through the new configuration file, and a copy is created for each operator to be updated. Next, the new weights are loaded into the operator copies, and the atomicity of the update process is ensured through the state marking mechanism.

[0025] A device for incremental deployment of a deep learning model comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the incremental deployment method of the deep learning model.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] The present invention adopts a deep learning model deployment strategy that can be incrementally deployed and hot-updated, which effectively reduces the difficulty of operation while ensuring the efficiency of reasoning services as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flow chart of an incremental deployment method for a deep learning model according to an embodiment of the present invention.

[0029] Figure 2A schematic diagram of a specific strategy selection process for an incremental deployment method of a deep learning model according to an embodiment of the present invention.

[0030] Figure 3 A schematic diagram of incremental deployment of a deep learning model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.

[0032] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0033] (1) Model deployment refers to the process of transferring a trained machine learning or deep learning model from a development environment to a production environment. In this process, the model is integrated into an application or service so that it can process real-time data and generate predictions or decisions. Model deployment typically involves selecting a suitable deployment platform, optimizing model performance, ensuring the stability and reliability of the model in a production environment, and monitoring the performance of the model.

[0034] (2) Model inference refers to the process of making predictions or decisions based on a trained machine learning or deep learning model. In this process, the model receives input data (usually in the form of tensors) and calculates output results through forward propagation. Model inference is the core function after the model is deployed, which determines the performance and effect of the model in practical applications. The inference process usually needs to be optimized to improve efficiency and responsiveness, especially when processing large-scale data or real-time data.

[0035] (3) Operator refers to a function or process that performs a specific mathematical operation or operation. An operator can be a simple mathematical operation (such as addition or multiplication) or a complex neural network layer (such as a convolutional layer or a fully connected layer). In a computational graph, operators are usually represented as edges that connect different nodes (variables) and define the computational relationship between these nodes. The selection and combination of operators determines the structure and function of the model.

[0036] (4) Weight refers to a type of model parameter that controls the transmission and transformation of input data in the model. In a neural network, weights are usually associated with the edges connecting neurons and represent the signal strength from one neuron to another. Weights are adjusted during model training through optimization algorithms (such as gradient descent) to minimize the loss function and improve the model's predictive ability. The value of the weight determines the model's learning ability and generalization ability.

[0037] like Figure 1 As shown, a method for incremental deployment of a deep learning model includes the following steps:

[0038] Step S101, define a model storage format based on a folder-based distributed storage structure, decouple the model's network structure from weight parameters, use independent files to store the weights of each layer, and manage the mapping relationship between weights and operators through configuration files.

[0039] In this embodiment, the network structure is stored in a graph and stored in a file. The weight files are uniformly placed in a folder and identified by name and id. The operators in the network structure can specify the weights required to participate in the calculation during the reasoning process by specifying the weight id, thereby realizing the decoupling of the network structure and weight parameters of the model. The mapping relationship between weights and operators is managed through configuration files.

[0040] Step S102, initializing loading, reading a configuration file, establishing a mapping relationship between operators and weights, and loading weights according to the mapping relationship.

[0041] In this embodiment, after the initialization loading is completed, the configuration file is read, the mapping relationship between the operator and the weight is established and the correctness of the mapping relationship is verified. If the mapping relationship is incorrect, an error is returned. If the mapping relationship is reasonable, a mapping relationship table is established and the weights required for reasoning are loaded into the memory.

[0042] Step S103, performing forward reasoning according to the existing mapping relationship and loading weights.

[0043] In this embodiment, the weights currently loaded by each operator are used for forward reasoning, so that there is no need to look up the mapping relationship table during the reasoning process.

[0044] Step S104, detecting and reading a new configuration file, and updating the pre-imported configuration file according to the type of the new configuration file.

[0045] In this embodiment, the configuration file update module detects whether there is a new configuration file to be imported. If there is no new configuration file to be imported, the reasoning is continued without interruption. If there is a new configuration file to be imported, the new configuration file type is checked and a pre-imported configuration file is formed. Check whether the new configuration file is an incremental configuration file. If it is an incremental configuration file, the incremental configuration file and the original configuration file are combined to form a pre-imported configuration file. If it is a modification of the mapping relationship of the original configuration file, the new configuration file is used as the pre-imported configuration file. If it does not meet the above two conditions, it means that the configuration file format is incorrect and cannot be correctly parsed, and the update request is rejected.

[0046] In some possible implementations, the incremental configuration file only contains the weight file and new mapping relationship information that need to be added by the incrementally modified operator, and the new configuration file modified by the original configuration file only contains the mapping relationship information that needs to be updated by the modified mapping operator.

[0047] Step S105 , performing hot update during inference according to the newly imported configuration file.

[0048] In this embodiment, five states are set for the operator: "initial state", "updating", "updated", "in use" and "to be recycled". When the new weight is loaded into the operator copy, the copy is marked as "updating"; after the loading is completed, it is marked as "updated"; the new operator is used in the next reasoning and marked as "in use"; at the same time, the original operator is marked as "to be recycled". The use of state marking can ensure that the reasoning weight update is completed without interrupting the reasoning. Finally, the resources occupied by the operator in the "to be recycled" state are recycled.

[0049] like Figure 2 As shown, the specific strategy selection of the deep learning model incremental deployment method in the embodiment of the present invention may include the following steps:

[0050] Step S201, establish a mapping relationship table between operators and weights according to the configuration file and load the weights of each layer according to the mapping relationship table.

[0051] In this embodiment, it is necessary to set an id and a flag bit for the weight of each operator respectively, so as to facilitate the determination of the state information related to the weight imported by each operator.

[0052] Step S202, importing data for forward reasoning based on the existing mapping relationship and loading weights.

[0053] In this embodiment, data is used according to the calculation graph to determine the operators used in sequence in the reasoning process. The operators determine the weights imported to participate in the calculation according to the mapping relationship table. The weights participating in the calculation need to be loaded into the memory, while the weights that do not need to participate in the calculation do not need to occupy memory.

[0054] Step S203, entering different stages according to whether a new configuration file is detected to be imported.

[0055] In this embodiment, if no new configuration file is detected to be imported, step S202 is still executed, that is, reasoning is continued based on the existing configuration file. If a new configuration file is detected to be imported, step S204 is entered.

[0056] Step S204, determining whether the newly imported configuration file is an incremental configuration file, if not, proceeding to step S205, continuing to determine whether the new configuration file is a mapping modification of the original configuration file; if it is an incremental configuration file, proceeding to step S206.

[0057] Step S205, determining whether the newly imported configuration file is a modification of the mapping relationship of the original configuration file. If it is not a modification of the mapping relationship of the original configuration file, reject the new configuration file and still execute step S202 according to the old configuration file; if it is a modification of the mapping relationship of the original configuration file, proceed to step S207.

[0058] Step S206: forming a new round of configuration files according to the newly imported incremental configuration files and the original configuration files.

[0059] In this embodiment, the newly imported incremental configuration file includes the incremental modification information of the configuration file and the weight corresponding to the incremental modification information. The original configuration file and the newly imported incremental configuration file will form a new round of configuration files and update the mapping relationship table accordingly.

[0060] Step S207: Use the new configuration file as a new round of configuration files.

[0061] Step S208, read a new round of configuration files, determine the set of operators to be updated, create copies for the operators to be updated, and load the new weights into the operator copies.

[0062] In this embodiment, the configuration file specifies the id of the new weight file required for the operator to be updated, and the weight file is determined in the weight file directory according to the id of the weight file and loaded into the operator copy.

[0063] Step S209: using a state marking mechanism to ensure the atomicity of the inference weight update process.

[0064] In this embodiment, five states are set for the operator: "initial state", "updating", "updated", "in use" and "to be recycled". When the new weight is loaded into the operator copy, the copy is marked as "updating"; after the loading is completed, it is marked as "updated"; the new operator is used in the next reasoning and marked as "in use"; at the same time, the original operator is marked as "to be recycled". The use of state marking can ensure that the reasoning weight update is completed without interrupting the reasoning.

[0065] Step S210: reclaiming resources occupied by operators in the "pending reclaim" state.

[0066] like Figure 3As shown, an incremental deployment of a deep learning model, the device includes: a file management module 301, a configuration file update module 302 and a hot update module 303, wherein:

[0067] The file management module 301 is used to define and parse the storage format of the deep learning model, including the mapping method between operators and weights, the storage method of weight files, and the formats of configuration files and incremental configuration files;

[0068] The configuration file updating module 302 is used to detect and read the newly imported configuration file and form a new round of configuration files to be imported according to the newly imported configuration file and the original configuration file;

[0069] The hot update module 303 is used to replace the inference weights of the deep learning model according to a new round of configuration files to be imported, and ensure the atomicity of operator updates and the uninterrupted nature of the inference process.

[0070] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for incremental deployment of a deep learning model, characterized in that: The following steps are involved: (1) Define a folder-based distributed storage structure for model storage, decouple the model's network structure from weight parameters, and use independent files to store the weights of each layer; Define a folder-based distributed storage structure for model storage, including: storing the network structure in a graph and storing it in a file, putting the weight files into a folder and identifying them with names and ids, and the operators in the network structure specify the weights that need to be calculated in the reasoning process by specifying the weight ids, thereby realizing the decoupling of the model's network structure and weight parameters; (2) Read the configuration file, establish the mapping relationship between operators and weights, and then load the weights of each layer according to the mapping relationship and perform forward reasoning; (3) When the deployment end receives an update request, it first downloads the new configuration file and parses the change information therein; if it is an incremental deployment configuration file, a new configuration file is formed with the original configuration file; if it is a configuration file that modifies the mapping relationship of the original configuration file, the modified new configuration file is used; otherwise, the update request will be rejected; (4) Determine the set of operators that need to be updated through the new configuration file and create a copy for each operator to be updated; load the new weights into the operator copies and ensure the atomicity of the update process through the state marking mechanism; The atomicity of the update process is ensured through the state marking mechanism, specifically: Set five states for the operator: initial state, updating, updating completed, in use, and waiting to be recycled; When new weights are loaded into an operator copy, the copy is marked as being updated; after loading is complete, it is marked as updated; the new operator is used in the next inference and marked as in use; at the same time, the original operator is marked as pending recycling; the state marking mechanism is used to ensure that the inference weight update is completed without interrupting the inference, and finally the resources occupied by the operator in the pending recycling state are recycled.

2. The incremental deployment method of the deep learning model according to claim 1, characterized in that: In step (2), when establishing the mapping relationship between the operator and the weight, the correctness of the mapping relationship is checked. If the mapping relationship is incorrect, an error is returned. If the mapping relationship is correct, a mapping relationship table is established and the weights required for reasoning are loaded into the memory.

3. The incremental deployment method of a deep learning model according to claim 1, characterized in that: In step (2), each operator has one or more weight files for performing forward reasoning. The weight files specifically used by each operator are loaded according to the mapping relationship and forward reasoning is performed.

4. The incremental deployment method of a deep learning model according to claim 1, characterized in that: In step (3), the configuration file monitoring module detects the update status of the configuration file in real time to ensure the timeliness of the update trigger.

5. An incremental deployment device for a deep learning model, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the incremental deployment method of the deep learning model described in any one of claims 1-4.

Citation Information

Patent Citations

  • Task hierarchical deployment method and device based on transfer learning and computer equipment

    CN113723518A

  • Model training and deployment method based on incremental learning and federated learning

    CN115952873A

  • Neural network structured sparse method based on incremental regularization

    CN110197257A

  • Super-large-scale distributed machine learning framework, method and device and storage medium

    CN114510351A