Model weight updating method, device, storage medium and program product
Patent Information
- Application Number
- PCT/IB2025/050049
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-17
Smart Images

Figure IB2025050049_17072025_PF_FP_ABST
Abstract
Description
[0001] TECHNICAL FIELD The present disclosure relates to the field of artificial intelligence technology, and more particularly to a model weight updating method, device, storage medium, and program product. BACKGROUND With the continuous development and improvement of computing power, data, and algorithms in human society, the training and deployment of deep learning models have become a hot research topic in the field of artificial intelligence. However, the massive consumption of computing resources and high deployment costs have become major obstacles to the development of deep learning models. To this end, cloud computing vendors have gradually introduced acceleration engines or frameworks for model inference, such as TensorRT. The TensorRT acceleration engine improves the model's operating efficiency on hardware resources such as graphics processing units (GPUs) by compiling and optimizing the model. However, the compilation and optimization time also seriously affects the efficiency of model deployment. In particular, the larger the model, the longer the compilation and optimization time, which can even reach several hours. This has led to the pain point of "one hour to compile, ten seconds to infer." Therefore, how to improve model deployment efficiency while achieving the performance benefits brought by compilation and optimization has become an urgent problem in the industry. SUMMARY OF THE INVENTION Various aspects of the present disclosure provide a model weight update method, device, storage medium, and program product for improving model deployment efficiency while achieving performance benefits from compilation optimization. Embodiments of the present disclosure provide a model weight update method, comprising: pre-compilation optimization of a first original model to obtain a target model without weight information, and saving the types and names of weights required by the first original model as weight description information adapted for the target model; when a second original model is isomorphic to the first original model is activated to perform a deep learning task, obtaining second weight information corresponding to the second original model, where the first original model has the first weight information; and injecting the second weight information into the target model based on the weight description information, so as to perform the deep learning task using the target model with the injected second weight information. Embodiments of the present disclosure also provide an electronic device, comprising: a memory and a processor; the memory storing a computer program, the processor coupled to the memory for executing the computer program to implement the steps of the aforementioned method. Embodiments of the present disclosure also provide a computer-readable storage medium storing the computer program, which, when executed by the processor, causes the processor to implement the steps of the aforementioned method. An embodiment of the present disclosure further provides a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the processor is enabled to implement the steps in the above method.In an embodiment of the present disclosure, a first original model is pre-compiled and optimized to obtain a target model without weight information. The types and names of the weights required by the first original model are stored as weight description information adapted for the target model. When a second original model is activated to perform a deep learning task, second weight information corresponding to the second original model is obtained. Based on the weight description information, the second weight information is injected into the target model, allowing the target model with the injected second weight information to perform the deep learning task. For isomorphic models, only one compilation optimization process is required, achieving the goal of eliminating compilation when switching between isomorphic models and saving compilation optimization time. Furthermore, while achieving the performance benefits of compilation optimization, injecting the second weight information into the target model to enable switching between isomorphic models can also improve model deployment efficiency. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are provided to provide a further understanding of the present disclosure and constitute a part of this disclosure. The illustrative embodiments of this disclosure and their descriptions are provided for illustrative purposes and are not intended to unduly limit this disclosure. In the accompanying drawings: Figure 1a is a schematic diagram of a model training and model inference process; Figure 1b is a schematic diagram of a model inference process provided by an exemplary embodiment of the present disclosure; Figure 1c is a schematic diagram of the internal structure of an inference acceleration engine provided by an exemplary embodiment of the present disclosure; Figure 2 is a schematic diagram of a model weight update method provided by an exemplary embodiment of the present disclosure; Figure 3a is a schematic diagram of a weight registration process based on a composite hash table provided by an exemplary embodiment of the present disclosure; Figure 3b is a schematic diagram of a hash table-based weight information acquisition and loading process provided by an exemplary embodiment of the present disclosure; Figure 4 is a schematic diagram of the structure of a model weight update device provided by an exemplary embodiment of the present disclosure; Figure 5 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION To further clarify the objectives, technical solutions, and advantages of the present disclosure, the technical solutions of the present disclosure will be described clearly and completely below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.In addition, the various models covered by this disclosure (including but not limited to language models and large models) comply with relevant laws and standards. A deep learning model is a machine learning model that uses a multi-layer neural network for training and prediction. The core of a deep learning model is a neural network, which consists of multiple neurons (or nodes). Each neuron is connected to the neurons in the previous layer. By adjusting the connection weights between neurons, the neural network can learn to extract useful information from input data during training and automatically learn to extract useful features from the input data during inference, which can then be used for tasks such as classification and regression. In the embodiments of the present disclosure, the type of deep learning model is not limited, and examples include but are not limited to: Restricted Boltzmann Machine (RBM), Autoencoder, Deep Belief Network (DBN), Deep Boltzmann Machine (DBM), Product Quantization Network (PQNet), Deep Perceptron (DP), Deep Feedforward Network (DFN), Convolutional Neural Networks (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Generative Adversarial Network (GAN), Auto-encoder (AE), Residual Neural Network (RNN), etc. ResNet) and attention mechanism (Attention.
[0002] This can include existing deep learning models such as the Deep Learning Mechanism (AM), or various deep learning models that may emerge in the future. Regardless of the type of deep learning model, its lifecycle consists of a training phase and an inference phase. The training phase involves using training samples to train the initial deep learning model, allowing it to grow into a model capable of solving a specific problem. The inference phase involves using the deep learning model obtained during the training phase to solve a specific problem. To improve model training efficiency, as shown in Figure 1a, a training acceleration engine can be used to train the model during the training phase. A training acceleration engine is a deep learning framework for model training, such as TensorFlow or PyTorch. By utilizing parallel computing and distributed computing technologies, as well as optimizing the hardware resources used for model training, it accelerates the training process of deep learning models and reduces training time. The hardware resources used for model training can include graphics processing units (GPUs) and tensor processing units (TPUs). As shown in Figure 1a, during the training phase, the model training process of the training acceleration engine involves at least model construction, model compilation, model optimization, and model serialization. Model construction involves the following: 1. Clarifying the research objective. Before model construction, the research objective and requirements must be determined. Clarifying the research objective helps select appropriate mathematical models and parameters. 2. Defining variables. Before converting the research object into a mathematical model, it must be abstracted into variables. These variables can be quantities, states, or features. Defining variables clarifies the characteristics and attributes of the research object. 3. Selecting a mathematical model. Based on the research objective and defined variables, an appropriate mathematical model is selected. Mathematical models can be linear, nonlinear, or probabilistic. The selection of a mathematical model requires comprehensive consideration of the research objective, variable characteristics, and data type. Model compilation involves selecting an appropriate loss function, optimizer, and evaluation metric for the model before training, and binding them to the model. Model compilation is a static process that only needs to be performed once before training begins. Model compilation involves the following: 1. Defining model parameters. Before compiling the model, you need to define the required model parameters, such as the learning rate and batch size. These parameters will affect the model's training process and final performance. 2. Compile the model.Use the Application Programming Interface (API) provided by a training acceleration engine (such as TensorFlow or PyTorch) to compile the defined model structure and parameters. Model optimization refers to improving model performance, accuracy, generalization, and training efficiency by adjusting model parameters and hyperparameters during training. Model optimization is a dynamic process that requires continuous adjustment of model parameters and hyperparameters to achieve optimal performance. Model optimization involves the following: 1. Setting the optimizer. The optimizer updates the model weights during training. Common optimizers include Stochastic Gradient Descent (SGD) and Adaptive Moment Estimation (Adam). You can generally choose an appropriate optimizer based on the model requirements and dataset characteristics. 2. Setting the loss function. The loss function measures the difference between the model's predictions and the true values. Depending on the specific task type (classification, regression, etc.), choose an appropriate loss function, such as cross-bench loss or mean squared error. 3. Setting the evaluation metric. In addition to the loss function, appropriate evaluation metrics are also required to assess model performance. For example, for classification tasks, metrics such as accuracy, precision, and recall can be used; for regression tasks, metrics such as mean squared error and root mean squared error can be used. 4. Perform model optimization. Optionally, model optimization involves, but is not limited to, the following: 1. Feature selection. Selecting features most relevant to the target variable can reduce the number of features and improve the model's generalization ability. 2. Model tuning. Model performance can be improved by adjusting model parameters or structure, such as adding or reducing the number of layers or changing the activation function. 3. Regularization. Adding regularization terms to the loss function constrains model complexity and prevents overfitting. 4. Data augmentation. Continuously training the model by generating new training data to improve its generalization ability. 5. Early stopping. Stopping training when performance on the validation set no longer improves can prevent overfitting. 6. Learning rate adjustment. Using techniques such as learning rate decay can improve convergence during training. 7. Multi-task learning. Allowing a model to solve multiple related tasks simultaneously can improve its generalization capabilities. Model serialization converts the model object running in memory into a binary sequence file. This binary sequence file is then stored on a persistent storage medium (such as a hard drive) to facilitate the flexible use of deep learning models.In various embodiments of the present disclosure, a binary sequence file can also be referred to as a serialized file of a deep learning model. The serialized file of a deep learning model includes information such as the model structure, the names of each network layer, and the names and values of the weights in each network layer. To improve the efficiency of model inference, as shown in FIG1a , an inference acceleration engine is used to perform model inference during the inference phase. Similar to a training acceleration engine, an inference acceleration engine is a deep learning framework for model inference, such as Torch + Xformers or TensorRT. The inference acceleration engine uses compilation optimization, optimization of hardware resources used for model inference, and model compression to accelerate the inference process of deep learning models and reduce model inference time. The inference acceleration engine and the training acceleration engine can be the same deep learning framework or different deep learning frameworks. For example, Torch is a deep learning training and inference framework that supports hardware resources such as CPUs and GPUs. Xformers is an acceleration library for the Multi-Head Attention (MHA) algorithm, which can be used for both model training and model inference. TensorRT is a tensor-oriented runtime acceleration engine primarily used for model inference. As shown in Figure 1a, the model inference process of the inference acceleration engine involves at least model deserialization, computational graph conversion, compilation optimization, and task execution. Model deserialization refers to the process of deserializing the binary sequence file of a deep learning model from persistent storage media (such as a hard drive) into memory to obtain a runnable model object. This allows the deep learning model to be used for self-deployed deep learning tasks. The model deserialization process can also be referred to as model loading. Computational graph conversion is the process of parsing the deep learning model into a computational graph. A computational graph is a data structure used to represent the computational process of a deep learning model. A computational graph consists of nodes and edges. Nodes represent the network layers in a deep learning model, each layer performing operations such as matrix multiplication and addition. Edges represent the data flow between the network layers represented by adjacent nodes. By traversing the structure of a deep learning model, the computations in each network layer are converted into nodes in the computation graph. Each node's inputs, outputs, and connections with other nodes are recorded. Compilation optimization refers to the process of optimizing the computation graph after conversion. These optimizations include node fusion, pruning, vectorization, memory optimization, and parallelization. This reduces computational complexity and memory usage, improving operational efficiency.Node fusion refers to merging multiple nodes in a computation graph into a single node, thereby reducing computational and communication overhead. For example, multiple convolution operations can be merged into a single convolution operation. Node fusion is also called operator fusion or graph fusion. Pruning refers to removing unnecessary nodes and edges in a computation graph to reduce computational complexity; the removed nodes and edges do not affect the model output. Vectorization converts loops and iterations in a computation graph into vectorized operations. Vectorized operations can leverage the parallel processing capabilities of hardware resources to improve computational efficiency. Memory optimization improves model execution efficiency by optimizing memory access patterns and reducing memory allocations. For example, cache optimization and memory alignment can be used to reduce memory access latency. Parallelization parallelizes operations in the computation graph to fully utilize the computing power of hardware resources used for model inference, such as multi-core CPUs or GPUs, thereby increasing model execution speed. The compilation optimization techniques listed above can be selected and applied based on specific needs and goals to achieve the desired performance optimization. Task execution refers to the process of using the compiled and optimized model to perform inference calculations for the corresponding task and output the calculation results. With the development of model applications, some application scenarios have emerged that require frequent switching between different deep learning models. For example, there are multiple tasks, such as Task A, Task B, and Task C, and different tasks require different deep learning models. "Different deep learning models" here include both different model structures and different weight information. When switching from Task A to Task B, the deep learning model used to execute Task A must be synchronously switched to the deep learning model used to execute Task B. In application scenarios involving frequent model switching, if the inference process shown in Figure 1a is followed, each deep learning model switch requires re-processing such as deserialization, computational graph conversion, and compilation optimization. Since optimization compilation is a time-consuming process, it can seriously affect model deployment or switching efficiency. Especially in scenarios with frequent switching and large model sizes, the extreme phenomenon of "one hour for compilation and ten seconds for inference" is prone to occur. In response to the above issues, the inventors of this case, after extensive research and analysis, have redefined deep learning models in application scenarios and proposed the concept of "isomorphic models." Isomorphic models are a general term for multiple models with the same structure but different weight information. Isomorphic models include multiple deep learning models with the same structure but different weight information. For these deep learning models with the same structure but different weight information, different training samples can be used to train them during the training process, thereby obtaining isomorphic models for performing different deep learning tasks.In other words, different models within a homogeneous model utilize different training samples during training. Using these different training samples results in different weight information, which is used to adapt to different deep learning tasks. For example, deep learning models A and B have the same model structure and are both used for object recognition. Deep learning model A is specifically used for part recognition, while deep learning model B is specifically used for animal recognition. The weight information of these two deep learning models is different, and thus deep learning models A and B are homogeneous models. Alternatively, deep learning models C and D have the same model structure and are both used for generating home decoration plans. Deep learning model C is specifically used for generating Nordic-style home decoration plans, while deep learning model D is specifically used for generating Classical-style home decoration plans. The weight information of these two deep learning models is different, and thus deep learning models C and D are homogeneous models. Of course, the model structures and model functions of different homogeneous models vary, and the present embodiments do not limit these. For ease of description and distinction, in the present embodiments, the deep learning models within a homogeneous model with the same model structure but different weight information are referred to as original models. Based on the definition of isomorphic models, embodiments of the present disclosure also provide a model weight update method for implementing model switching or deployment. Specifically, for isomorphic models, a compilation optimization process is performed on any of the original models to obtain a compiled and optimized model structure. Subsequently, when a particular original model is needed for model inference, the compilation optimization process is skipped and the weight information of the original model is injected into the compiled and optimized model structure according to task requirements to perform a weight update. This results in a compiled and optimized model capable of executing the corresponding task, thereby implementing model deployment or switching. In other words, for isomorphic models, only one compilation optimization process is required, achieving the goal of eliminating compilation during isomorphic model switching and saving compilation optimization time. Furthermore, when another original model in the isomorphic model is enabled to perform a deep learning task, the weight information of the other original model can be directly injected into the target model, allowing the target model to execute the deep learning task based on the injected weight information. While achieving the performance benefits of compilation optimization, the injection of weight information into the target model enables isomorphic model switching or deployment, thereby improving model deployment efficiency. As shown in FIG1b , the model inference stage provided by this embodiment includes: step b1, a process of pre-compiling and optimizing the first original model to obtain a target model; and step b2, a process of injecting the weight information of the second original model into the target model according to the requirements of the deep learning task to implement model switching or deployment.The first and second original models are isomorphic models. The model structure and model functions of the isomorphic models are not limited. Furthermore, the training process for each original model in the isomorphic model is also not limited. The model training methods described in the previous embodiments may be employed, but are not limited to those described. In subsequent embodiments of the present disclosure, assuming that each original model in the isomorphic model has been trained, the focus will be on how to quickly switch or deploy between isomorphic models based on the requirements of deep learning tasks. The first original model has first weight information, and the second original model has second weight information. The first original model can be any original model in the isomorphic model to which it belongs. Furthermore, the second original model can be the same as the first original model or a different model in the isomorphic model. This is not limited and depends on the deep learning task to be performed. When the first and second original models are the same, the first and second weight information are the same. As shown in FIG1b , the implementation process of compiling and optimizing the first original model to obtain a target model is as follows: Step b11: Deserializing the serialized file of the first original model. The deserialization process can obtain the first original model and load the first original model into memory. Step b12: Performing computational graph conversion on the first original model loaded into memory to obtain the computational graph of the first original model. Step b13: Compiling and optimizing the computational graph to obtain a compiled and optimized model. The compiled and optimized model includes model structure information of the isomorphic model to which the first original model belongs and first weight information corresponding to the first original model. Step b14: Serializing the model structure information obtained through compilation and optimization to obtain a target model without weight information. The serialized file of the first original model is a file suitable for persistent storage, such as, but not limited to, a binary file. The file includes: model structure information of the first original model, the names of each network layer included in the first original model, the types and names of weights in each network layer, and the corresponding first weight information. All original models in a homogeneous model have the same model structure information. Specifically, the model structure information of a first original model is identical to that of other original models in the homogeneous model, and can also be referred to as the model structure information of the homogeneous model to which it belongs. The purpose of serializing the compiled model obtained through the aforementioned compilation and optimization is to convert the model structure information into a serialized file for persistent storage. This allows subsequent use without further compilation and optimization. Instead, the model structure information can be directly deserialized from the persistently stored serialized file, improving model inference efficiency.In this embodiment, the serialized file of the target model is also a file suitable for persistent storage. For example, it can be, but is not limited to, a binary file. This file includes the aforementioned model structure information, but does not include the first weight information. The model structure information primarily includes the number of network layers contained in the first original model or its isomorphic model, the types of compute nodes involved in these network layers, and the connection relationships between these network layers or compute nodes. The "weight information" referred to in various embodiments of this disclosure primarily refers to weight values, not weight types and names. Furthermore, although only the model structure information is serialized and persistently stored, the physical computing resource object executing the model weight update method also stores a compiled model in its memory. This compiled model contains both the model structure information and the first weight information. Therefore, model inference can also be performed using the compiled model. Model inference based on the compiled model and serializing the compiled model are two independent operations, with no ordering restriction. For ease of distinction and description, the physical computing resource object used to execute the deep learning model on the electronic device is referred to as the first physical computing resource object, and the physical computing resource object responsible for executing the model weight update method is referred to as the second physical computing resource object. Physical computing resource objects on an electronic device include CPUs, GPUs, and data processing units (DPUs). The first and second physical computing resource objects can be the same, for example, both CPUs or both GPUs. Alternatively, the first and second physical computing resource objects can be different, for example, the first physical computing resource object can be, but is not limited to, a GPU, and the second physical computing resource object can be a CPU. In embodiments of the present disclosure, to facilitate the successful injection of second weight information into a target model, during the process of compiling and optimizing a first original model to obtain a target model, the types and names of weights required by the first original model are saved as weight description information required by the target model. This weight description information is used to describe the types and names of weights required by the target model. Based on this information, weight information from various original models can be injected into the target model, enabling weight information transfer. This weight description information can be serialized into the target model's serialized file as part of the target model. In the embodiment of the present disclosure, the serialized file of the target model includes not only the aforementioned model structure information, but also the weight description information herein. The weight description information herein mainly refers to the type and name of the weights in each network layer in the first original model or its isomorphic model, but does not include weight information (i.e., weight value).In an optional embodiment, during the computation graph conversion process, the weight types and names required by the first original model can be saved as weight description information required by the target model. Specifically, the first original model includes multiple network layers, each of which has its own weight type and name. During the computation graph conversion process for the first original model, the computation graph conversion process can be performed on each network layer one by one. During the computation graph conversion process, the weight type and name of each network layer can be saved as the weight description information of the target model. One implementation process for converting the computation graph of the first original model to obtain the computation graph corresponding to the first original model is as follows: parsing the first original model to obtain each network layer included in the first original model, operators involved in each network layer, inputs and outputs of the operators in each network layer, weight types and names of each network layer, and weight information corresponding to each weight name (i.e., first weight information), etc., wherein operators include at least addition, subtraction, multiplication, division, convolution, pooling, normalization, etc., and generally, one network layer corresponds to one node; further, mapping the ordered connection relationships between the network layers into directed edges between the nodes corresponding to the network layers, and mapping the inputs and outputs between the network layers into inputs and outputs between the nodes corresponding to the network layers; further, obtaining the computation graph of the first original model based on the nodes, directed edges, and input and output relationships between the nodes. In conjunction with the computational graph conversion process described above, the target network layer being converted can be converted into a target node in the computational graph and the weight types and names required by the target network layer can be saved. The target network layer can be any network layer in the first original model. Furthermore, the method for saving the weight types and names required by the target network layer varies depending on how the target network layer is constructed. If the target network layer is constructed using the plugin method, the weight types and names of the target network layer are saved to the target node via parameter passing. If the target network layer is constructed using the built-in method, a first data structure is created, and the identification information of the target network layer and the weight types and names of the target network layer are stored in the first data structure. The built-in method and the plugin method are two different ways to construct network layers in a deep learning model. The built-in method can be understood as coupling the construction methods of the network layers. The coupling relationship between the network layers is determined during construction and cannot be easily modified.The plug-in approach can be understood as a decoupled construction method between network layers. Relationships between network layers can be established through plug-in and unplugging. Plugin network layers can be custom network layers and include optimization operators. Accordingly, the network layers in a deep learning model can be divided into built-in network layers and plug-in network layers. A model can include either type of network layer or both. Plugin network layers have lower implementation costs than built-in network layers. In this embodiment, the first original model may include a built-in network layer by default. Furthermore, whether the first original model includes a plug-in network layer depends on specific needs. In the disclosed embodiments, the implementation of the first data structure is not limited and may be an array, a list, a hash table, or other forms. In an optional embodiment, the first data structure is implemented as a composite hash table. A detailed description of this implementation is provided in the subsequent embodiments. As shown in FIG1b , the process of injecting the weight information of the second original model into the target model to achieve model switching or deployment according to the requirements of the deep learning task is as follows: Step b21: When the second original model is enabled to perform the deep learning task, second weight information corresponding to the second original model is obtained; Step b22: Injecting the second weight information into the target model. Specifically, after executing step b14, the second weight information is injected into the target model through model deserialization or model resident memory; Step b23: Performing the deep learning task using the target model injected with the second weight information. Further optionally, as shown in FIG1b , obtaining the second weight information corresponding to the second original model includes: Step b211: When loading a second original model isomorphic to the first original model, deserializing the serialized file of the second original model, deserializing to obtain the second original model, and loading the second original model into memory; Step b212: Obtaining a weight dictionary for the second original model, wherein the weight dictionary contains the second weight information of the second original model. The serialized file of the second original model is a file suitable for persistent storage, such as, but not limited to, a binary file. The file includes: model structure information of the second original model, the names of each network layer included in the second original model, the type and name of the weights in each network layer, and the corresponding second weight information. Regarding the information type contained in the serialized files, the serialized files of the first and second original models contain the same information type. However, the serialized files of the target model contain different information types from those of the first and second original models. The primary difference is that the serialized files of the target model do not contain specific weight information.Further optionally, the implementation of injecting the second weight information into the target model includes: injecting the second weight information into the target model based on weight description information corresponding to the target model. Further optionally, preprocessing the second weight information to obtain third weight information, and relocating the third weight information from the second memory to the first memory; and reading the third weight information from the first memory based on the weight description information and injecting it into the target model. The second target weight information is stored in the memory of the second physical computing resource object, whereas the first physical computing resource object is responsible for running the model. To facilitate injection of the second weight information into the target model, the second weight information needs to be relocated from the memory of the second physical computing resource object to the memory of the first physical computing resource object. In the disclosed embodiments, the memory of the first physical computing resource object may be referred to as the first memory, and the memory of the second physical computing resource object may be referred to as the second memory. Furthermore, before migrating the second weight information from the second memory to the first memory, the second weight information may be preprocessed, such as by merging or renaming the weight information. The preprocessed weight information is then referred to as third weight information. The third weight information is then moved from the second memory to the first memory, and the third weight information in the second memory is injected into the target model based on the weight description information. In some embodiments, reading the third weight information from the first memory and injecting it into the target model based on the weight description information includes: constructing a second data structure compatible with the target model, the second data structure being used to store weight names and the storage addresses of their corresponding weight information; adding the storage addresses of the third weight information in the first memory to the second data structure based on the weight names; and reading the third weight information from the first memory and injecting it into the target model based on the weight names stored in the second data structure and the storage addresses of their corresponding weight information in the first memory. In the embodiments of the present disclosure, the implementation of the second data structure is not limited; for example, it may be, but is not limited to, an array, a list, or a hash table. In an optional embodiment, the second data structure is implemented as a hash table including weight names and weight information storage addresses. For details, see the description of subsequent embodiments. In some embodiments, the target model is persistently resident in the memory of the second physical computing resource object, referred to as resident memory. When the target model is persistently resident in the memory of the second physical computing resource object, the third weight information can be read from the first memory and injected into the target model based on the weight names and corresponding weight information stored in the second data structure and the storage addresses in the first memory.Specifically, according to the weight names stored in the weight description information and the storage addresses of the corresponding weight information in the first memory, the third weight information corresponding to each network layer in the target model can be read from the first memory layer by layer in a preset reading order and written into the weight space of the corresponding network layer in the target model in the first memory; or, according to the weight names stored in the weight description information and the storage addresses of the corresponding weight information in the first memory, the third weight information corresponding to each network layer in the target model can be read from the first memory at one time and written into the weight space of the corresponding network layer in the target model in the first memory. In some embodiments, based on the weight names and corresponding weight information stored in the weight description information and the storage addresses in the first memory of the weight information, third weight information corresponding to each network layer in the target model is read from the first memory layer by layer and written into the weight space of the corresponding network layer in the target model in the first memory. This includes: if the target model is not in the memory of the second physical computing resource object, immediately loading the serialized file corresponding to the target model from a persistent storage medium and deserializing the serialized file corresponding to the target model to obtain the target model. The memory of the second physical computing resource object is the memory for executing the model weight update method, such as a CPU. During the deserialization process, for any network layer in the deserialized target model, based on the weight names and corresponding weight information stored in the weight description information and the second data structure, third weight information corresponding to any network layer is read from the first memory. Weight space is allocated for any network layer in the first memory, and the third weight information is written into the weight space allocated for any network layer. Further, optionally, during the compilation and optimization of the first original model, optimization strategy information used for compilation and optimization of each network layer in the first original model can be saved. Accordingly, writing the third weight information into the weight space in the first memory corresponding to the network layer in the target model includes: reordering the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain reordered weight information; and writing the reordered weight information into the weight space in the first memory corresponding to the network layer in the target model to facilitate injection of the reordered weight information into the target model. The reordering of the third weight information is similar to the reordering operation performed on the first weight information during the compilation optimization process, including, but not limited to, operations such as merging, deleting, and sorting the weight information, depending on the compilation optimization strategy.It should be noted that the model weight update methods provided in the above or following embodiments of the present disclosure can be executed by various inference acceleration engines. To implement the model weight update methods provided in the embodiments of the present disclosure in an inference acceleration engine, as shown in FIG1c , the following functional modules can be added to the inference acceleration engine: a weight registration module 21, a weight acquisition module 22, and a weight loading module 23. Of course, the inference acceleration engine also includes a model loading module 11, a computation graph conversion module 12, a compilation optimization module 13, a serialization module 14, and a model inference module 15. The coordinated cooperation of these modules provides users with seamless weight update capabilities without requiring compilation, enabling rapid model switching or deployment. The model loading module 11 is configured to deserialize the serialized file of the first original model to obtain the first original model in memory. The computation graph conversion module 12 is configured to perform computation graph conversion on the first original model to obtain a computation graph corresponding to the first original model. The compilation and optimization module 13 is configured to compile and optimize the computation graph to obtain a compiled model, wherein the compiled model includes the model structure information and first weight information of the first original model. The model structure information in the compiled model serves as the target model. The serialization module 14 is configured to serialize the target model to obtain a serialized file corresponding to the target model and persistently store the serialized file corresponding to the target model. The weight registration module 21 is configured to obtain the type and name of the weights in each network layer of the first original model during the computation graph conversion process performed by the computation graph conversion module. Based on the construction method of each network layer, the weight type and name in each network layer are associated and stored in an appropriate storage method. This process is also called the weight registration process. It should be noted that the registered weight types and names, as weight description information, are serialized into the target model's serialization file as part of the target model. This weight description information is reread when switching to a second original model to perform a deep learning task, and the second weight information is injected accordingly, thereby achieving continuous tracking of weight information. The weight registration module 21 records and identifies weight description information, helping each network layer record its own weight types and names. The weight acquisition module 22 is used to obtain the second weight information corresponding to the second original model when a second original model is isomorphic to the first original model and is used to perform a deep learning task.Specifically, the weight acquisition module 22 is configured to construct a second data structure compatible with the target model. The second data structure is configured to store weight names and the storage addresses of their corresponding weight information. The module also provides a function interface for external modules to update the second data structure, allowing external modules to update the weight names and the storage addresses of their corresponding weight information stored in the second data structure through the function interface. The second data structure may be, but is not limited to, a hash table containing weight names and weight data addresses. The aforementioned external module primarily refers to the model inference module 15. When the second original model is enabled to perform a deep learning task, the model inference module 15 is configured to load a serialized file corresponding to the second original model from a persistent storage medium; deserialize the serialized file corresponding to the second original model to obtain the second original model; and, during the deserialization process, obtain the second weight information corresponding to any network layer in the deserialized second original model. The module then writes the name and storage location of the second weight information corresponding to any network layer into the second data structure by calling the function interface provided by the weight acquisition module, allowing the weight acquisition module to obtain the second weight information. Further optionally, the model inference module 15 writes the name and storage location of the second weight information corresponding to any network layer into the second data structure, including: preprocessing the second weight information to obtain third weight information, and relocating the third weight information from the second memory to the first memory; calling the function interface provided by the weight acquisition module to add the storage address of the third weight information in the first memory to the second data structure according to the weight name. Based on the weight registration module 21 and the weight acquisition module 22, the weight loading module 23 allows each network layer in the target model to obtain the new weight information corresponding to itself during weight injection, so as to perform the corresponding deep learning task based on the injected new weight information. Specifically, the weight loading module 23 is configured to inject the second weight information into the target model based on the weight description information, and to perform the deep learning task using the target model injected with the second weight information.Further optionally, the weight loading module 23 reads the weight description information from the first data structure, reads the storage address of the third weight information in the first memory from the second data structure based on the weight type and name in the weight description information, reads the third weight information from the first memory based on the storage address of the third weight information in the first memory, and performs rearrangement operations such as merging, deleting, and sorting on the third weight information according to a pre-stored compilation optimization strategy. The module then injects the rearranged third weight information into the target model, and controls the target model injected with the rearranged third weight information to execute the deep learning task originally intended for the second original model. Further optionally, if the target model is not resident in memory, the weight loading module 23 is further configured to deserialize the serialized file of the target model to obtain the target model, and then inject the second weight information into the target model using the aforementioned operations, so that the target model injected with the second weight information can execute the deep learning task. In the disclosed embodiments, the model inference process is significantly simplified compared to the inference process in traditional model switching scenarios. This process primarily includes model switching, deserialization, and weight update operations, eliminating the compilation and optimization process. This significantly improves model switching and inference efficiency while retaining the performance advantages of compilation and optimization. The beneficial effects of the disclosed embodiments are particularly significant in scenarios with a large number of models, significantly unleashing user AI productivity and enhancing the core competitiveness of various cloud computing instances that provide cloud computing services based on deep learning models, such as the Elastic Compute Service (ECS), among similar cloud vendor offerings. Figure 2 is a flow diagram of a weight update method provided by an exemplary embodiment of the present disclosure. This method is applied to a model inference engine. As shown in Figure 2, the method includes the following steps:
[0003] 201. Pre-compile and optimize the first original model to obtain a target model without weight information, and save the type and name of the weights required by the first original model as weight description information adapted to the target model;
[0004] 202. When a second original model is isomorphic to the first original model and is enabled to perform a deep learning task, obtain second weight information corresponding to the second original model, where the first original model has first weight information;
[0005] 203. Inject the second weight information into the target model based on the weight description information, so as to perform the deep learning task using the target model with the injected second weight information. In this embodiment, the first and second primitive models are isomorphic models. This embodiment does not limit the model structure or model function of the isomorphic models to which the first and second primitive models belong. For example, the first and second primitive models may both be used for object recognition, with the first primitive model specifically used for part recognition and the second primitive model specifically used for animal recognition. Alternatively, the first and second primitive models may both be used for generating home decoration plans, with the first primitive model specifically used for generating Nordic-style home decoration plans and the second primitive model specifically used for generating Classical-style home decoration plans. The above embodiments are merely illustrative and do not limit the technical solutions of this disclosure. In this embodiment, any primitive model under the isomorphic model can be selected as the first primitive model, and the first primitive model can be compiled and optimized to obtain a target model without weight information, but with only model structure information. Since the target model is compiled from the first primitive model, the weight types and names required by the target model are the same as those required by the first primitive model. Therefore, the types and names of the weights required by the first original model can be saved as weight description information adapted to the target model. This weight description information is used to describe the weight information required by the target model and to provide conditions for subsequently injecting the weight information into the target model. It should be noted that compiling and optimizing the first original model can be performed after a deep learning task has been received. In this way, while obtaining the target model, the compiled and optimized model (hereinafter referred to as the compiled model) can also be used to execute the deep learning task. Alternatively, compiling and optimizing the first original model can be performed before a deep learning task has been received, with the primary purpose of obtaining the target model. Furthermore, optionally, before compiling and optimizing the first original model, the method further includes: deserializing the serialized file of the first original model to obtain the first original model. The serialized file of the first original model is stored in a persistent storage medium. The process of deserializing the serialized file of the first original model is to load the serialized file of the first original model from the persistent storage medium into a memory, and deserialize the serialized file in the memory to obtain an executable model object, which is the first original model.Before loading the serialized file of the first original model from the persistent storage medium, a model training and serialization process is also included. These processes can be found in the relevant descriptions of the above embodiments and are not further described here. In this embodiment, compiling and optimizing the first original model to obtain a target model without weight information includes: performing computation graph conversion on the first original model to obtain a computation graph corresponding to the first original model; compiling and optimizing the computation graph to obtain a compiled model, the compiled model including model structure information and first weight information; and using the model structure information in the compiled model as the target model. In some embodiments, a computation graph conversion is performed on the first original model to obtain a computation graph corresponding to the first original model in one implementation process as follows: the first original model is parsed to obtain each network layer included in the first original model, operators involved in each network layer, inputs and outputs of operators in each network layer, and weight types and names of each network layer, as well as weight information corresponding to each weight name (i.e., first weight information), etc., wherein operators include at least addition, subtraction, multiplication, division, convolution, pooling, normalization, etc., and generally, one network layer corresponds to one node; further, based on the ordered connection relationship between each network layer, the layers are mapped into directed edges between the nodes corresponding to each network layer, and the inputs and outputs between each network layer are mapped into inputs and outputs between the nodes corresponding to each network layer; further, based on each node, directed edges, and input and output relationships between each node, a computation graph of the first original model is obtained. In this embodiment, after obtaining the weight types and names in each network layer of the first original model, the weight types and names required by the first original model are saved as weight description information adapted for the target model. This can be understood as a registration process for the weight description information of the first original model. Based on this weight description information, the mapping relationship between the target model and the weight information can be dynamically updated during the weight injection process, thereby enabling the transfer of weight information in each network layer. The weight description information registration process varies depending on the construction method of the network layer of the first original model. The weight description information registration process can be implemented during the computation graph conversion process.Specifically, during the computation graph conversion process, the computation graph conversion can be performed layer by layer. For the target network layer currently being converted, which is any network layer in the first original model, the target network layer is converted into a target node in the computation graph. Furthermore, the construction method of the target network layer can be identified based on its attribute information. If the target network layer is constructed as a plug-in, the type and name of the weights in the target network layer are saved to the target node via parameter passing, specifically implementing the type and name of the weights in the target network layer as the attribute information of the target node. If the target network layer is constructed as a built-in layer, a first data structure is created, and the identification information of the target network layer and the type and name of the weights in the target network layer are correspondingly stored in the first data structure. The first data structure is associated with the target model and is used to store the weight types and names of each network layer required by the target model. In the embodiments of the present disclosure, the implementation method of the first data structure is not limited and can be an array, a list, a hash table, etc. As shown in Figure 3a, the first data structure is a composite hash table. This composite hash table includes a correspondence between network layers (Layer) and weight lists (List). The weight list includes a tuple consisting of a weight type (WeightsKind) and a weight name (WeightsName), simply represented as Layer->List[Tuple(WeightsKind, WeightsName)]. In Figure 3a, a user-built network, also known as the first original model, is constructed. For example, the first original model has three network layers, denoted as Convolution-1 > Convolution-2 > Layer-1. Convolution-1 and Convolution-2 represent two convolutional layers, and Layer-1 represents a normalization layer. After weight registration for Convolution-1 > Convolution-2 > Layer-1, a computational graph with weight registration information is obtained.Convolution-1 includes two weights, named Conv.1.kernel and Conv.1.bias; Convolution-2 also includes two weights, named Conv.2.kernel and Conv.2.bias; and Layernorm-1 includes one weight, named Layernorm.1.scale. The weight types and names in the order Convolution-1 > Convolution-2 > Normalization Layernorm-1 are stored as tuples in a List. A correspondence is established between each network layer and the List, achieving the goal of storing the weight types and names of each network layer in a composite hash table data structure. Then, based on the computational graph with weight registration information, model compilation, optimization, and serialization are performed to obtain the target model serialization file. In Figure 3a, the weight information corresponding to each weight name is set to the default value, for example, a vector with a value of 0. In some embodiments, after obtaining the computation graph of the first original model, the computation graph can be compiled and optimized. Specifically, each node in the computation graph can be optimized using an optimization strategy compatible with the first original model to obtain a compiled model. The compiled model includes a model structure and first weight information, where the first weight information is the value of the first weight of the first original model. Furthermore, to ensure the persistence of the target model, the target model can be persistently stored. Because running memory space is limited, the target model can be serialized to obtain a serialized file corresponding to the target model, and the serialized file corresponding to the target model can be persistently stored. In this embodiment, when a deep learning task requiring a second original model isomorphic to the first original model is received, the second original model is activated to execute the deep learning task, and the second weight information corresponding to the second original model is obtained.Optionally, when enabling a second original model isomorphic to the first original model to perform a deep learning task, obtaining second weight information corresponding to the second original model includes: loading a serialized file corresponding to the second original model from a persistent storage medium when enabling the second original model to perform the deep learning task; deserializing the serialized file corresponding to the second original model to obtain the second original model; and, during the deserialization process, obtaining second weight information corresponding to any network layer in the deserialized second original model. Deserializing the serialized file involves loading the serialized file from the persistent storage medium into memory, deserializing the serialized file in memory to obtain an executable model object, and storing the second original model in the memory as a model object. It should be noted that the storage and injection of the model weight information are both performed in the memory of a second physical computing resource object that executes the weight update method. The second physical computing resource object is preferably a CPU. Furthermore, the second weight information can be injected into a target model based on the weight description information, so that the target model with the injected second weight information can execute the deep learning task. The target model is a model compiled and optimized based on the first original model, and has high performance. For a second original model that is isomorphic to the first original model, when executing a deep learning task, the second original model does not need to be compiled and optimized; instead, the second weight information needs to be injected into the target model. From the perspective of the second original model, this maintains the performance advantages of compilation optimization while also improving model deployment efficiency. In some embodiments, injecting the second weight information into the target model based on weight description information includes: preprocessing the second weight information to obtain third weight information, and relocating the third weight information from the second memory to the first memory; and reading the third weight information from the first memory based on the weight description information and injecting it into the target model. The first memory is the memory of a first physical computing resource object (e.g., a GPU) running the target model, while the second weight information is stored in the memory (i.e., the second memory) of a second physical computing resource object (e.g., a CPU) executing the model weight update method. The preprocessing may include operations such as merging and renaming some weights.In some embodiments, reading third weight information from a first memory and injecting it into a target model based on weight description information includes: constructing a second data structure compatible with the target model, the second data structure being used to store weight names and storage addresses of corresponding weight information; adding the storage addresses of the third weight information in the first memory to the second data structure corresponding to the weight names; and reading the third weight information from the first memory and injecting it into the target model based on the weight description information and the second data structure. In some embodiments, the target model is persistently resident in the memory of a second physical computing resource object (hereinafter referred to as resident memory). If the target model is persistently resident in the memory of the second physical computing resource object, the third weight information can be read from the first memory and injected into the target model based on the weight names and storage addresses of the corresponding weight information in the first memory stored in the weight description information and the second data structure. Specifically, according to the weight names stored in the weight description information and the storage addresses of the corresponding weight information in the first memory, the third weight information corresponding to each network layer in the target model can be read from the first memory layer by layer in a preset reading order and written into the weight space of the corresponding network layer in the target model in the first memory; or, according to the weight names stored in the weight description information and the storage addresses of the corresponding weight information in the first memory, the third weight information corresponding to each network layer in the target model can be read from the first memory at one time and written into the weight space of the corresponding network layer in the target model in the first memory. In some embodiments, according to the weight name and the storage address of the corresponding weight information stored in the weight description information and the second data structure in the first memory, the third weight information corresponding to each network layer in the target model is read from the first memory layer by layer and written into the weight space of the corresponding network layer in the target model in the first memory, including: if the target model is not in the memory of the second physical computing resource object, the serialized file corresponding to the target model is immediately loaded from the persistent storage medium, and the serialized file corresponding to the target model is deserialized to obtain the target model. The memory of the second physical computing resource object is the memory for executing the model weight update method, such as the CPU; during the deserialization process, for any network layer in the deserialized target model, according to the weight name and the storage address of the corresponding weight information stored in the weight description information and the second data structure in the first memory, the third weight information corresponding to any network layer is read from the first memory; a weight space is allocated to any network layer in the first memory, and the third weight information is written into the weight space allocated to any network layer.In this embodiment, based on the first and second data structures, each network layer can obtain its own corresponding new weights during deserialization. However, these new weights require weight reordering. Because each network layer uses a different optimization strategy during compilation optimization, the corresponding weight distribution is also different. Therefore, weight reordering is a major challenge in the weight update process. To address this issue, the weight reordering process can be separated into an independent function interface. Each network layer will record its own reordering algorithm and, after obtaining a new weight data pointer, call its corresponding reordering algorithm to achieve weight reordering and loading. Based on this, during the compilation optimization process of the first original model, the optimization strategy information used for the compilation optimization of each network layer in the first original model can be saved. Accordingly, writing the third weight information into the weight space of the corresponding network layer in the target model in the first memory includes: reordering the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain the reordered weight information; and writing the reordered weight information into the weight space of the corresponding network layer in the target model in the first memory. It should be noted that compiling and optimizing the first original model essentially compiles and optimizes the computation graph corresponding to the first original model. Therefore, the optimization strategy information used for compiling and optimizing each network layer in the first original model is the optimization strategy information used for compiling and optimizing each node in the computation graph. For ease of understanding, the following describes in detail the weight information update process, using the example of the first original model being the M model and the second original model being the N model, in conjunction with Figures 3a and 3b. Assume that the M model includes two Buildin convolutional layers and one Plugin normalization layer. First, the user constructs the M model network, with three network layers: Convolution-1, Convolution-2, and Layernorm-1. Weights are registered for each of the three network layers during the computation graph conversion process for the M model. The network layer Layernorm-1 sets the weight name to Layernorm.1.scale through parameter passing; and the Convolution-1 and Convolution-2 nodes are automatically associated with the corresponding network layers through the weight name in the form of a composite hash table. The weight names of each network layer are Conv.1.kerneR Conv.1.bias, Conv.2.kernel, and Conv.2.bias respectively.After the model is compiled and optimized, the weight types and names of each network layer are serialized into the model's serialization file along with the target model obtained from the compilation optimization. In Figure 3a, taking the nodes in the computational graph as an example, three network layers, Convolution-1, Convolution-2, and Layernorm-1, are illustrated. Further, when it is necessary to switch from model M to model N in the isomorphic models, model N also contains two Buildin convolutional layers and one Plugin normalization layer. Each network layer has its own weight information. During the weight acquisition process, the user sets the weights and sets the storage addresses of the weight information of each network layer in model N in the form of key pairs into the weight data hash table shown in Figure 3b. The new storage addresses of the five weight information, Conv.1.kernel, Conv.1.bias, Conv.2.kernel, Conv.2.bias, and Layernorm.1.scale, are 0x0000010, 0x0000020, 0x0000030, 0x0000040, and 0x0000050 respectively. Then, during the weight loading process, the deserialization operation is performed on the target model serialization file obtained from the compilation optimization. During the deserialization process, each network layer will look for the corresponding storage address in the hash table shown in Figure 3b according to its associated weight name. For example, the Convolution-1 node will rearrange and update the weight information at the 0x0000010 address into the Conv.1.kernel weight, and rearrange and update the weight information at the 0x0000020 address into the Conv.2.bias weight. And so on until all nodes are deserialized. At this time, the weight update work is synchronously completed, and the model with updated weights is obtained. Throughout the process, what the user actually feels is a model deserialization process, and such a process does not change the user's usage habits and is relatively imperceptible. At the same time, this weight update mechanism breaks the limitations in fusion, enabling the user to enjoy better performance acceleration while implementing weight updates. Figure 4 is a schematic structural diagram of the model weight update device provided by an exemplary embodiment of the present disclosure.As shown in FIG4 , the apparatus includes: a compilation and optimization module 41 configured to pre-compile and optimize a first original model to obtain a target model without weight information, and to store the types and names of weights required by the first original model as weight description information adapted for the target model; an acquisition module 42 configured to, when a second original model isomorphic to the first original model is enabled to perform a deep learning task, obtain second weight information corresponding to the second original model, where the first original model has the first weight information; and an injection module 43 configured to inject the second weight information into the target model based on the weight description information, so as to perform the deep learning task using the target model injected with the second weight information. In this embodiment, when the compilation and optimization module 41 is configured to pre-compile and optimize the first original model to obtain a target model without weight information, the module is specifically configured to: perform computation graph conversion on the first original model to obtain a computation graph corresponding to the first original model; perform compilation and optimization processing on the computation graph to obtain a compiled model, where the compiled model includes model structure information and first weight information; and use the model structure information in the compiled model as the target model. Further optionally, the system further includes: a processing module and a storage module, wherein the processing module is configured to serialize the target model to obtain a serialized file corresponding to the target model; and the storage module is configured to persistently store the serialized file corresponding to the target model. In this embodiment, when the compilation and optimization module 41 is configured to store the type and name of the weights required by the first original model as weight description information adapted for the target model, the system is specifically configured to: during the computation graph conversion process, determine the target network layer currently being converted, where the target network layer is any network layer in the first original model, and the target network layer is converted into a target node in the computation graph; if the target network layer is constructed in a plug-in manner, save the type and name of the weights in the target network layer to the target node via parameter passing; if the target network layer is constructed in a built-in manner, create a first data structure, and store the identification information of the target network layer and the type and name of the weights in the target network layer in the first data structure.In this embodiment, when a second original model is isomorphic to the first original model and a second weight information corresponding to the second original model is enabled to perform a deep learning task, the acquisition module 42 is specifically configured to: load a serialized file corresponding to the second original model from a persistent storage medium when the second original model is enabled to perform the deep learning task; deserialize the serialized file corresponding to the second original model to obtain the second original model; and, during the deserialization process, obtain the second weight information corresponding to any network layer in the deserialized second original model. In this embodiment, when the injection module 43 is configured to inject the second weight information into the target model based on the weight description information, the injection module 43 is specifically configured to: preprocess the second weight information to obtain third weight information, and relocate the third weight information from a second memory to a first memory, where the first memory is the memory of a first physical computing resource object running the target model, and the second memory is the memory of a second physical computing resource object running the model weight update method; and read the third weight information from the first memory based on the weight description information and inject it into the target model. Optionally, when the injection module 43 is configured to read the third weight information from the first memory and inject it into the target model according to the weight description information, it is specifically configured to: construct a second data structure adapted to the target model, the second data structure being used to store the weight name and the storage address of the corresponding weight information; according to the weight name, add the storage address of the third weight information in the first memory to the second data structure; and according to the weight description information and the second data structure, read the third weight information from the first memory and inject it into the target model. Optionally, when the injection module 43 is configured to read the third weight information from the first memory and inject it into the target model according to the weight description information and the second data structure, it is specifically configured to: read the third weight information corresponding to each network layer in the target model from the first memory layer by layer according to the weight description information and the second data structure and write it into the weight space of the corresponding network layer in the target model in the first memory; or read the third weight information corresponding to each network layer in the target model from the first memory at one time according to the weight description information and the second data structure and write it into the weight space allocated to the corresponding network layer in the target model in the first memory.Optionally, when the injection module 43 is configured to read, layer by layer, the third weight information corresponding to each network layer in the target model from the first memory based on the weight description information and the second data structure, and write the third weight information corresponding to the corresponding network layer in the target model into the weight space allocated in the first memory for the corresponding network layer in the target model, the injection module 43 is specifically configured to: if the target model is not in the memory of the second physical computing resource object, load the serialized file corresponding to the target model from the persistent storage medium, and deserialize the serialized file corresponding to the target model to obtain the target model; during the deserialization process, for any network layer in the deserialized target model, read the third weight information corresponding to the any network layer from the first memory based on the weight description information and the second data structure; allocate weight space for the any network layer in the first memory, and write the third weight information into the weight space allocated for the any network layer. Further optionally, the storage module is further configured to, during the compilation and optimization process of the first original model, store optimization strategy information used for compiling and optimizing each network layer in the first original model. Accordingly, when the injection module 43 is configured to write the third weight information into the weight space in the first memory corresponding to the network layer in the target model, it is specifically configured to: rearrange the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain rearranged weight information; and write the rearranged weight information into the weight space in the first memory corresponding to the network layer in the target model. The detailed implementation and beneficial effects of each module in this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated on here. Figure 5 is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. As shown in Figure 5, a memory 50a and a processor 50b; the memory 50a stores a computer program, and the processor 50b is coupled to the memory 50a for executing the computer program to implement the following steps: pre-compiling and optimizing the first original model to obtain a target model without weight information, and saving the type and name of the weights required by the first original model as weight description information adapted to the target model; when enabling a second original model isomorphic to the first original model to perform a deep learning task, obtaining second weight information corresponding to the second original model, where the first original model has the first weight information; and injecting the second weight information into the target model according to the weight description information, so as to perform the deep learning task using the target model injected with the second weight information.In this embodiment, when the processor 50b is configured to pre-compile and optimize the first original model to obtain a target model without weight information, it is specifically configured to: perform computation graph conversion on the first original model to obtain a computation graph corresponding to the first original model; compile and optimize the computation graph to obtain a compiled model, where the compiled model includes model structure information and first weight information; and use the model structure information in the compiled model as the target model. Furthermore, optionally, the processor 50b is further configured to serialize the target model to obtain a serialized file corresponding to the target model; and persistently store the serialized file corresponding to the target model. In this embodiment, when the processor 50b is used to save the type and name of the weights required by the first original model as weight description information adapted to the target model, it is specifically used to: during the calculation graph conversion process, determine the target network layer in the current conversion, where the target network layer is any network layer in the first original model, and the target network layer is converted into a target node in the calculation graph; if the target network layer is constructed in a plug-in manner, save the type and name of the weights in the target network layer to the target node by parameter passing; if the target network layer is constructed in a built-in manner, create a first data structure, and store the identification information of the target network layer and the type and name of the weights in the target network layer in the first data structure. In this embodiment, the processor 50b is used to obtain the second weight information corresponding to the second original model when enabling the second original model to perform a deep learning task. Specifically, it is used to: load the serialized file corresponding to the second original model from the persistent storage medium when enabling the second original model to perform the deep learning task; deserialize the serialized file corresponding to the second original model to obtain the second original model; and during the deserialization process, obtain the second weight information corresponding to any network layer in the deserialized second original model.In this embodiment, when the processor 50b is configured to inject the second weight information into the target model based on the weight description information, it is specifically configured to: preprocess the second weight information to obtain third weight information, and relocate the third weight information from a second memory to a first memory, where the first memory is the memory of a first physical computing resource object running the target model; and the second memory is the memory of a second physical computing resource object running the model weight update method; read the third weight information from the first memory based on the weight description information and inject it into the target model. Optionally, when the processor 50b is configured to read the third weight information from the first memory based on the weight description information and inject it into the target model, it is specifically configured to: construct a second data structure compatible with the target model, the second data structure being configured to store weight names and storage addresses of corresponding weight information; add the storage addresses of the third weight information in the first memory to the second data structure based on the weight names; and read the third weight information from the first memory based on the weight description information and the second data structure and inject it into the target model. Optionally, when the processor 50b is used to read the third weight information from the first memory and inject it into the target model according to the weight description information and the second data structure, it is specifically used to: read the third weight information corresponding to each network layer in the target model from the first memory layer by layer according to the weight description information and the second data structure and write it into the weight space allocated to the corresponding network layer in the target model in the first memory; or read the third weight information corresponding to each network layer in the target model from the first memory at one time according to the weight description information and the second data structure and write it into the weight space allocated to the corresponding network layer in the target model in the first memory.Optionally, when the processor 50b is used to read the third weight information corresponding to each network layer in the target model from the first memory layer by layer and write it into the weight space allocated to the corresponding network layer in the target model in the first memory according to the weight description information and the second data structure, it is specifically used to: if the target model is not in the memory of the second physical computing resource object, load the serialized file corresponding to the target model from the persistent storage medium, and deserialize the serialized file corresponding to the target model to obtain the target model; during the deserialization process, for any network layer in the deserialized target model, read the third weight information corresponding to any network layer from the first memory according to the weight description information and the second data structure; allocate weight space for any network layer in the first memory, and write the third weight information into the weight space allocated to any network layer. Further optionally, the processor 50b is further configured to, during the compilation and optimization process of the first original model, store optimization strategy information used for compiling and optimizing each network layer in the first original model. Accordingly, when writing the third weight information into the weight space of the corresponding network layer in the target model in the first memory, the processor 50b is specifically configured to: rearrange the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain rearranged weight information; and write the rearranged weight information into the weight space of the corresponding network layer in the target model in the first memory. Furthermore, as shown in FIG5 , the electronic device further includes other components, such as a communication component 50c, a display 50d, a power supply component 50e, and an audio component 50f. FIG5 schematically illustrates only some components and does not imply that the electronic device includes only the components shown in FIG5 . The detailed implementation and beneficial effects of each module in the method of this embodiment have been described in detail in the previous embodiments and will not be elaborated upon here. Accordingly, an exemplary embodiment of the present disclosure further provides a computer-readable storage medium storing a computer program / instructions. When the computer program / instructions are executed by a processor, the processor is enabled to implement each step in the above-described method embodiment. Accordingly, an exemplary embodiment of the present disclosure further provides a computer program product. The computer program product includes the computer program / instructions. When the computer program is executed by a processor, the processor is enabled to implement each step in the above-described method embodiment.The aforementioned memory can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G, or other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID), infrared data association (IrDA), ultra-wideband (UWB), Bluetooth (BT), and other technologies. The aforementioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only detect the boundaries of a touch or slide action, but also the duration and pressure associated with the touch or slide action. The aforementioned power supply component provides power to the various components of the device in which the power supply component is located.The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides. The aforementioned audio component may be configured to output and / or input audio signals. For example, the audio component may include a microphone (MIC) configured to receive external audio signals when the device in which the audio component resides is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals may be further stored in a memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals. Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program code. The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process flow and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing device, produce a device for implementing the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions can also be stored in a computer-readable memory capable of directing the computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.In a typical configuration, a computing device includes one or more processors (Central Processing Units, CPUs), input / output interfaces, network interfaces, and memory. Memory may take the form of non-persistent storage in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. oMemory is an example of computer-readable media. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmitting medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves. It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus comprising a list of elements may include not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, product, or apparatus. Without further limitation, the phrase "comprising a..." does not preclude the presence of additional identical elements in the process, method, product, or apparatus comprising the elements. The foregoing are merely examples of the present disclosure and are not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations of the present disclosure are possible. Any modifications, equivalent substitutions, improvements, and the like made within the spirit and principles of the present disclosure are intended to be encompassed by the claims of the present disclosure. INDUSTRIAL APPLICABILITY In an embodiment of the present disclosure, a first original model is pre-compiled and optimized to obtain a target model without weight information, and the types and names of the weights required by the first original model are saved as weight description information adapted to the target model. When a second original model is isomorphic to the first original model and is enabled to perform a deep learning task, second weight information corresponding to the second original model is obtained. Based on the weight description information, the second weight information is injected into the target model, so that the target model with the injected second weight information performs the deep learning task.For isomorphic models, only one compilation optimization process is required, achieving the goal of eliminating compilation when switching between isomorphic models and saving compilation optimization time. Furthermore, while achieving the performance benefits brought by compilation optimization, switching between isomorphic models can also be achieved by injecting the second weight information into the target model, thereby improving model deployment efficiency.
Claims
Claims 1. A method for updating model weights, applied to a model inference engine, the method comprising: Pre-compile and optimize the first original model to obtain a target model without weight information, and save the type and name of the weights required by the first original model as weight description information adapted to the target model; when enabling a second original model isomorphic to the first original model to execute a deep learning task, obtain the second weight information corresponding to the second original model, where the first original model has first weight information; according to the weight description information, inject the second weight information into the target model to use the target model injected with the second weight information to execute the deep learning task.
2. The method according to claim 1, wherein Pre-compiling and optimizing the first original model to obtain a target model without weight information includes: performing a computational graph transformation on the first original model to obtain a computational graph corresponding to the first original model; performing a compilation optimization process on the computational graph to obtain a compiled model, where the compiled model includes model structure information and first weight information; using the model structure information in the compiled model as the target model.
3. The method according to claim 2, wherein It further includes: serializing the target model to obtain a serialized file corresponding to the target model; persistently storing the serialized file corresponding to the target model.
4. The method according to claim 2, wherein Saving the type and name of the weights required by the first original model as weight description information adapted to the target model includes: during the computational graph transformation process, determining the target network layer in the current transformation, where the target network layer is any network layer in the first original model and is transformed into a target node in the computational graph; if the construction method of the target network layer is the plug-in method, save the type and name of the weights in the target network layer to the target node in the form of parameter passing; if the construction method of the target network layer is the built-in method, create a first data structure and store the identification information of the target network layer and the type and name of the weights in the target network layer correspondingly in the first data structure.
5. The method according to claim 1, wherein When enabling a second original model isomorphic to the first original model to execute a deep learning task, obtaining the second weight information corresponding to the second original model includes: when enabling the second original model to execute a deep learning task, loading the serialized file corresponding to the second original model from a persistent storage medium; deserializing the serialized file corresponding to the second original model to obtain the second original model; and during the deserialization process, for any network layer in the deserialized second original model, obtaining the second weight information corresponding to the any network layer.
6. The method according to any one of claims 1-5, wherein According to the weight description information, the second Inject the second weight information into the target model, including: preprocessing the second weight information to obtain third weight information, and migrating the third weight information from the second memory to the first memory; according to the weight description information, reading the third weight information from the first memory and injecting it into the target model; wherein, the first memory is the memory of the first physical computing resource object running the target model, and the second memory is the memory of the second physical computing resource object for executing the weight update method.
7. The method according to claim 6, wherein According to the weight description information, reading the third weight information from the first memory and injecting it into the target model, including: constructing a second data structure adapted to the target model, where the second data structure is used to store the weight name and the storage address of its corresponding weight information; according to the weight name, correspondingly adding the storage address of the third weight information in the first memory to the second data structure; according to the weight description information and the second data structure, reading the third weight information from the first memory and injecting it into the target model.
8. The method according to claim 7, wherein, According to the weight description information and the second data structure, reading the third weight information from the first memory and injecting it into the target model, including: according to the weight description information and the second data structure, layer by layer reading the third weight information corresponding to each network layer in the target model from the first memory and writing it into the weight space allocated to the corresponding network layer in the first memory for the target model; or according to the weight description information and the second data structure, reading at once the third weight information corresponding to each network layer in the target model from the first memory and writing it into the weight space allocated to the corresponding network layer in the first memory for the target model.
9. The method according to claim 8, wherein, According to the weight description information and the second data structure, layer by layer reading the third weight information corresponding to each network layer in the target model from the first memory and writing it into the weight space allocated to the corresponding network layer in the first memory for the target model, including: if the target model is not in the second memory, loading the serialized file corresponding to the target model from the persistent storage medium, and deserializing the serialized file corresponding to the target model to obtain the target model; during the deserialization process, for any network layer deserialized from the target model, according to the weight description information and the second data structure, reading the third weight information corresponding to the any network layer from the first memory; allocating a weight space for the any network layer in the first memory, and writing the third weight information into the weight space allocated for the any network layer.
10. The method according to claim 8, wherein, Further included: During the compilation optimization of the first original model the optimization strategy information used for the compilation optimization of each network layer in the first original model is saved; Correspondingly, writing the third weight information into the weight space of the corresponding network layer in the first memory of the target model, including: rearranging the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain the rearranged weight information; Writing the rearranged weight information into the weight space of the corresponding network layer in the first memory of the target model.
11. An electronic device, comprising: A memory and a processor; The memory stores a computer program, and the processor is coupled to the memory and is configured to execute the computer program to implement the steps in the method according to any one of claims 1-10.
12. A computer-readable storage medium storing a computer program / instruction, when the computer program / instruction is executed by a processor, causing the processor to be able to implement the steps in the method according to any one of claims 1-10.
13. A computer program product, the computer program product includes a computer program / instruction, when the computer program / instruction is executed by a processor, causing the processor to be able to implement the steps in the method according to any one of claims 1-10.
Citation Information
Patent Citations
TensorRT-based target detection model acceleration method and device
CN112668672A
Novel method for identifying identity of underground coal mine personnel with weak characteristics
CN117152849A
Deep learning model inference acceleration method and system, and device and medium
WO2022057468A1