Model weight updating method and device, storage medium and program product

Through the concept of isomorphic model, the use of weight description information to achieve rapid switching and deployment of deep learning models, solving the problem of too long compilation optimization time and improving model deployment efficiency and switching speed.

CN120297348APending Publication Date: 2025-07-11HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410044771.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the compilation and optimization time of deep learning models is too long, resulting in low model deployment efficiency. Especially when the model is large, the compilation time may be as long as several hours, while the inference time is only ten seconds, which affects the practical application efficiency of the model.

Method used

The concept of isomorphic model is proposed. By compiling and optimizing the first original model, it saves its weight type and name as weight description information, and directly injects corresponding weight information when switching to the isomorphic model, avoiding repeated compilation and optimization processes, and achieving rapid switching and deployment of the model.

Benefits of technology

While maintaining the performance benefits brought by compilation optimization, it significantly improves model deployment efficiency, especially in scenarios with huge models, reducing compilation optimization time and improving model switching and deployment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297348A_ABST
    Figure CN120297348A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model weight updating method and device, a storage medium and a program product. In the embodiment of the invention, the first original model is compiled and optimized in advance to obtain the target model without weight information, and the second original model isomorphic with the first original model is started to execute the deep learning task, so that the second weight information corresponding to the second original model can be directly injected into the target model; the target model can execute the deep learning task on the basis of the injected second weight information, and for the isomorphic model, only one compiling optimization process needs to be executed, so that the purpose of free compiling during isomorphic model switching is achieved, and the compiling optimization time is saved; furthermore, the second weight information is injected into the target model to realize the switching of the isomorphic model while the performance benefit brought by compiling optimization is obtained, and the model deployment efficiency can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, storage medium, and program product for updating model weights. Background Art

[0002] With the continuous development and improvement of computing power, data, and algorithms in human society, the training and deployment of deep learning models have become a hot research topic in the current field of artificial intelligence. However, the consumption of massive computing resources and the high deployment cost have become a major factor hindering the development of deep learning models.

[0003] To this end, each cloud computing manufacturer has gradually proposed some acceleration engines or frameworks for model inference, such as TensorRT. The TensorRT acceleration engine improves the running efficiency of the model on hardware resources such as Graphics Processing Unit (GPU) to a certain extent by compiling and optimizing the model. However, the compilation and optimization time also seriously affects the model deployment efficiency. Especially for larger model sizes, the compilation and optimization time can be as long as several hours, resulting in the pain point problem of "one hour of compilation and ten seconds of inference".

[0004] Therefore, how to improve the model deployment efficiency while obtaining the performance benefits brought by compilation and optimization has become an urgent problem to be solved in the industry. Summary of the Invention

[0005] Multiple aspects of this application provide a method, device, storage medium, and program product for updating model weights to improve the model deployment efficiency while obtaining the performance benefits brought by compilation and optimization.

[0006] An embodiment of this application provides a method for updating model weights, including: pre-compiling and optimizing a first original model to obtain a target model without weight information, and saving the type and name of the weights required by the first original model as weight description information adapted to the target model; when enabling a second original model isomorphic to the first original model to execute a deep learning task, obtaining second weight information corresponding to the second original model, where the first original model has first weight information; injecting the second weight information into the target model according to the weight description information, so as to use the target model with the injected second weight information to execute the deep learning task.

[0007] An embodiment of this application also provides an electronic device, including: a memory and a processor; a computer program is stored in the memory, and the processor is coupled to the memory and is used to execute the computer program to implement the steps in the above method.

[0008] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above method.

[0009] The embodiments of the present application further provide a computer program product, which includes a computer program / instructions that, when executed by a processor, enable the processor to implement the steps in the above method.

[0010] In the embodiments of the present application, the first original model is pre-compiled and optimized to obtain a target model without weight information, and the types and names of the weights required by the first original model are saved as weight description information adapted to the target model; when enabling the second original model isomorphic to the first original model to execute a deep learning task, the second weight information corresponding to the second original model is obtained; according to the weight description information, the second weight information is injected into the target model, so that the target model using the injected second weight information executes the deep learning task. For isomorphic models, only one compilation and optimization process needs to be performed, achieving the purpose of avoiding compilation when switching isomorphic models and saving the time of compilation and optimization; further, while obtaining the performance benefits brought by compilation and optimization, by injecting the second weight information into the target model to realize the switching of isomorphic models, the model deployment efficiency can also be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0012] Figure 1a is a schematic flowchart of a model training and model inference;

[0013] Figure 1b is a schematic flowchart of a model inference provided by an exemplary embodiment of the present application;

[0014] Figure 1c is an internal schematic diagram of an inference acceleration engine provided by an exemplary embodiment of the present application;

[0015] Figure 2 is a schematic flowchart of a model weight update method provided by an exemplary embodiment of the present application;

[0016] Figure 3a is a schematic diagram of a weight registration process based on a composite hash table provided by an exemplary embodiment of the present application;

[0017] Figure 3b is a schematic diagram of a weight information acquisition and loading process based on a hash table provided by an exemplary embodiment of the present application;

[0018] Figure 4 A structural schematic diagram of a model weight update device provided by an exemplary embodiment of the present application;

[0019] Figure 5 A structural schematic diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0020] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0021] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse. In addition, various models (including but not limited to language models or large models) involved in the present application comply with relevant laws and standards.

[0022] A deep learning model is a machine learning model trained and predicted through a multi-layer neural network. The core of a deep learning model is a neural network, which consists of multiple neurons (or nodes). Each neuron is connected to the neurons in the previous layer. By adjusting the connection weights between neurons, the neural network can learn the ability to extract useful information from input data during the training process and automatically learn to extract useful features from input data during the inference process, and then be used for tasks such as classification and regression.

[0023] In the embodiments of the present application, the type of the deep learning model is not limited. For example, it includes but is not limited to: existing deep learning models such as Restricted Boltzmann Machine (RBM), Autoencoder, Deep Belief Network (DBN), Deep Boltzmann Machine (DBM), Product Quantization Network (PQNet), Deep Perceptron (DP), Deep Feedforward Network (DFN), Convolutional Neural Networks (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Generative Adversarial Network (GAN), Auto-encoder (AE), Residual Neural Network (ResNet), and Attention Mechanism (AM), or various deep learning models that may emerge in the future.

[0024] Regardless of the type of deep learning model, the life cycle of the deep learning model includes a training stage and an inference stage. The training stage refers to the process of training an initial deep learning model using training samples to enable the deep learning model to learn and grow into a model that can solve specific problems; the inference stage refers to the process of using the deep learning model obtained in the training stage to solve specific problems.

[0025] To improve the efficiency of model training, as Figure 1a shown, a training acceleration engine can be used to train the model in the training stage. Among them, the training acceleration engine is a deep learning framework for model training, such as TensorFlow or PyTorch, etc. By adopting technologies such as parallel computing and distributed computing, and optimizing the hardware resources for model training, it is used to accelerate the training process of the deep learning model and reduce the training time. Among them, the hardware resources for model training can be Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), etc. As Figure 1aAs shown, the model training process using a training acceleration engine involves at least operations such as model construction, model compilation, model optimization, and model serialization.

[0026] Among them, model construction involves the following: 1. Clearly define the research purpose. Before constructing a model, it is necessary to first determine the purpose and requirements of the research. Clearly defining the research purpose helps to select appropriate mathematical models and parameters. 2. Define variables. Before transforming the research object into a mathematical model, it needs to be abstracted into some variables. Variables can be quantities, states, or characteristics, etc. Defining variables can clarify the characteristics and attributes of the research object. 3. Select a mathematical model. Based on the research purpose and defined variables, select an appropriate mathematical model. Mathematical models can be linear models, non-linear models, probability models, etc. Selecting a mathematical model requires comprehensive consideration of the research purpose, variable characteristics, and data types, etc.

[0027] Model compilation refers to selecting appropriate loss functions, optimizers, and evaluation metrics for the model before training the model and binding them to the model. Model compilation is a static process and only needs to be performed once before starting training. Model compilation involves the following: 1. Define model parameters. Before compiling the model, it is necessary to first define the parameters required by the model, such as the learning rate, batch size, etc. These parameters will affect the training process and final performance of the model. 2. Compile the model. Use the application programming interface (API) provided by the training acceleration engine (such as TensorFlow or PyTorch) to compile the defined model structure and parameters.

[0028] Model optimization refers to the process of improving the performance of a model by adjusting its parameters and hyperparameters during training, thereby enhancing the model's accuracy, generalization ability, and training efficiency. Model optimization is a dynamic process that requires continuous adjustment of the model's parameters and hyperparameters during training to achieve better performance. Model optimization involves the following aspects: 1. Setting the optimizer. The optimizer is used to update the weights of the model during training. Common optimizers include Stochastic Gradient Descent (SGD) optimizer, Adaptive Moment Estimation (Adam) optimizer, etc. Usually, an appropriate optimizer can be selected according to the requirements of the model and the characteristics of the dataset. 2. Setting the loss function. The loss function is used to measure the gap between the model's prediction results and the true values. According to the specific task type (classification, regression, etc.), an appropriate loss function is selected, such as cross-entropy loss, mean squared error, etc. 3. Setting the evaluation metrics. In addition to the loss function, appropriate evaluation metrics also need to be set to evaluate the performance of the model. For example, for classification tasks, metrics such as accuracy, precision, recall, etc. can be used; for regression tasks, metrics such as mean squared error, root mean squared error, etc. can be used. 4. Executing the model optimization process.

[0029] Optionally, model optimization involves but is not limited to the following: 1. Feature selection. Selecting the features most relevant to the target variable can reduce the number of features and improve the model's generalization ability. 2. Model adjustment. By adjusting the parameters or structure of the model, such as increasing or decreasing the number of layers, changing the activation function, etc., to improve the model's performance. 3. Regularization. By adding a regularization term to the loss function, the complexity of the model is constrained to avoid overfitting. 4. Data augmentation. Continuously training the model by generating new training data to increase the model's generalization ability. 5. Early stopping. Stopping the training when the performance on the validation set no longer improves can avoid overfitting. 6. Learning rate adjustment. Using techniques such as learning rate decay can achieve better convergence during training. 7. Multi-task learning. Letting the model solve multiple related tasks simultaneously can improve the model's generalization ability.

[0030] Model serialization refers to converting the model object running in memory into a binary sequence file and storing it in a persistent storage medium (such as a hard disk) in the form of a binary sequence file for flexible use of the deep learning model. In the embodiments of the present application, the binary sequence file can also be referred to as the serialization file of the deep learning model. The serialization file of the deep learning model includes information such as the model structure information of the deep learning model, the names of each network layer, the names and weight values of the weights in each network layer, etc.

[0031] To improve the efficiency of model inference, such as Figure 1aAs shown, during the inference phase, an inference acceleration engine is used to perform inference on the model. Similar to the training acceleration engine, the inference acceleration engine is a deep learning framework for model inference, such as torch+

[0032] xformers, TensorRT, etc. The inference acceleration engine uses compilation optimization, optimizes the hardware resources for model inference, compresses the model, etc., to accelerate the inference process of the deep learning model and reduce the model inference time. Among them, the inference acceleration engine and the training acceleration engine can be the same deep learning framework or different deep learning frameworks. For example, torch is a deep learning training and inference framework that supports hardware resources such as CPUs and GPUs. xformers is an acceleration library for the multi-head attention (MHA) algorithm and can be used for both model training and model inference. TensorRT is a tensor-oriented runtime acceleration engine mainly used for model inference. As Figure 1a shown, the model inference process using the inference acceleration engine involves at least steps such as model deserialization, computational graph conversion, compilation optimization, and task execution.

[0033] Among them, model deserialization refers to the process of deserializing the binary sequence file of the deep learning model in the persistent storage medium (such as a hard disk) into memory to obtain a model object that can be run, so that the deep learning model can perform deep learning tasks on its own. Among them, the model deserialization process can also be called the model loading process.

[0034] Computational graph conversion is the process of parsing the deep learning model into the form of a computational graph. A computational graph is a data structure used to represent the computational process of a deep learning model. The computational graph includes nodes and edges. Nodes represent each network layer in the deep learning model, and each network layer contains operations such as matrix multiplication and addition. Edges represent the data flow between the network layers represented by adjacent nodes. By traversing the structure of the deep learning model, the computations in each network layer of the deep learning model are converted into nodes in the computational graph, and the input, output, and connection relationships with other nodes of each node are recorded.

[0035] Compilation optimization refers to the process of optimizing the computational graph after its transformation, such as node fusion, pruning, vectorization, memory optimization, parallelization, etc., to reduce the computational load and memory occupancy and improve the running efficiency. Among them, node fusion means merging multiple nodes in the computational graph into one node, thereby reducing the computational and communication overheads. For example, multiple convolution operations can be merged into one convolution operation. Node fusion can also be called operator fusion or graph fusion; pruning means removing unnecessary nodes and edges in the computational graph to reduce the computational load, where the removed nodes and edges do not affect the model output results. Vectorization means converting loops and iterations in the computational graph into vectorized operations, and vectorized operations can utilize the parallel processing capabilities of hardware resources to improve the computational efficiency. Memory optimization means that by optimizing the memory access pattern and reducing the number of memory allocations, the running efficiency of the model can be improved. For example, techniques such as cache optimization and memory alignment can be used to reduce the memory access latency. Parallelization means that by parallelizing the operations in the computational graph, the computational capabilities of hardware resources such as multi-core CPUs or GPUs used for model inference can be fully utilized to improve the running speed of the model. The above-listed compilation optimization techniques can be selected and applied according to specific requirements and goals to achieve the desired performance optimization effect.

[0036] The task execution process refers to the process of using the optimized-compiled model to perform the inference calculation of the corresponding task and output the calculation result.

[0037] With the development of model applications, there have emerged some application scenarios that require frequent switching between different deep learning models. For example, there are multiple tasks, such as Task A, Task B, Task C, etc., and different tasks require different deep learning models respectively. The "different deep learning models" mentioned here include both different model structures and different weight information. When switching from Task A to Task B, the deep learning model used to execute Task A needs to be synchronously switched to the deep learning model used to execute Task B. In application scenarios involving frequent model switching, if Figure 1a the inference process shown is followed, after each deep learning model switch, processes such as deserialization, computational graph transformation, and compilation optimization need to be gone through again. Since optimization compilation is a time-consuming process, it will seriously affect the model deployment or switching efficiency. Especially in the case of frequent switching and large model sizes, the extreme phenomenon of "compiling for one hour and inferring for ten seconds" is likely to occur.

[0038] In view of the above problems, after long-term research and analysis, the inventors of this case have redefined the deep learning model in the application scenario and proposed the concept of "isomorphic model". An isomorphic model refers to a collective term for multiple models with the same model structure but different weight information. The isomorphic model contains multiple deep learning models with the same model structure but different weight information. For these deep learning models with the same model structure but different weight information, different training samples can be used to train the models during the training process, and then isomorphic models for performing different deep learning tasks can be obtained. That is to say, different models under the isomorphic model use different training samples during the training process, and different weight information is obtained by training with different training samples to adapt to different deep learning tasks. For example, deep learning models A and B have the same model structure and are both used for object recognition. Deep learning model A is specifically used for face recognition, and deep learning model B is specifically used for animal recognition. The weight information of the two deep learning models is different, and deep learning models A and B belong to isomorphic models. Or, deep learning models C and D have the same model structure and are both used for generating home decoration plans. Deep learning model C is specifically used for generating Nordic-style home decoration plans, and deep learning model D is specifically used for generating classical-style home decoration plans. The weight information of the two deep learning models is different, and deep learning models C and D belong to isomorphic models. Of course, the model structures and model functions of different isomorphic models are different, and the embodiments of this application do not limit the model structures and model functions of the isomorphic models. For the convenience of description and distinction, in the embodiments of this application, the deep learning models with the same model structure but different weight information in the isomorphic model are called original models.

[0039] Based on the definition of the isomorphic model, the embodiments of this application also provide a method for updating model weights. This method is used to implement model switching or deployment. That is, for an isomorphic model, a compiled and optimized model structure is obtained by performing a compilation and optimization process on any one of the original models. Then, during the model inference process that requires a certain original model, the compilation and optimization link can be directly skipped, and the weight information of the certain original model is injected into the compiled and optimized model structure according to the task requirements for weight update, so as to obtain a compiled and optimized model that can perform the corresponding task and achieve model deployment or switching. That is to say, for an isomorphic model, only one compilation and optimization process needs to be performed, achieving the purpose of avoiding compilation when switching isomorphic models and saving the time of compilation and optimization. Further, when enabling other original models in the isomorphic model to perform deep learning tasks, the weight information of other original models can be directly injected into the target model, so that the target model can perform deep learning tasks based on the injected weight information. While obtaining the performance benefits brought by compilation and optimization, by injecting the weight information into the target model to achieve the switching or deployment of the isomorphic model, the model deployment efficiency can also be improved.

[0040] As Figure 1b shown, the model inference stage provided in this embodiment includes: b1. A process of pre-compiling and optimizing the first original model to obtain a target model; and b2. A process of injecting the weight information of the second original model into the target model according to the deep learning task requirements to implement model switching or deployment. Among them, the first original model and the second original model belong to isomorphic models. The model structure and model function of the isomorphic model are not limited. In addition, the training process of each original model in the isomorphic model is not limited, and the model training method described in the foregoing embodiments may be used but is not limited thereto. In subsequent embodiments of the present application, on the premise that each original model in the isomorphic model has been trained, the focus is on how to quickly switch or deploy between isomorphic models according to the requirements of deep learning tasks.

[0041] The first original model has first weight information, and the second original model has second weight information. The first original model can be any original model in its isomorphic model. In addition, the second original model can be the same as the first original model or other models in the isomorphic model that are different from the first original model, which is not limited and is specifically determined according to the deep learning task to be executed. When the first original model is the same as the second original model, the first weight information is the same as the second weight information.

[0042] As Figure 1b shown, the implementation process of compiling and optimizing the first original model to obtain a target model is as follows: Step b11. Deserialize the serialized file of the first original model. The deserialization process can obtain the first original model and load the first original model into the memory; Step b12. Perform a computational graph conversion on the first original model loaded into the memory to obtain the computational graph of the first original model; Step b13. Compile and optimize the computational graph to obtain a compiled and optimized model. The compiled and optimized model includes the model structure information of the isomorphic model to which the first original model belongs and the first weight information corresponding to the first original model; Step b14. Serialize the model structure information in the compiled and optimized model to obtain a target model without weight information. Among them, the serialized file of the first original model is a file suitable for persistent storage, such as but not limited to a binary file. This file includes: the model structure information of the first original model, the names of each network layer included in the first original model, the types and names of the weights in each network layer, and the corresponding first weight information. For each original model under the isomorphic model, they have the same model structure information. That is to say, the model structure information of the first original model is the same as the model structure information of other original models belonging to the same isomorphic model, and can also be called the model structure information of its isomorphic model.

[0043] It should be noted here that for the compiled model obtained through the above compilation optimization, the purpose of serializing the compiled model is to convert the model structure information into a serialized file for persistent storage. In this way, when used later, instead of performing the model compilation optimization again, the model structure information can be obtained directly by deserializing the persistently stored serialized file, thereby improving the model inference efficiency. In this embodiment, the serialized file of the target model is also a file suitable for persistent storage, such as but not limited to a binary file. This file includes the aforementioned model structure information but does not include the first weight information. The model structure information mainly includes how many network layers the first original model or its isomorphic model includes, what types of computing nodes these network layers involve, and the connection relationships between these network layers or computing nodes. The "weight information" involved in the embodiments of the present application mainly refers to the weight values, rather than the weight types and names. Additionally, although only the model structure information is persistently stored through serialization, the compiled model is also simultaneously stored in the memory of the physical computing resource object that executes the model weight update method. This compiled model includes both the model structure information and the first weight information. Therefore, the compiled model can also be used for model inference. Among them, performing model inference based on the compiled model and serializing the compiled model are two independent operations, and there is no limitation on the order between them. For the sake of easy distinction and description, the physical computing resource object on the electronic device used to execute the deep learning model is referred to as the first physical computing resource object, and the physical computing resource object responsible for executing the model weight update method is referred to as the second physical computing resource object. The physical computing resource objects on the electronic device include the CPU, GPU, Data Processing Unit (DPU), etc.; among them, the first physical computing resource object and the second physical computing resource object can be the same, such as both being the CPU or both being the GPU; or, the first physical computing resource object and the second physical computing resource object are different. For example, the first physical computing resource object can be but not limited to the GPU, and the second physical computing resource object can be the CPU.

[0044] In an embodiment of the present application, in order to facilitate the successful injection of the second weight information into the target model, during the process of obtaining the target model by compiling and optimizing the first original model, the types and names of the weights required by the first original model are saved as the weight description information required by the target model. This weight description information is used to describe the types and names of the weights required by the target model. Based on this, various weight information of the original models can be injected into the target model to achieve the transmission of weight information. This weight description information can be used as a part of the target model and serialized into the serialization file of the target model together with the target model. In an embodiment of the present application, the serialization file of the target model not only includes the aforementioned model structure information, but also includes the weight description information here. The weight description information here mainly refers to the types and names of the weights in each network layer of the first original model or its isomorphic model, and does not include the weight information (i.e., the weight value).

[0045] In an alternative embodiment, during the computational graph conversion process, the types and names of the weights required by the first original model can be saved as the weight description information required by the target model. Specifically, the first original model includes multiple network layers, and each network layer has its own weight type and name. During the computational graph conversion process of the first original model, the computational graph conversion process can be performed on each network layer one by one, and during the computational graph conversion process, the weight types and names of each network layer are saved as the weight description information of the target model.

[0046] Among them, an implementation process of obtaining the computational graph corresponding to the first original model by performing computational graph conversion on the first original model is as follows: parsing the first original model to obtain each network layer included in the first original model, the operators involved in each network layer, the inputs and outputs of the operators in each network layer, and the weight types, names, and the weight information (i.e., the first weight information) corresponding to each weight name of each network layer, etc. Among them, the operators at least include addition, subtraction, multiplication, division, convolution, pooling, normalization, etc. Usually, one network layer corresponds to one node; further, according to the ordered connection relationship between each network layer, it is mapped to the directed edges between the nodes corresponding to each network layer, and the inputs and outputs between each network layer are mapped to the inputs and outputs between the nodes corresponding to each network layer; further, based on each node, the directed edge, and the input and output relationships between each node, the computational graph of the first original model is obtained.

[0047] Combined with the above-mentioned computational graph conversion process, for the target network layer in the current conversion, on the one hand, the target network layer can be converted into a target node in the computational graph, and on the other hand, the weight types and names required by the target network layer can be saved; the target network layer can be any network layer in the first original model. Further, according to the different construction methods of the target network layer, the methods of saving the weight types and names required by the target network layer will be different. If the construction method of the target network layer is the Plugin method, the types and names of the weights in the target network layer are saved into the target node in the form of parameter passing; if the construction method of the target network layer is the BuildIn method, a first data structure is created, and the identification information of the target network layer and the types and names of the weights in the target network layer are correspondingly stored in the first data structure.

[0048] Among them, the BuildIn method and the Plugin method are two methods for constructing each network layer in the deep learning model. The BuildIn method can be understood as that the construction methods between network layers are coupled. When constructing the network layer, the coupling relationship between network layers has been determined and is not easy to modify. The Plugin method can be understood as that the construction methods between network layers are decoupled, and the relationship between network layers can be established by plugging and unplugging. And the plugin network layer can be some custom network layers, and these network layers can contain some optimization operators. Correspondingly, the network layers in the deep learning model can be divided into BuildIn network layers and Plugin network layers. A model can only contain any one type of network layer, or can contain both types of network layers at the same time. Compared with the BuildIn network layer, the implementation cost of the Plugin network layer is lower. In this embodiment, the first original model can default to contain BuildIn network layers. In addition, whether the first original model contains Plugin network layers can be determined according to specific requirements. In the embodiments of the present application, the implementation method of the first data structure is not limited, and it can be an array, a list, a hash table, etc. In an optional embodiment, the first data structure is implemented as a composite hash table. For a detailed description of this implementation method, reference can be made to the subsequent embodiments.

[0049] As Figure 1b shown, the process of injecting the weight information of the second original model into the target model according to the deep learning task requirements to implement model switching or deployment is as follows: Step b21, when enabling the second original model to execute the deep learning task, obtain the second weight information corresponding to the second original model; Step b22, inject the second weight information into the target model; Step b23, use the target model injected with the second weight information to execute the deep learning task.

[0050] Further optionally, as Figure 1bAs shown in the figure, obtaining the second weight information corresponding to the second original model includes: Step b211, when loading the second original model isomorphic to the first original model, deserializing the serialization file of the second original model. Through deserialization, the second original model can be obtained and loaded into the memory; Step b212, obtaining the weight dictionary of the second original model, and the weight dictionary contains the second weight information of the second original model. Among them, the serialization file of the second original model is a file suitable for persistent storage. For example, it can be, but is not limited to, a binary file. This file includes: the model structure information of the second original model, the names of each network layer included in the second original model, the types and names of the weights in each network layer, and the corresponding second weight information.

[0051] Here, from the perspective of the information types included in the serialization file, the information types included in the serialization files of the first original model and the second original model are the same. However, the information type included in the serialization file of the target model is different from the information types in the serialization files of the first original model and the second original model. The biggest difference is that: the serialization file of the target model does not contain specific weight information.

[0052] Further optionally, the implementation manner of injecting the second weight information into the target model includes: injecting the second weight information into the target model based on the weight description information corresponding to the target model.

[0053] Further optionally, preprocess the second weight information to obtain the third weight information, and relocate the third weight information from the second memory to the first memory; according to the weight description information, read the third weight information from the first memory and inject it into the target model. Among them, the second target weight information is stored in the memory of the second physical computing resource object. However, it is the first physical computing resource object that is responsible for model operation. To facilitate injecting the second weight information into the target model, it is necessary to relocate the second weight information from the memory of the second physical computing resource object to the memory of the first physical computing resource object. In the embodiments of this application, the memory of the first physical computing resource object can be referred to as the first memory, and the memory of the second physical computing resource object can be referred to as the second memory. Further, before relocating the second weight information from the second memory to the first memory, the second weight information can also be preprocessed, such as operations like merging and renaming the weight information, and the preprocessed weight information is called the third weight information. Then, relocate the third weight information from the second memory to the first memory, and according to the weight description information, inject the third weight information in the second memory into the target model.

[0054] In some embodiments, reading the third weight information from the first memory and injecting it into the target model according to the weight description information includes: constructing a second data structure adapted to the target model, where the second data structure is used to store the weight names and the storage addresses of their corresponding weight information; adding the storage addresses of the third weight information in the first memory corresponding to the weight names to the second data structure according to the weight names; and reading the third weight information from the first memory and injecting it into the target model according to the weight description information and the weight names stored in the second data structure and the storage addresses of their corresponding weight information in the first memory. In the embodiments of the present application, the implementation manner of the second data structure is not limited. For example, it may be, but is not limited to, an array, a list, a hash table, etc. In an optional embodiment, the second data structure is implemented as a hash table including weight names and weight information storage addresses, and specific details can be found in the description of subsequent embodiments.

[0055] In some embodiments, the target model persists in the memory of the second physical computing resource object, simply referred to as resident in memory. In the case where the target model persists in the memory of the second physical computing resource object, the third weight information can be read from the first memory and injected into the target model according to the weight description information and the weight names stored in the second data structure and the storage addresses of their corresponding weight information in the first memory. Specifically, according to the weight description information and the weight names stored in the second data structure and the storage addresses of their corresponding weight information in the first memory, the third weight information corresponding to each network layer in the target model can be read layer by layer from the first memory in a preset reading order and written into the weight space of the corresponding network layer in the first memory of the target model; or, according to the weight description information and the weight names stored in the second data structure and the storage addresses of their corresponding weight information in the first memory, the third weight information corresponding to each network layer in the target model can be read from the first memory at one time and written into the weight space of the corresponding network layer in the first memory of the target model.

[0056] In some embodiments, according to the weight description information and the storage addresses of the weight names and their corresponding weight information stored in the second data structure in the first memory, the third weight information corresponding to each network layer in the target model is read layer by layer from the first memory and written into the weight space of the corresponding network layer in the first memory of the target model, including: if the target model is not in the memory of the second physical computing resource object, the serialized file corresponding to the target model is immediately loaded from the persistent storage medium, and the serialized file corresponding to the target model is deserialized to obtain the target model. The memory of the second physical computing resource object is the memory for executing the weight update method of this model, such as the CPU; during the deserialization process, for any network layer deserialized from the target model, according to the weight description information and the storage addresses of the weight names and their corresponding weight information stored in the second data structure in the first memory, the third weight information corresponding to any network layer is read from the first memory; a weight space is allocated for any network layer in the first memory, and the third weight information is written into the weight space allocated for any network layer.

[0057] Further optionally, during the compilation and optimization process of the first original model, the optimization strategy information used for the compilation and optimization of each network layer in the first original model can be saved. Correspondingly, writing the third weight information into the weight space of the corresponding network layer in the first memory of the target model includes: rearranging the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain the rearranged weight information; writing the rearranged weight information into the weight space of the corresponding network layer in the first memory of the target model, so as to facilitate injecting the rearranged weight information into the target model. The rearrangement of the third weight information is the same as the rearrangement operation performed on the first weight information during the compilation and optimization process, for example, including but not limited to: operations such as merging, deleting, and sorting of weight information, depending on the compilation and optimization strategy.

[0058] It should be noted here that the model weight update method provided in the above or below embodiments of this application can be executed by various inference acceleration engines. In order to implement the model weight update method provided in the embodiments of this application in the inference acceleration engine, as Figure 1c shown, the following functional modules can be added to the inference acceleration engine: a weight registration module 21, a weight acquisition module 22, and a weight loading module 23. Of course, the inference acceleration engine also includes a model loading module 11, a computation graph conversion module 12, a compilation optimization module 13, a serialization module 14, and a model inference module 15. With the cooperation of these modules, the user can be provided with the ability of seamless weight update without compilation, and the model switching or deployment can be quickly realized.

[0059] Among them, the model loading module 11 is used to deserialize the serialized file of the first original model to obtain the first original model in memory; the computation graph conversion module 12 is used to perform computation graph conversion on the first original model to obtain the computation graph corresponding to the first original model; the compilation and optimization module 13 is used to perform compilation and optimization processing on the computation graph to obtain a compiled model, and the compiled model includes the model structure information and the first weight information of the first original model; among them, the model structure information in the compiled model is used as the target model; the serialization module 14 is used to serialize the target model to obtain the serialized file corresponding to the target model; and persistently store the serialized file corresponding to the target model.

[0060] The weight registration module 21 is used to, during the process of the computation graph conversion module performing computation graph conversion on the first original model, obtain the types and names of the weights in each network layer of the first original model, and select an appropriate storage method to associatively store the weight types and names in each network layer with the corresponding network layer according to the construction method of each network layer. This process is also called the weight registration process. It should be noted that the registered weight types and names, as weight description information, will be serialized into the serialized file of the target model as part of the target model; these weight description information will be reread and used to inject the second weight information when it is necessary to switch to the second original model to execute the deep learning task, so as to achieve continuous tracking of the weight information.

[0061] The weight registration module 21 realizes the recording and identification of the weight description information, and can help each network layer record the weight types and names involved by itself. The weight acquisition module 22 is used to obtain the second weight information corresponding to the second original model when enabling the second original model isomorphic to the first original model to execute the deep learning task. Specifically, the weight acquisition module 22 is used to construct a second data structure adapted to the target model, and the second data structure is used to store the weight names and the storage addresses of the corresponding weight information; and provide a function interface for external modules to update the second data structure, so that external modules can update the weight names and the storage addresses of the corresponding weight information stored in the second data structure through this function interface. Among them, the second data structure can be a hash table containing the weight names and the weight data addresses, but is not limited thereto.

[0062] The above-mentioned external module mainly refers to the model inference module 15. The model inference module 15 is used to load the serialized file corresponding to the second original model from the persistent storage medium when enabling the second original model to perform a deep learning task; deserialize the serialized file corresponding to the second original model to obtain the second original model; and during the deserialization process, for any network layer in the deserialized second original model, obtain the second weight information corresponding to the any network layer, and write the name and its storage location of the second weight information corresponding to the any network layer into the second data structure by calling the function interface provided by the weight acquisition module, so as to facilitate the weight acquisition module to obtain the second weight information.

[0063] Further optionally, the model inference module 15 writes the name and its storage location of the second weight information corresponding to any network layer into the second data structure, including: preprocessing the second weight information to obtain the third weight information, and relocating the third weight information from the second memory to the first memory; calling the function interface provided by the weight acquisition module, and adding the storage address of the third weight information in the first memory to the second data structure according to the weight name.

[0064] Based on the weight registration module 21 and the weight acquisition module 22, the weight loading module 23 can enable each network layer in the target model to obtain the new weight information corresponding to itself during weight injection, so as to perform the corresponding deep learning task based on the injected new weight information. Specifically, the weight loading module 23 is used to inject the second weight information into the target model according to the weight description information, and use the target model injected with the second weight information to perform the deep learning task.

[0065] Further optionally, the weight loading module 23 reads the weight description information from the first data structure, reads the storage address of the third weight information in the first memory from the second data structure according to the weight type and name in the weight description information, reads the third weight information from the first memory according to the storage address of the third weight information in the first memory, and performs rearrangement operations such as merging, deleting, and sorting on the third weight information according to the pre-saved compilation optimization strategy, and injects the rearranged third weight information into the target model, and controls the target model injected with the rearranged third weight information to perform the deep learning task that should originally be performed by the second original model.

[0066] Further optionally, when the target model is not resident in memory, the weight loading module 23 is also used to deserialize the serialized file of the target model to obtain the target model, and then inject the second weight information into the target model by using the above operations, so as to use the target model injected with the second weight information to perform the deep learning task.

[0067] In the embodiments of the present application, compared with the inference process in the traditional model switching scenario, the model inference process of the embodiments of the present application is greatly simplified, mainly including model switching operation, deserialization operation, and weight update operation, omitting the compilation optimization process, which can greatly improve the model switching efficiency and inference efficiency. At the same time, the performance advantages brought by compilation optimization can still be retained. Especially in the scenario where the number of models is huge, the beneficial effects of the embodiments of the present application are particularly remarkable, which can greatly release the user's AI productivity and enhance the core competitiveness of various cloud computing instances that provide cloud computing services based on deep learning models, such as Elastic Compute Service (ECS) among the similar products of cloud providers.

[0068] Figure 2 It is a schematic flowchart of a weight update method provided by an exemplary embodiment of the present application. This method is applied to a model inference engine, such as Figure 2 as shown, this method includes:

[0069] 201. Pre-compile and optimize the first original model to obtain a target model without weight information, and save the type and name of the weights required by the first original model as weight description information adapted to the target model;

[0070] 202. When enabling the second original model isomorphic to the first original model to execute a deep learning task, obtain the second weight information corresponding to the second original model, and the first original model has the first weight information;

[0071] 203. According to the weight description information, inject the second weight information into the target model, so as to use the target model with the injected second weight information to execute the above deep learning task.

[0072] In this embodiment, the first original model and the second original model belong to isomorphic models. The embodiments of the present application do not limit the model structure of the isomorphic models to which the first original model and the second original model belong, nor do they limit the model functions of the isomorphic models. For example, both the first original model and the second original model are used for target recognition. The first original model is specifically used for face recognition, and the second original model is specifically used for animal recognition. Or, both the first original model and the second original model are used for generating home decoration plans. The first original model is specifically used for generating Nordic-style home decoration plans, and the second original model is specifically used for generating classical-style home decoration plans. The above embodiments are only for illustrative purposes and do not limit the technical solutions of the present application.

[0073] In this embodiment, any original model under the isomorphic model can be selected as the first original model, and the first original model is compiled and optimized to obtain a target model without weight information but only with model structure information. Since the target model is compiled from the first original model, the weight types and names required by the target model are the same as those required by the first original model. Therefore, the types and names of the weights required by the first original model can be saved as weight description information adapted to the target model, and this weight description information is used to describe which weight information the target model needs, providing conditions for injecting weight information into the target model subsequently.

[0074] It should be noted that the compilation and optimization of the first original model can be carried out when a deep learning task has been received, so that while obtaining the target model, the compiled and optimized model (referred to as the compiled model) can also be used to execute the deep learning task; or, it can also be carried out when no deep learning task has been received, with the main purpose of obtaining the target model.

[0075] Further optionally, before compiling and optimizing the first original model, it also includes: deserializing the serialized file of the first original model to obtain the first original model. Among them, the serialized file of the first original model is stored in a persistent storage medium. The process of deserializing the serialized file of the first original model is to load the serialized file of the first original model from the persistent storage medium into memory, and perform deserialization processing on the serialized file in memory to obtain a runnable model object, which is the first original model. Before loading the serialized file of the first original model from the persistent storage medium, it also includes the model training and serialization processes, and these processes can refer to the relevant descriptions of the above embodiments and will not be elaborated here.

[0076] In this embodiment, compiling and optimizing the first original model to obtain a target model without weight information includes: converting the computational graph of the first original model to obtain the computational graph corresponding to the first original model; performing compilation and optimization processing on the computational graph to obtain a compiled model, and the compiled model includes model structure information and first weight information; using the model structure information in the compiled model as the target model.

[0077] In some embodiments, the implementation process of obtaining the computational graph corresponding to the first original model by performing computational graph transformation on the first original model is as follows: Parse the first original model to obtain each network layer included in the first original model, each operator involved in each network layer, the inputs and outputs of the operators in each network layer, as well as the weight types, names of each network layer, and the weight information corresponding to each weight name (i.e., the first weight information), etc. Among them, the operators at least include addition, subtraction, multiplication, division, convolution, pooling, normalization, etc. Usually, one network layer corresponds to one node; further, according to the ordered connection relationship between each network layer, map it to the directed edges between the nodes corresponding to each network layer, and map the inputs and outputs between each network layer to the inputs and outputs between the nodes corresponding to each network layer; further, based on each node, directed edge, and the input and output relationships between each node, obtain the computational graph of the first original model.

[0078] In this embodiment, after obtaining the weight types and names in each network layer of the first original model, save the types and names of the weights required by the first original model as the weight description information adapted to the target model. It can be understood as the registration process of the weight description information of the first original model. According to this weight description information, the mapping relationship between the target model and the weight information can be dynamically updated during the weight injection process, so as to realize the transmission of the weight information in each network layer. Among them, the construction methods of the network layers of the first original model are different, and the corresponding registration processes of the weight description information are also different. The registration process of the weight description information can be implemented during the computational graph transformation process.

[0079] Specifically, during the computational graph transformation process, the computational graph transformation can be performed layer by layer for each network layer. For the target network layer in the current transformation, the target network layer is any network layer in the first original model. On the one hand, the target network layer is transformed into the target node in the computational graph; on the other hand, the construction method of the target network layer can be identified through the attribute information of the target network layer; if the construction method of the target network layer is the plug-in method, save the types and names of the weights in the target network layer to the target node in the form of parameter passing, specifically, implement the types and names of the weights in the target network layer as the attribute information of the target node; if the construction method of the target network layer is the built-in method, create the first data structure, and store the identification information of the target network layer and the types and names of the weights in the target network layer in the first data structure. The first data structure is associated with the target model and is used to store the weight types and weight names of each network layer required by the target model.

[0080] In the embodiments of the present application, the implementation method of the first data structure is not limited, and it can be an array, a list, a hash table, etc. For example Figure 3aAs shown, taking the first data structure as a composite hash table as an example, the composite hash table includes the correspondence between the network layer (Layer) and the weight list (List). The weight list includes tuples formed by the weight type (WeightsKind) and the weight name (WeightsName), simply represented as Layer->List[Tuple(WeightsKind, WeightsName)]. In Figure 3a Figure 3a , taking the first original model having three network layers as an example, the three network layers are respectively denoted as Convolution-1, Convolution-2, and Layemorm-1. Convolution-1 and Convolution-2 represent two convolutional layers, and Layemorm-1 represents a normalization layer. The convolutional layer Convolution-1 includes two weights, and the names of the two weights are respectively Conv.1.kernel and Conv.1.bias; Convolution-2 also includes two weights, and the names of the two weights are respectively Conv.2.kernel and Conv.2.bias; Layernorm-1 includes one weight, and the name of the weight is Layernorm.1.scale. Among them, the weight types and names in the convolutional layer Convolution-1, the convolutional layer Convolution-2, and the normalization layer Layernorm-1 are stored in the List in the form of tuples respectively, and the correspondence between each network layer and the List is established, so as to achieve the purpose of storing the weight types and weight names of each network layer respectively in the form of a composite hash table data structure. In Figure 3a Figure 3a , the weight information corresponding to each weight name is set to a default value, for example, a vector with a value of 0.

[0081] In some embodiments, after obtaining the computational graph of the first original model, the computational graph can be further compiled and optimized. Specifically, an optimization strategy adapted to the first original model can be used to optimize each node in the computational graph to obtain a compiled model. The compiled model includes a model structure and first weight information, and the first weight information is the weight value of the first weight of the first original model.

[0082] Furthermore, in order for the target model to be persistently used, the target model can be persistently stored. Since the running memory space is limited, the target model can be serialized to obtain a serialized file corresponding to the target model, and the serialized file corresponding to the target model can be persistently stored.

[0083] In this embodiment, when receiving a deep learning task to be executed by a second original model isomorphic to the first original model, the second original model isomorphic to the first original model is enabled to execute the deep learning task, and the second weight information corresponding to the second original model is obtained. Optionally, when enabling the second original model isomorphic to the first original model to execute the deep learning task, obtaining the second weight information corresponding to the second original model includes: when enabling the second original model to execute the deep learning task, loading the serialized file corresponding to the second original model from the persistent storage medium; deserializing the serialized file corresponding to the second original model to obtain the second original model; and during the deserialization process, for any network layer in the deserialized second original model, obtaining the second weight information corresponding to any network layer. Among them, deserializing the serialized file is to load the serialized file from the persistent storage medium into the memory, and deserializing the serialized file in the memory to obtain a runnable model object, that is, storing the second original model in the form of a model object in the memory. It should be noted that the storage and injection of the model weight information are both carried out in the memory of the second physical computing resource object that executes the weight update method, and the second physical computing resource object is preferably a CPU.

[0084] Furthermore, according to the weight description information, the second weight information can be injected into the target model to use the target model with the injected second weight information to execute the deep learning task. Among them, the target model is a model compiled and optimized based on the first original model and has high performance. For the second original model isomorphic to the first original model, when a deep learning task needs to be executed, there is no need to compile and optimize the second original model. Just inject the second weight information into the target model. From the perspective of the second original model, while maintaining the performance advantages brought by compilation and optimization, the model deployment efficiency is also improved.

[0085] In some embodiments, injecting the second weight information into the target model according to the weight description information includes: preprocessing the second weight information to obtain third weight information, and relocating the third weight information from the second memory to the first memory; according to the weight description information, reading the third weight information from the first memory and injecting it into the target model. Among them, the first memory is the memory of the first physical computing resource object (such as a GPU) that runs the target model, and the second weight information is stored in the memory (i.e., the second memory) of the second physical computing resource object (such as a CPU) that executes the model weight update method. Among them, the preprocessing can be operations such as merging and renaming some weights.

[0086] In some embodiments, reading the third weight information from the first memory and injecting it into the target model according to the weight description information includes: constructing a second data structure adapted to the target model, where the second data structure is used to store the weight names and the storage addresses of their corresponding weight information; adding the storage addresses of the third weight information in the first memory corresponding to the weight names to the second data structure according to the weight names; and reading the third weight information from the first memory and injecting it into the target model according to the weight description information and the second data structure.

[0087] In some embodiments, the target model persists in the memory of the second physical computing resource object, which is abbreviated as persistent in memory. When the target model persists in the memory of the second physical computing resource object, the third weight information can be read from the first memory and injected into the target model according to the weight description information and the storage addresses of the weight names and their corresponding weight information stored in the second data structure. Specifically, according to the weight description information and the storage addresses of the weight names and their corresponding weight information stored in the second data structure in the first memory, the third weight information corresponding to each network layer in the target model can be read layer by layer from the first memory in a preset reading order and written into the weight space of the corresponding network layer in the first memory of the target model; or, according to the weight description information and the storage addresses of the weight names and their corresponding weight information stored in the second data structure in the first memory, the third weight information corresponding to each network layer in the target model can be read from the first memory at one time and written into the weight space of the corresponding network layer in the first memory of the target model.

[0088] In some embodiments, reading the third weight information corresponding to each network layer in the target model layer by layer from the first memory and writing it into the weight space of the corresponding network layer in the first memory of the target model according to the weight description information and the storage addresses of the weight names and their corresponding weight information stored in the second data structure includes: if the target model is not in the memory of the second physical computing resource object, immediately loading the serialized file corresponding to the target model from the persistent storage medium and deserializing the serialized file corresponding to the target model to obtain the target model. The memory of the second physical computing resource object is the memory for executing the weight update method of this model, such as the CPU; during the deserialization process, for any network layer deserialized from the target model, reading the third weight information corresponding to any network layer from the first memory according to the weight description information and the storage addresses of the weight names and their corresponding weight information stored in the second data structure; allocating a weight space for any network layer in the first memory and writing the third weight information into the weight space allocated for any network layer.

[0089] In this embodiment, based on the first data structure and the second data structure, a new weight corresponding to itself can be obtained during the deserialization of each network layer. However, the new weights need to be rearranged. Since different optimization strategies are used during the compilation optimization of each network layer, the corresponding weight arrangements are also different. Therefore, weight rearrangement is also a major difficulty in the weight update process. To address this issue, the weight rearrangement process can be separated into independent function interfaces. Each network layer will record the rearrangement algorithm it uses. After obtaining the pointer to the new weight data, it will call its corresponding rearrangement algorithm to achieve the rearrangement and loading of the weights.

[0090] Based on this, during the compilation optimization of the first original model, the optimization strategy information used for the compilation optimization of each network layer in the first original model can be saved. Correspondingly, writing the third weight information into the weight space of the corresponding network layer in the first memory of the target model includes: rearranging the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain the rearranged weight information; writing the rearranged weight information into the weight space of the corresponding network layer in the first memory of the target model.

[0091] It should be noted that the compilation optimization of the first original model is essentially the compilation optimization of the computational graph corresponding to the first original model. Then, the optimization strategy information used for the compilation optimization of each network layer in the first original model is the optimization strategy information used for the compilation optimization of each node in the computational graph.

[0092] For the sake of easy understanding, the following combines Figure 3a and Figure 3b Taking the first original model as the M model and the second original model as the N model as an example, the weight information update process is described in detail.

[0093] Suppose the M model contains two BuildIn convolutional layers and one Plugin normalization layer. First, the user constructs the M model network, with three network layers being Convolution-1, Convolution-2, and Layernorm-1 respectively, and weight registration is performed on each of the three network layers during the computational graph conversion of the M model. Among them, for the network layer Layernorm-1, the weight name is set to Layernorm.1.scale through parameter passing; while for the Convolution-1 and Convolution-2 nodes, they are automatically associated with the corresponding network layers in the form of a composite hash table, and the weight names of each network layer are Conv.1.kernel, Conv.1.bias, Conv.2.kernel, and Conv.2.bias respectively. After the model is compiled and optimized, the weight types and names of each network layer are serialized into the model's serialization file along with the target model obtained from the compilation optimization. In Figure 3a , taking the nodes in the computational graph as examples, the three network layers Convolution-1, Convolution-2, and Layernorm-1 are illustrated.

[0094] Furthermore, when it is necessary to switch from the M model to the N model in the isomorphic model, the N model also contains two BuildIn convolutional layers and one Plugin normalization layer, and each network layer has its own weight information. The storage addresses of the weight information of each network layer in the N model are set in the form of key-value pairs to the Figure 3b shown hash table. The new storage addresses of the 5 weight information are 0x0000010, 0x0000020, 0x0000030, 0x0000040, and 0x0000050 respectively. Then, a deserialization operation is performed on the target model obtained from the compilation optimization. During the deserialization process, each network layer will look for the corresponding storage address in the Figure 3b shown hash table according to its associated weight name. For example, the Convolution-1 node will rearrange and update the weight information at the 0x0000010 address into the Conv.1.kernel weight, and rearrange and update the weight information at the 0x0000020 address into the Conv.2.bias weight. And so on, until all nodes are deserialized, and at this time, the weight update work is completed synchronously.

[0095] Throughout the process, what the user actually experiences is a model deserialization process. Such a process does not change the user's usage habits and is relatively imperceptible. At the same time, this weight update mechanism breaks the limitations in fusion, enabling the user to enjoy better performance acceleration while implementing weight updates.

[0096] Figure 4 The structural schematic diagram of the model weight update device provided by an exemplary embodiment of the present application. As Figure 4 shown, the device includes:

[0097] A compilation optimization module 41, configured to pre-compile and optimize a first original model to obtain a target model without weight information, and save the type and name of the weights required by the first original model as weight description information adapted to the target model;

[0098] An acquisition module 42, configured to obtain second weight information corresponding to the second original model when enabling a second original model isomorphic to the first original model to execute a deep learning task, where the first original model has first weight information;

[0099] An injection module 43, configured to inject the second weight information into the target model according to the weight description information, so as to execute the deep learning task by using the target model injected with the second weight information.

[0100] In this embodiment, when the compilation optimization module 41 is used to pre-compile and optimize a first original model to obtain a target model without weight information, it is specifically configured to: perform a computational graph conversion on the first original model to obtain a computational graph corresponding to the first original model; perform a compilation optimization process on the computational graph to obtain a compiled model, where the compiled model includes model structure information and first weight information; use the model structure information in the compiled model as the target model.

[0101] Further optionally, it further includes: a processing module and a storage module, where the processing module is configured to perform serialization processing on the target model to obtain a serialized file corresponding to the target model; the storage module is configured to persistently store the serialized file corresponding to the target model.

[0102] In this embodiment, when the compilation optimization module 41 is used to save the type and name of the weights required by the first original model as weight description information adapted to the target model, it is specifically configured to: during the computational graph conversion process, determine a target network layer in the current conversion, where the target network layer is any network layer in the first original model, and the target network layer is converted into a target node in the computational graph; if the construction method of the target network layer is the plug-in method, save the type and name of the weights in the target network layer to the target node in a parameter passing manner; if the construction method of the target network layer is the built-in method, create a first data structure, and correspondingly store the identification information of the target network layer and the type and name of the weights in the target network layer in the first data structure.

[0103] In this embodiment, when the obtaining module 42 is used to obtain the second weight information corresponding to the second original model when enabling the second original model isomorphic to the first original model to execute a deep learning task, it is specifically used for: when enabling the second original model to execute a deep learning task, loading the serialized file corresponding to the second original model from the persistent storage medium; deserializing the serialized file corresponding to the second original model to obtain the second original model; and during the deserialization process, for any network layer in the deserialized second original model, obtaining the second weight information corresponding to the any network layer.

[0104] In this embodiment, when the injection module 43 is used to inject the second weight information into the target model according to the weight description information, it is specifically used for: preprocessing the second weight information to obtain third weight information, and relocating the third weight information from the second memory to the first memory, where the first memory is the memory of the first physical computing resource object running the target model, and the second memory is the memory of the second physical computing resource object running the model weight update method; reading the third weight information from the first memory according to the weight description information and injecting it into the target model.

[0105] Optionally, when the injection module 43 is used to read the third weight information from the first memory and inject it into the target model according to the weight description information, it is specifically used for: constructing a second data structure adapted to the target model, where the second data structure is used to store the weight name and the storage address of its corresponding weight information; adding the storage address of the third weight information in the first memory to the second data structure according to the weight name; reading the third weight information from the first memory according to the weight description information and the second data structure and injecting it into the target model.

[0106] Optionally, when the injection module 43 is used to read the third weight information from the first memory and inject it into the target model according to the weight description information and the second data structure, it is specifically used for: layer by layer reading the third weight information corresponding to each network layer in the target model from the first memory according to the weight description information and the second data structure and writing it into the weight space of the corresponding network layer in the first memory of the target model; or reading the third weight information corresponding to each network layer in the target model from the first memory at one time according to the weight description information and the second data structure and writing it into the weight space allocated to the corresponding network layer in the first memory of the target model.

[0107] Optionally, when the injection module 43 is used to layer-by-layer read the third weight information corresponding to each network layer in the target model from the first memory according to the weight description information and the second data structure and write it into the weight space allocated to the corresponding network layer in the first memory for the target model, it is specifically used for: if the target model is not in the memory of the second physical computing resource object, loading the serialized file corresponding to the target model from the persistent storage medium and deserializing the serialized file corresponding to the target model to obtain the target model; during the deserialization process, for any network layer deserialized from the target model, reading the third weight information corresponding to the any network layer from the first memory according to the weight description information and the second data structure; allocating a weight space for the any network layer in the first memory, and writing the third weight information into the weight space allocated for the any network layer.

[0108] Further optionally, the saving module is further used to save the optimization strategy information used for compiling and optimizing each network layer in the first original model during the process of compiling and optimizing the first original model. Correspondingly, when the injection module 43 is used to write the third weight information into the weight space of the corresponding network layer in the first memory for the target model, it is specifically used for: rearranging the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain the rearranged weight information; writing the rearranged weight information into the weight space of the corresponding network layer in the first memory for the target model.

[0109] The detailed implementation manners and beneficial effects of each module in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.

[0110] Figure 5 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. As Figure 5 shown, a memory 50a and a processor 50b; a computer program is stored in the memory 50a, and the processor 50b is coupled to the memory 50a for executing the computer program to be used to implement the following steps:

[0111] Pre-compile and optimize the first original model in advance to obtain a target model without weight information, and save the type and name of the weights required by the first original model as the weight description information adapted to the target model; when enabling the second original model isomorphic to the first original model to execute a deep learning task, obtain the second weight information corresponding to the second original model, and the first original model has first weight information; according to the weight description information, inject the second weight information into the target model to execute the deep learning task by using the target model injected with the second weight information.

[0112] In this embodiment, when the processor 50b is used to pre-compile and optimize the first original model to obtain a target model without weight information, it is specifically used for: performing a computation graph transformation on the first original model to obtain a computation graph corresponding to the first original model; performing a compilation optimization process on the computation graph to obtain a compiled model, where the compiled model includes model structure information and first weight information; and using the model structure information in the compiled model as the target model.

[0113] Further optionally, the processor 50b is also used to perform a serialization process on the target model to obtain a serialized file corresponding to the target model; and persistently store the serialized file corresponding to the target model.

[0114] In this embodiment, when the processor 50b is used to save the type and name of the weights required for the first original model as weight description information adapted to the target model, it is specifically used for: during the computation graph transformation process, determining a target network layer in the current transformation, where the target network layer is any network layer in the first original model, and the target network layer is transformed into a target node in the computation graph; if the construction method of the target network layer is the plug-in method, saving the type and name of the weights in the target network layer to the target node in the form of parameter passing; if the construction method of the target network layer is the built-in method, creating a first data structure, and correspondingly storing the identification information of the target network layer and the type and name of the weights in the target network layer into the first data structure.

[0115] In this embodiment, when the processor 50b is used to obtain the second weight information corresponding to the second original model when enabling the second original model isomorphic to the first original model to execute a deep learning task, it is specifically used for: when enabling the second original model to execute a deep learning task, loading the serialized file corresponding to the second original model from the persistent storage medium; deserializing the serialized file corresponding to the second original model to obtain the second original model; and during the deserialization process, for any network layer in the deserialized second original model, obtaining the second weight information corresponding to the any network layer.

[0116] In this embodiment, when the processor 50b is used to inject the second weight information into the target model according to the weight description information, it is specifically configured to: preprocess the second weight information to obtain third weight information, and relocate the third weight information from the second memory to the first memory, where the first memory is the memory of the first physical computing resource object that runs the target model; the second memory is the memory of the second physical computing resource object that runs the model weight update method; read the third weight information from the first memory and inject it into the target model according to the weight description information.

[0117] Optionally, when the processor 50b is used to read the third weight information from the first memory and inject it into the target model according to the weight description information, it is specifically configured to: construct a second data structure adapted to the target model, where the second data structure is used to store the weight name and the storage address of its corresponding weight information; add the storage address of the third weight information in the first memory to the second data structure according to the weight name; read the third weight information from the first memory and inject it into the target model according to the weight description information and the second data structure.

[0118] Optionally, when the processor 50b is used to read the third weight information from the first memory and inject it into the target model according to the weight description information and the second data structure, it is specifically configured to: layer by layer read the third weight information corresponding to each network layer in the target model from the first memory according to the weight description information and the second data structure and write it into the weight space allocated to the corresponding network layer in the first memory of the target model; or read the third weight information corresponding to each network layer in the target model from the first memory at one time according to the weight description information and the second data structure and write it into the weight space allocated to the corresponding network layer in the first memory of the target model.

[0119] Optionally, when the processor 50b is used to layer - by - layer read the third weight information corresponding to each network layer in the target model from the first memory according to the weight description information and the second data structure and write it into the weight space allocated to the corresponding network layer in the first memory for the target model, it is specifically used for: if the target model is not in the memory of the second physical computing resource object, loading the serialized file corresponding to the target model from the persistent storage medium and deserializing the serialized file corresponding to the target model to obtain the target model; during the deserialization process, for any network layer deserialized from the target model, reading the third weight information corresponding to the any network layer from the first memory according to the weight description information and the second data structure; allocating a weight space for the any network layer in the first memory and writing the third weight information into the weight space allocated for the any network layer.

[0120] Further optionally, the processor 50b is further used to save the optimization strategy information used for compiling and optimizing each network layer in the first original model during the process of compiling and optimizing the first original model; correspondingly, when the processor 50b is used to write the third weight information into the weight space of the corresponding network layer in the first memory for the target model, it is specifically used for: rearranging the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain the rearranged weight information; writing the rearranged weight information into the weight space of the corresponding network layer in the first memory for the target model.

[0121] In addition, as Figure 5 shown, the electronic device further includes: other components such as a communication component 50c, a display 50d, a power supply component 50e, and an audio component 50f. Figure 5 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 5 the components shown.

[0122] Regarding the detailed implementation manners and beneficial effects of each module in the method of this embodiment, they have been described in detail in the foregoing embodiments, and will not be elaborated here.

[0123] Correspondingly, an exemplary embodiment of the present application further provides a computer - readable storage medium storing computer programs / instructions, which when executed by a processor, cause the processor to be able to implement the steps in the above - mentioned method embodiments.

[0124] Correspondingly, an exemplary embodiment of the present application further provides a computer program product, the computer program product includes computer programs / instructions, which when executed by a processor, cause the processor to be able to implement the steps in the above - mentioned method embodiments.

[0125] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0126] The above-mentioned communication component is configured to facilitate communication, in a wired or wireless manner, between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G, and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0127] The above-mentioned display includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.

[0128] The above power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0129] The above audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in a memory or sent via a communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0130] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, Compact Disc Read-Only Memory (CD-ROM), optical memory, etc.) that contain computer-usable program code.

[0131] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0132] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, thereby providing steps for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0134] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and memory.

[0135] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (Random Access Memory, RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.

[0136] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change random access memory (Phase-change Random Access Memory, PRAM), static random access memory (SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (Digital Video Disc, DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0137] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0138] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for updating model weights, characterized in that, Applied to a model inference engine, the method includes: Pre-compiling and optimizing a first original model to obtain a target model without weight information, and saving the types and names of the weights required by the first original model as weight description information adapted to the target model; When enabling a second original model isomorphic to the first original model to execute a deep learning task, obtaining second weight information corresponding to the second original model, where the first original model has first weight information; According to the weight description information, injecting the second weight information into the target model to execute the deep learning task by using the target model injected with the second weight information.

2. The method according to claim 1, wherein Pre-compiling and optimizing a first original model to obtain a target model without weight information includes: Performing a computational graph transformation on the first original model to obtain a computational graph corresponding to the first original model; Performing a compilation and optimization process on the computational graph to obtain a compiled model, where the compiled model includes model structure information and first weight information; Using the model structure information in the compiled model as the target model.

3. The method according to claim 2, characterized in that It also includes: Performing serialization processing on the target model to obtain a serialized file corresponding to the target model; Persistently storing the serialized file corresponding to the target model.

4. The method according to claim 2, wherein Saving the types and names of the weights required by the first original model as weight description information adapted to the target model includes: During the computational graph transformation process, determining a target network layer in the current transformation, where the target network layer is any network layer in the first original model, and the target network layer is transformed into a target node in the computational graph; If the construction method of the target network layer is a plugin method, saving the types and names of the weights in the target network layer to the target node in a parameter passing manner; If the construction method of the target network layer is a built-in method, creating a first data structure and storing the identification information of the target network layer and the types and names of the weights in the target network layer in the first data structure in a corresponding manner.

5. The method according to claim 1, wherein When enabling a second original model isomorphic to the first original model to execute a deep learning task, obtaining second weight information corresponding to the second original model includes: When enabling the second original model to execute a deep learning task, loading the serialized file corresponding to the second original model from a persistent storage medium; Deserializing the serialized file corresponding to the second original model to obtain the second original model; and During the deserialization process, for any network layer in the deserialized second original model, obtaining the second weight information corresponding to the any network layer.

6. The method according to any one of claims 1-5, characterized in that, According to the weight description information, injecting the second weight information into the target model includes: Performing preprocessing on the second weight information to obtain third weight information, and relocating the third weight information from a second memory to a first memory; According to the weight description information, reading the third weight information from the first memory and injecting it into the target model. Among them, the first memory is the memory of the first physical computing resource object that runs the target model, and the second memory is the memory of the second physical computing resource object for executing the weight update method.

7. The method according to claim 6, wherein Reading the third weight information from the first memory according to the weight description information and injecting it into the target model includes: Constructing a second data structure adapted to the target model, where the second data structure is used to store the weight name and the storage address of its corresponding weight information; Adding the storage address of the third weight information in the first memory to the second data structure according to the weight name; Reading the third weight information from the first memory according to the weight description information and the second data structure and injecting it into the target model.

8. The method according to claim 7, wherein Reading the third weight information from the first memory according to the weight description information and the second data structure and injecting it into the target model includes: Reading the third weight information corresponding to each network layer in the target model layer by layer from the first memory according to the weight description information and the second data structure and writing it into the weight space allocated to the corresponding network layer in the first memory of the target model; Or Reading the third weight information corresponding to each network layer in the target model at one time from the first memory according to the weight description information and the second data structure and writing it into the weight space allocated to the corresponding network layer in the first memory of the target model.

9. The method according to claim 8, characterized in that Reading the third weight information corresponding to each network layer in the target model layer by layer from the first memory according to the weight description information and the second data structure and writing it into the weight space allocated to the corresponding network layer in the first memory of the target model includes: If the target model is not in the second memory, loading the serialized file corresponding to the target model from the persistent storage medium and deserializing the serialized file corresponding to the target model to obtain the target model; During the deserialization process, for any network layer deserialized from the target model, reading the third weight information corresponding to the any network layer from the first memory according to the weight description information and the second data structure; Allocating a weight space for the any network layer in the first memory and writing the third weight information into the weight space allocated for the any network layer.

10. The method according to claim 8, characterized in that It further includes: During the compilation and optimization process of the first original model, saving the optimization strategy information used for compiling and optimizing each network layer in the first original model; Correspondingly, writing the third weight information into the weight space of the corresponding network layer in the first memory of the target model includes: Rearranging the third weight information corresponding to each network layer according to the optimization strategy information used by each network layer to obtain the rearranged weight information; Writing the rearranged weight information into the weight space of the corresponding network layer in the first memory of the target model.

11. An electronic device, characterized in that, It includes: A memory and a processor; A computer program is stored in the memory, and the processor is coupled to the memory and configured to execute the computer program to implement the steps in the method according to any one of claims 1-10.

12. A computer-readable storage medium storing computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the processor is caused to be able to implement the steps in the method according to any one of claims 1-10.

13. A computer program product, characterized in that, The computer program product includes computer program / instructions which, when executed by the processor, cause the processor to be able to implement the steps in the method according to any one of claims 1-10.

Citation Information

Cited By

  • Memory management system and method, compiling method, equipment, medium and product

    CN121255475A