Neural network model optimization method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202411837803.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-09
AI Technical Summary
The efficiency of switching of neural network models in different domains, and the large model scale leads to increased hardware requirements of computing devices and reduced output efficiency and performance.
By compressing the neural network model, the model scale is reduced, and the model conversion device is used to convert the model based on the data of the source domain and the target domain to obtain the target domain weight matrix adapted to the target domain, thereby quickly performing domain adaptation.
It improves the adaptability and output efficiency of neural network models, reduces the cost of model conversion, and avoids the time and computational cost of retraining the model.
Smart Images

Figure CN119962576A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine learning, and in particular to a method and device for optimizing a neural network model, an electronic device, and a storage medium. Background Art
[0002] With the continuous development of machine learning technology and the increasing complexity of actual scenarios for machine learning applications, the structure of neural network models has become more and more complex, and the number of network layers has continued to increase, resulting in increasing hardware requirements for computing devices running neural network models. This not only increases the hardware requirements for computing devices required to run neural network models, but may also lead to a decrease in the output efficiency and performance of neural network models.
[0003] In addition, the neural network model trained on the source domain data is difficult to adapt to the processing tasks in the target domain. In order to ensure that the neural network model can work effectively in different domains, it is necessary to perform domain adaptation on the neural network model, but this process is usually complicated and time-consuming. The above factors together lead to the problem of low efficiency in switching between different domains of the neural network model. Summary of the invention
[0004] In view of this, the present application provides a method and device for optimizing a neural network model, an electronic device and a storage medium, which can not only reduce the model scale of the neural network model, but also enable the neural network model to quickly perform domain adaptation, thereby improving the adaptability and output efficiency of the neural network model.
[0005] In a first aspect, an embodiment of the present application provides a method for optimizing a neural network model, comprising: compressing a first neural network model to obtain a second neural network model; obtaining first historical data in a source domain and second historical data in a target domain; training the second neural network model through the first historical data to obtain the second neural network model in the source domain; obtaining a target weight matrix of the target domain according to the second neural network model in the source domain and the second historical data through a model conversion device; the model conversion device performs model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain a target neural network model in the target domain.
[0006] According to the above method of the embodiment of the present application, the following additional technical features may also be provided:
[0007] In one possible implementation, the compression processing of the first neural network model to obtain the second neural network model includes: determining multiple residual units in the first neural network model, each residual unit including an input, an output and a shortcut connection, and the shortcut connection is used to connect the output of one residual unit with the input of an adjacent residual unit; according to the Adams method, determining the weight coefficient weighted by the shortcut connection between adjacent residual units; reconstructing the first neural network model according to the output of the previous residual unit, the input of the current residual unit and the weight coefficient; and training the reconstructed first neural network model to obtain the second neural network model.
[0008] In one possible implementation, reconstructing the first neural network model based on the output of the previous residual unit, the input of the current residual unit and the weight coefficient includes: multiplying the output of the previous residual unit by the corresponding weight coefficient and adding the result to the input of the current residual unit to obtain the updated input of the current residual unit; traversing each residual unit and obtaining the corresponding updated input of the residual unit; and reconstructing the first neural network model based on each updated input and the structure of the corresponding residual unit.
[0009] In a possible implementation, the method of obtaining the target weight matrix of the target domain based on the second neural network model and the second historical data in the source domain through a model conversion device includes: using the weight matrix of the second neural network model in the source domain as the initial weight matrix of the target neural network model in the target domain; updating the initial weight matrix based on the second historical data and the LM least squares algorithm to obtain the target weight matrix of the target domain.
[0010] In a possible implementation, the initial weights are updated according to the second historical data and the LM least squares algorithm to obtain the target weight matrix of the target domain, including: training the initial weight matrix using the second historical data, and updating the trained weight matrix through the LM least squares algorithm; repeatedly training the updated weight matrix using the second historical data, and updating the weight matrix again through the LM least squares algorithm until a preset convergence condition is reached, thereby obtaining the target weight matrix of the target domain.
[0011] In a possible implementation, the model conversion device performs model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain the target neural network model in the target domain, including: loading the target weight matrix of the target domain into the second neural network model in the source domain; in the target domain, adjusting the network structure of the second neural network model according to the target weight matrix; and testing and verifying the second neural network model in the target domain after updating the network structure to obtain the target neural network model.
[0012] In a possible implementation, the second neural network model in the target domain after the updated network structure is tested and verified to obtain the target neural network model, including: dividing the second historical data into a training set and a test set; using the training set to train the second neural network model to obtain a third neural network model in the target domain; inputting the test set into the third neural network model in the target domain to obtain a first output result; inputting the test set into the second neural network model in the target domain to obtain a second output result; if the difference between the first output result and the second output result is within a preset error range, it is determined that the target neural network model in the target domain is obtained.
[0013] In a second aspect, an embodiment of the present application provides an optimization device for a neural network model, comprising: a compression unit, used to compress a first neural network model to obtain a second neural network model; an acquisition unit, used to obtain first historical data in a source domain, and second historical data in a target domain; a training unit, used to train the second neural network model through the first historical data to obtain the second neural network model in the source domain; a processing unit, used to obtain a target weight matrix of the target domain according to the second neural network model in the source domain and the second historical data through a model conversion device; the processing unit is also used for the model conversion device to perform model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain the target neural network model in the target domain.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method of the first aspect are implemented.
[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method of the first aspect are implemented.
[0016] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method of the first aspect.
[0017] Optionally, as an implementation method, the chip may also include a memory, in which instructions are stored, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the neural network model optimization method described in the first aspect.
[0018] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method of the first aspect.
[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0021] Figure 1 A schematic diagram showing a flow chart of a method for optimizing a neural network model provided in an embodiment of the present application is shown;
[0022] Figure 2 A schematic diagram of shortcut paths in a first neural network model and a second neural network model in an optimization method of a neural network model provided in an embodiment of the present application is shown;
[0023] Figure 3 A schematic diagram of the structure of an optimization device for a neural network model provided in an embodiment of the present application is shown;
[0024] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0026] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0027] The following is a detailed description of the optimization method of the neural network model provided in the embodiment of the present application through specific embodiments and their application scenarios in conjunction with the accompanying drawings. The following embodiments and features in the embodiments may be combined with each other unless there is a conflict.
[0028] The present application embodiment provides a method for optimizing a neural network model, such as Figure 1 As shown, the method includes steps 101 to 105:
[0029] Step 101: compressing the first neural network model to obtain a second neural network model;
[0030] The first neural network model may be a neural network model with an attention mechanism, such as a neural network model corresponding to task scenarios such as target detection, target classification, image classification, machine translation, speech recognition, and text recognition.
[0031] Specifically, the neural network model provided in the embodiment of the present application can be applied to the process industry. The process industry is an industry with attributes such as continuous production, material flow, and process control, such as: power, smelting, chemical, pharmaceutical and other industrial industries. The first neural network model can be used to analyze and predict certain indicators in the process industry.
[0032] The first neural network model may be a residual network model (Residual Network, ResNet model). The residual network model has a residual connection structure, which enables the residual network model to more easily propagate gradients during training, thereby avoiding the gradient vanishing problem of the neural network model. In addition, this residual connection structure can more easily identify and retain important feature information, thereby making the residual network model more conducive to compression processing.
[0033] The second neural network model is a model obtained by compressing the first neural network model. The type of the second neural network model may remain unchanged, but the model scale of the second neural network model is smaller than that of the first neural network model.
[0034] Specifically, the number of layers of the second neural network model may be less than the number of layers of the first neural network model. Therefore, the second neural network model has a better data processing speed, which helps to improve the model optimization efficiency of the neural network model optimization method provided in the embodiment of the present application.
[0035] Step 102: Acquire first historical data in the source domain, and acquire second historical data in the target domain;
[0036] The source domain and the target domain include training samples respectively. The source domain may include training samples obtained in one environment, i.e., first historical data. The target domain includes training samples similar to the training samples in the source domain, and the target domain may include training samples obtained in another similar environment, i.e., second historical data. The second historical data may also be used to further adjust or optimize the model based on the second neural network model trained in the source domain, so that the second neural network model trained in the source domain environment is more easily adapted to the target domain.
[0037] In actual application scenarios, the embodiment of the present application uses the optimization method of the neural network model provided in the mineral processing industry for exemplary description. The first historical data may be the historical production data of the first filter press in the mineral processing plant, and the second historical data may be the historical production data of the second filter press in the same mineral processing plant (to ensure the consistency of production process, process control, etc.).
[0038] Optionally, in order to improve data quality, the historical production data of the first filter press and the second filter press may be preprocessed by eliminating redundancy, removing abnormal data, interpolating default data, etc., to obtain preprocessed first historical data and second historical data.
[0039] Step 103: training a second neural network model using the first historical data to obtain a second neural network model in the source domain;
[0040] The first historical data may include a training set and a test set, the training set may be used for training the second neural network model, and the test set may be used for verifying the output performance of the second neural network model.
[0041] Furthermore, the input parameters and output parameters of the second neural network model in the source domain can be determined according to the production characteristics of the first filter press, the focus indicators and other factors.
[0042] Exemplarily, when training the second neural network model in the source domain, the historical data of the first filter press such as the frame filling time, diaphragm time, and drying time can be used as input parameters, and the historical data of the first filter press about the machine hours (the amount of production tasks completed by the equipment in unit time) can be used as output parameters.
[0043] It can be understood that by training the second neural network model with the first historical data collected in the source domain, a neural network model adapted to the source domain, i.e., the "second neural network model under the source domain" can be obtained. The second neural network model under the source domain can make predictions based on the data obtained in the source domain and output inference results.
[0044] Step 104: obtaining a target weight matrix of the target domain according to the second neural network model in the source domain and the second historical data through a model conversion device;
[0045] Specifically, the model conversion device may be a fine-tuned transferred model (FTTM). The model conversion device is used to fine-tune the second neural network model in the source domain so that it can better adapt to the data characteristics and task requirements of the target domain.
[0046] Since the second historical data is obtained in the target domain, and the trained second neural network model is trained under the source domain conditions, if the second historical data is directly loaded into the second neural network model trained under the source domain conditions, the second neural network model in the source domain will produce large reasoning errors when processing target domain tasks.
[0047] Therefore, the optimization method of the neural network model provided in the embodiment of the present application needs to obtain the target weight matrix of the second neural network model adapted to the target domain based on the second historical data and the second neural network model in the source domain, so as to provide a basis for transferring the second neural network model in the source domain to the target domain.
[0048] Specifically, the model conversion device can use the second historical data to perform a limited number of training cycles on the weight matrix of the second neural network model in the source domain, thereby obtaining a target weight matrix, so that the converted second neural network model can perform better in the target domain.
[0049] Step 105: The model conversion device performs model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain the target neural network model in the target domain.
[0050] It is worth noting that the basis for the model conversion device to convert the second neural network model in the source domain is the target weight matrix of the target domain, and there is no need to adjust the hyperparameters of the second neural network model. Since the adjustment of hyperparameters requires multiple manual verifications and readjustments. Therefore, the model conversion device can effectively improve the efficiency of model conversion, simplify the model conversion process, and avoid the extra time and computational cost of adjusting hyperparameters during the model conversion process by transferring the model through the weight matrix in the target domain, thereby quickly obtaining the target neural network model in the target domain.
[0051] Of course, if the second historical data in the target domain has been obtained, the second neural network model can also be trained using the second historical data to obtain the third neural network model in the target domain. However, the computational cost and time cost of training the second neural network model again using the second historical data are much higher than the computational cost and time cost of transferring the second neural network model in the source domain to obtain the target neural network model in the target domain. This model transfer method can avoid retraining the model and save time cost and computing resources.
[0052] In this way, the optimization method of the neural network model provided in the embodiment of the present application preliminarily reduces the model scale of the first neural network model to obtain the second neural network model by compressing the first neural network model as the basic model, and uses the first historical data to train the second neural network model to obtain the second neural network model adapted to the source domain. In the case where it is necessary to perform domain adaptation on the trained second neural network model in the source domain, the model conversion device can quickly convert the second neural network model in the source domain into the target neural network model in the target domain by determining the target weight matrix of the target domain, which can avoid retraining the model to save time cost and improve the adaptability of the second neural network model in the target domain. Through the above process of optimizing the neural network model, the complexity of the second neural network model is effectively reduced, the adaptability of the second neural network model is improved, and the model conversion cost is reduced, providing a new solution for indicator prediction in the process industry.
[0053] In some embodiments of the present application, the above step 101 may specifically include the following steps 1011 to 1014:
[0054] Step 1011: Determine a plurality of residual units in the first neural network model.
[0055] Each residual unit includes an input, an output, and a shortcut connection, and the shortcut connection is used to connect the output of one residual unit with the input of an adjacent residual unit.
[0056] As mentioned above, the first neural network model is a residual network model, which includes multiple residual units. The residual unit is the basic building block of the residual network model. Each residual unit includes input, output, shortcut connection, trainable parameters (weight coefficients) and activation function.
[0057] The input of the residual unit receives the output from the previous residual unit or the input data from the input layer. After the residual unit is calculated, the output result is passed to the next residual unit or the output layer. The shortcut connection directly connects the output of the previous residual unit to the input of the current residual unit.
[0058] Step 1012: According to the Adams method, determine the weight coefficients of the shortcut connections between adjacent residual units.
[0059] In the first neural network model, the connection between each residual unit may not be a simple direct connection, but a shortcut connection to transfer weight coefficients. The determination of the weight coefficient is crucial to the performance of the model.
[0060] The method provided in the embodiment of the present application uses the Adams method to calculate the weight coefficients transmitted between adjacent residual units through shortcut connections. The Adams method is a numerical differential equation solving method that gradually approximates the solution of the differential equation in an iterative manner.
[0061] Specifically, the first neural network model constructs a complex transformation by composing a sequence to transform the input to the hidden layer:
[0062] h t+11 =h t +f(h t +θ t ); (1)
[0063] where h t ∈R D is the output at time t, D represents the dimension of the output, t∈{0, 1, …, T}, T represents the output layer of the first neural network model at the final moment, θ t is the weight of the network layer at time t. These iterations can be viewed as the Euler discretization of the continuous transformation.
[0064] When the number of layers is large enough and the step size is small enough, we take the limit case on this basis and use the ordinary differential equation (ODE) corresponding to the neural network to parameterize the hidden units in the continuous dynamic system:
[0065]
[0066] Starting from the input layer h0, we can define hT Output.
[0067] The first neural network model is associated with the initial value solution problem of ODE, considering the form of Adams method as follows:
[0068]
[0069] Where Δh is the step size, f n+i is f(h t+i +θ t+i ) abbreviation, β i f n+i The corresponding weights and satisfy
[0070] Let m = 2, β in equation (3) m = 0, then the adjacent network layers of the hidden layer can be constructed as shown in formula (4):
[0071] h n+2 =h n+1 +k n f n +(1-k n )f n+1 (4)
[0072] k n ∈R is the trainable parameter in the corresponding hidden layer, that is, the weight coefficient that can be trained. In particular, when m = 1, it is the Euler method in the differential numerical solution.
[0073] Based on this, a weight coefficient can be obtained, which can be used to merge multiple residual units and construct shortcut connections of the Adams method between the residual units of the first neural network model. The shortcut connection can transfer the weight coefficient to adjacent residual units, thereby reducing the number of layers of the first neural network model.
[0074] Step 1013: Reconstruct the first neural network model according to the output of the previous residual unit, the input of the current residual unit and the weight coefficient.
[0075] Specifically, based on the output of the previous residual unit, the input of the current residual unit, and the weight coefficients transferred by the shortcut connection, the reconstruction strategy of the neural network model can be determined. For example, the residual units whose functions can be effectively simulated by shortcut connections can be removed or merged, thereby reducing the number of layers of the model and changing the internal structure of the residual unit (such as increasing or decreasing the depth and width of the convolutional layer).
[0076] Furthermore, the parameters of the retained residual units are migrated and initialized, thereby reconstructing the first neural network model. Therefore, the model size of the reconstructed first neural network model is compressed.
[0077] In some possible embodiments, step 1013 may specifically include the following steps 1013a to 1013c: Step 1013a: multiply the output of the previous residual unit by the corresponding weight coefficient, and add it to the input of the current residual unit to obtain the updated input of the current residual unit.
[0078] According to the weight coefficient calculated by the Adams method, the corresponding weight coefficient is determined for each residual unit, and then the output of the previous residual unit is multiplied by the corresponding weight coefficient to obtain the weighted output. The weighted output is then added to the original input of the current residual unit to obtain the updated input of the current residual unit.
[0079] Step 1013b: traverse each residual unit and obtain the updated input of the corresponding residual unit.
[0080] The updated inputs of all residual units are set as the original inputs, and then each residual unit is traversed to calculate the updated input of each residual unit according to the method of step 1013a.
[0081] Step 1013c: Reconstruct the first neural network model according to each updated input and the structure of the corresponding residual unit.
[0082] According to each updated input and the structure of the corresponding residual unit, the residual unit is removed or merged to determine the reconstruction strategy. According to the reconstruction strategy, the structure of the first neural network model is adjusted. For example: deleting some convolutional layers, merging some convolutional layers, adjusting the parameters of the convolutional layers, etc., and then migrating the parameters of the retained residual units and initializing the newly added parameters to obtain the reconstructed first neural network model.
[0083] Step 1014: Train the reconstructed first neural network model to obtain a second neural network model.
[0084] In order to improve the performance of the first neural network model, the first historical data may be used again to train and test the first neural network model.
[0085] The training process of the first neural network model may include: calculating the gradient of the loss function to the model weights through a back-propagation algorithm, and using an optimizer to update the model weights according to the gradient to improve the performance of the model.
[0086] When the trained first neural network model passes the test, the second neural network model is obtained. Figure 2 (a) in FIG. 1 shows the shortcut connection structure of the first neural network model. Figure 2(b) shows that the second neural network model is different from the shortcut connection structure of the first neural network model. The "second neural network model" is a neural network model structure based on the first neural network model, which uses the Adams method to perform shortcut connections and compress the model size.
[0087] In this way, by training and updating the weight coefficients of the residual units in the first neural network model, reconstructing the first neural network model based on the weight coefficients, and training the reconstructed first neural network model, a second neural network model with a smaller model scale is obtained, thereby improving the learning accuracy and output efficiency of the second neural network model, and preliminarily optimizing the neural network model.
[0088] Exemplarily, in order to prove that the output efficiency of the second neural network model is improved, a grid search hyperparameter optimization method can be used to traverse the hyperparameter combinations in the first neural network model and the second neural network model, and grid search is used to traverse and solve, and CC (representing the Pearson colocalization coefficient, indicating the degree of match between the predicted and actual values, 1 indicates a perfect match) is used to judge the output performance of the two models.
[0089] Table 1
[0090] Batch size Network Layers First filter press 32 44 Second filter press 32 56
[0091] Table 2
[0092] Batch size Network Layers First filter press 32 34 Second filter press 32 50
[0093] Table 3
[0094]
[0095] Table 1 shows the optimal hyperparameters of the first neural network model for the source domain and the target domain, and Table 2 shows the optimal hyperparameters of the second neural network for the source domain and the target domain. Table 3 shows the output parameters of the first neural network model and the second neural network model.
[0096] It can be seen that in the source domain environment, the output results of the first neural network model and the second neural network model for MAE (mean absolute error), RMSE (root mean square error), and CC are similar, but the algorithm solution rate of the second neural network model is faster. The first neural network model runs for 45 seconds, while the second neural network model obtained after compression using the Adam method only runs for 30 seconds. It proves that the second neural network model can effectively improve the output efficiency while ensuring the output accuracy.
[0097] In some embodiments, step 104 may include step 1041 and step 1042:
[0098] Step 1041: Using the weight matrix of the second neural network model in the source domain as the initial weight matrix of the target neural network model in the target domain.
[0099] The second neural network model in the source domain can be expressed as y=f(w,h), where w represents the weight matrix and h represents the hyperparameter of the model. The process of the model transfer device migrating the second neural network model in the source domain can be:
[0100] Load f s (w s ,h s );
[0101] Load x T ,y T ;
[0102] Use x T ,y T Fine-tune f s (w s ,h s );
[0103] Test f T (w T ,h T ).
[0104] Specifically, the features of the second neural network model in the source domain are contained in the vector w s and h s In which the subscript S represents the source domain. S ,y S ) is the first historical data in the source domain; the second historical data in the target domain is the target domain dataset (i.e. x T ,y T ), due to w s and h s is the result obtained through training, so f s (w s ,h s ) as the initial weight matrix and initial hyperparameters of the neural network model desired in the target domain.
[0105] Step 1042: Update the initial weight matrix according to the second historical data and the LM least squares algorithm to obtain the target weight matrix of the target domain.
[0106] Use the second historical data x of the target domain T ,y T As data input f s (w s ,h s ), use the LM least squares algorithm to calculate ws Perform a finite number of training cycles to obtain the target weight matrix w for the new target domain T In addition, since the selection of hyperparameters takes a lot of time, using h s Substituting for the hyperparameters of the target domain, the target domain function after a finite number of training cycles is denoted as f T (w T ,h T ). In this way, the time cost of manually selecting hyperparameters can be avoided, further improving the efficiency of the neural network model for domain adaptation.
[0107] In some embodiments, step 1042 may further include step 1042a and step 1042b:
[0108] Step 1042a: Use the second historical data to train the initial weight matrix, and update the trained weight matrix using the LM least squares algorithm.
[0109] Specifically, the weight matrix w of the second neural network model in the source domain can be s and the hyperparameter h s As the initial weight matrix and hyperparameters of the target neural network model in the target domain. Then use the second historical data xT, yT to input into the second neural network model in the source domain to obtain the output of the second neural network model in the source domain. Then, use the difference between the output and the true value yT to obtain the loss function value.
[0110] Furthermore, the LM least squares algorithm is used to calculate the gradient of the loss function with respect to the weight matrix, and the weight matrix w is updated according to the gradient s , in order to reduce the loss function value.
[0111] Step 1042b: Repeatedly use the second historical data to train the updated weight matrix, and update the weight matrix again through the LM least squares algorithm until a preset convergence condition or a preset number of iterations is reached, thereby obtaining a target weight matrix of the target domain.
[0112] Repeat the process of step 1042a, using the updated weight matrix w in each iteration s After each iteration, it is determined whether the preset convergence condition is met. If the preset convergence condition is met, the target weight matrix under the target domain is determined. The final target weight matrix w T , which is the target weight matrix adapted to the target domain, w t It will be used to build the target neural network model under the target domain.
[0113] Exemplarily, the preset convergence condition may be any one of the following situations: the decrease amplitude of the loss function is less than a preset threshold, the update amplitude of the weight matrix is less than a preset threshold, and a preset number of iterations is reached.
[0114] In some embodiments, step 105 may include the following steps 1051 to 1053:
[0115] Step 1051: Load the target weight matrix of the target domain into the second neural network model under the source domain.
[0116] The target weight matrix w obtained in the above process T The value of is copied to the corresponding weight parameter of the second neural network model under the source domain. After loading is completed, a forward propagation verification can be performed to ensure that the target weight matrix w T It has been correctly loaded into the second neural network model under the source domain, and the second neural network model can work normally.
[0117] Step 1052: In the target domain, the network structure of the second neural network model in the source domain is adjusted according to the target weight matrix to obtain the target neural network model in the target domain.
[0118] According to the data characteristics and task requirements of the target domain, adjust the convolutional layers that need to be added or reduced in the second neural network model. And according to the target weight matrix w T According to the characteristics of the convolutional layer, the parameters of the convolutional layer, such as filter size, step size, padding, etc., are adjusted to obtain the network structure of the second neural network model in the target domain.
[0119] Optionally, the second historical data of the target domain can be used again to fine-tune the network structure of the second neural network model under the adjusted target domain to further optimize the model performance.
[0120] Step 1053: Test and verify the second neural network model in the target domain after the network structure is updated to obtain the target neural network model.
[0121] The adjusted second neural network model under the target domain is verified to determine whether it meets the verification conditions. If the verification passes, it means that the model performance of the second neural network model under the target domain meets the requirements, and the target neural network model under the target domain is obtained. If the verification fails, it is necessary to repeat step 1052 to adjust the second neural network model under the target domain again.
[0122] In some embodiments, step 1053 may include the following steps 1053a to 1053d:
[0123] Step 1053a: Divide the second historical data into a training set and a test set.
[0124] The second historical data is divided into two parts, namely a training set and a test set. The training set accounts for most of the second historical data and is used to train the second neural network model in the target domain, while the test set is used to evaluate the performance of the second neural network model in the target domain.
[0125] Step 1053b: Use the training set to train the second neural network model to obtain a third neural network model in the target domain.
[0126] It should be noted that the training market and computing cost of training the second neural network using the second historical data are high, but the prediction accuracy of the third neural network model trained by this method is higher. The third neural network model is not the target product of the neural network model provided in the embodiment of the present application, but is only used to prove that the use of the model transfer device can ensure the prediction performance of the second neural network model in the source domain after domain adaptation.
[0127] It can be understood that after one verification, it means that in the optimization method of the neural network model provided in the embodiment of the present application, the method of using the model transfer device for domain adaptation is feasible, and testing and verification is not required after each domain adaptation.
[0128] Step 1053c: input the test set into the third neural network model under the target domain to obtain a first output result; input the test set into the second neural network model under the target domain to obtain a second output result.
[0129] Exemplarily, the first output result and the second output result may be the operating hours of the filter press disclosed above.
[0130] Step 1053d: If the difference between the first output result and the second output result is within a preset error range, it is determined that a target neural network model under the target domain is obtained.
[0131] Table 4
[0132]
[0133]
[0134] Exemplarily, Table 4 shows the output parameters of the third neural network and the second neural network model in the target domain after network structure adjustment under the target domain environment. Compared with the two, the output results of the Pearson colocalization coefficient (CC), MAE (mean absolute error) and RMSE (root mean square error) of the second neural network model in the target domain are closer to the results of the third neural network model and can meet the requirements. In addition, the error between the CC value output by the second neural network model in the target domain and the CC value output by the third neural network model is only 8%, which is within the preset error range (which can be set to 10%), indicating that at this time, the second neural network model in the target domain has been verified and the target neural network model in the target domain has been obtained.
[0135] In addition, Table 4 also shows the output results of the second neural network model if the test machine of the second historical data is directly loaded into the source domain. Taking the output results of the third neural network model as a benchmark, in comparison, the output results of the second neural network model in the source domain for MAE (mean absolute error) and RMSE (root mean square error) are significantly different from those of the third neural network model, and the error between the CC value output by the second neural network model in the source domain and the CC value output by the third neural network model increases to 20%. It can be seen that the target neural network model in the target domain obtained after the transfer by the model transfer device is optimized in output accuracy compared with the second neural network model directly in the source domain, and the reasoning accuracy in the target domain environment is higher.
[0136] In addition, the data in Table 4 also show that the output efficiency (45s) of the second neural network model in the target domain, that is, the target neural network model, is also higher than the output efficiency (55s) of the second neural network model in the source domain, and the output efficiency is accelerated, indicating that the target neural network model after the model conversion device is used for model transfer is more adaptable to the target domain. The final target neural network model takes into account both the output efficiency and the output accuracy, so it can significantly improve the prediction accuracy and prediction efficiency of the target indicators after it is applied to the process industry.
[0137] like Figure 3As shown, the embodiment of the present application also provides an optimization device 300 for a neural network model. The optimization device 300 for the neural network model in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook or a personal digital assistant (Personal Digital Assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0138] The optimization device 300 of the neural network model provided in the embodiment of the present application includes:
[0139] A compression unit 301, configured to compress the first neural network model to obtain a second neural network model;
[0140] An acquisition unit 302, configured to acquire first historical data in a source domain and acquire second historical data in a target domain;
[0141] A training unit 303 is used to train the second neural network model through the first historical data to obtain a second neural network model in the source domain;
[0142] A processing unit 304 is used to obtain a target weight matrix of a target domain according to a second neural network model and second historical data in a source domain through a model conversion device;
[0143] The processing unit 304 is also used for the model conversion device to perform model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain the target neural network model in the target domain.
[0144] In some embodiments, the compression unit 301 is specifically used to: determine multiple residual units in the first neural network model, each residual unit includes an input, an output and a shortcut connection, and the shortcut connection is used to connect the output of a residual unit with the input of an adjacent residual unit; according to the Adams method, determine the weight coefficient weighted by the shortcut connection between adjacent residual units; reconstruct the first neural network model according to the output of the previous residual unit, the input of the current residual unit and the weight coefficient; train the reconstructed first neural network model to obtain the second neural network model.
[0145] In some embodiments, the compression unit 301 is also specifically used to: multiply the output of the previous residual unit by the corresponding weight coefficient, and add it to the input of the current residual unit to obtain the updated input of the current residual unit; traverse each residual unit and obtain the updated input of the corresponding residual unit; reconstruct the first neural network model according to the structure of each updated input and the corresponding residual unit.
[0146] In some embodiments, the processing unit 304 is also specifically used to: use the weight matrix of the second neural network model in the source domain as the initial weight matrix of the target neural network model in the target domain; update the initial weight matrix according to the second historical data and the LM least squares algorithm to obtain the target weight matrix of the target domain.
[0147] In some embodiments, the processing unit 304 is also specifically used to: use the second historical data to train the initial weight matrix, and update the trained weight matrix through the LM least squares algorithm; repeatedly use the second historical data to train the updated weight matrix, and update the weight matrix again through the LM least squares algorithm until the preset convergence condition is reached, thereby obtaining the target weight matrix of the target domain.
[0148] In some embodiments, the processing unit 304 is further specifically configured to: load the target weight matrix of the target domain into the second neural network model under the source domain;
[0149] In the target domain, the network structure of the second neural network model is adjusted according to the target weight matrix; the second neural network model in the target domain after the updated network structure is tested and verified to obtain the target neural network model.
[0150] In some embodiments, the processing unit 304 is further specifically used to: divide the second historical data into a training set and a test set; use the training set to train the second neural network model to obtain a third neural network model under the target domain; input the test set into the third neural network model under the target domain to obtain a first output result; input the test set into the second neural network model under the target domain to obtain a second output result; if the difference between the first output result and the second output result is within a preset error range, it is determined that the target neural network model under the target domain is obtained.
[0151] The neural network model optimization device 300 provided in the embodiment of the present application can implement each process implemented by the neural network model optimization method embodiment in any of the above-mentioned embodiments. To avoid repetition, the beneficial effects will not be described again here.
[0152] The present application also provides an electronic device, such as Figure 4 As shown, the electronic device 400 includes a processor 401 and a memory 402. The memory 402 stores programs or instructions that can be executed on the processor 401. When the program or instruction is executed by the processor 401, the various steps of the optimization method embodiment of the above-mentioned neural network model are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0153] The memory 402 can be used to store software programs and various data. The memory 402 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 402 may include a volatile memory or a non-volatile memory, or the memory 402 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 402 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0154] The processor 401 may include one or more processing units; optionally, the processor 401 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 401.
[0155] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned neural network model optimization method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0156] The embodiment of the present application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned neural network model optimization method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0157] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0158] The embodiment of the present application also provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the optimization method embodiment of the neural network model as described above, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0159] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0160] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A method for optimizing a neural network model, characterized in that: include: Compressing the first neural network model to obtain a second neural network model; Acquiring first historical data in a source domain, and acquiring second historical data in a target domain; Training the second neural network model by using the first historical data to obtain the second neural network model in the source domain; Obtaining a target weight matrix of the target domain according to the second neural network model and the second historical data in the source domain by a model conversion device; The model conversion device performs model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain a target neural network model in the target domain.
2. The optimization method of the neural network model according to claim 1, characterized in that: The step of compressing the first neural network model to obtain the second neural network model includes: Determine a plurality of residual units in the first neural network model, each of the residual units comprising an input, an output, and a shortcut connection, wherein the shortcut connection is used to connect the output of one residual unit with the input of an adjacent residual unit; Determine, according to the Adams method, weight coefficients weighted by the shortcut connections between adjacent residual units; Reconstructing the first neural network model according to the output of the previous residual unit, the input of the current residual unit and the weight coefficient; The reconstructed first neural network model is trained to obtain the second neural network model.
3. The optimization method of the neural network model according to claim 2, characterized in that: The reconstructing the first neural network model according to the output of the previous residual unit, the input of the current residual unit and the weight coefficient comprises: Multiplying the output of the previous residual unit by the corresponding weight coefficient and adding the result to the input of the current residual unit to obtain an updated input of the current residual unit; Traversing each of the residual units and obtaining the updated input of the corresponding residual unit; Reconstruct the first neural network model according to each updated input and the corresponding structure of the residual unit.
4. The method for optimizing a neural network model according to claim 1, characterized in that: The step of obtaining the target weight matrix of the target domain according to the second neural network model and the second historical data in the source domain by the model conversion device includes: Using the weight matrix of the second neural network model in the source domain as the initial weight matrix of the target neural network model in the target domain; The initial weight matrix is updated according to the second historical data and the LM least squares algorithm to obtain the target weight matrix of the target domain.
5. The method for optimizing a neural network model according to claim 4, characterized in that: The updating of the initial weights according to the second historical data and the LM least squares algorithm to obtain the target weight matrix of the target domain includes: Using the second historical data to train the initial weight matrix, and updating the trained weight matrix by using the LM least squares algorithm; The updated weight matrix is trained repeatedly using the second historical data, and the weight matrix is updated again by the LM least squares algorithm until a preset convergence condition is reached, thereby obtaining the target weight matrix of the target domain.
6. The method for optimizing a neural network model according to claim 1, characterized in that: The model conversion device performs model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain a target neural network model in the target domain, including: Loading the target weight matrix of the target domain into the second neural network model under the source domain; Under the target domain, adjusting the network structure of the second neural network model according to the target weight matrix; The second neural network model in the target domain after the updated network structure is tested and verified to obtain the target neural network model.
7. The method for optimizing a neural network model according to claim 6, characterized in that: The testing and verifying of the second neural network model in the target domain after the updated network structure to obtain the target neural network model includes: Dividing the second historical data into a training set and a test set; Using the training set to train the second neural network model to obtain a third neural network model under the target domain; Input the test set into the third neural network model under the target domain to obtain a first output result; input the test set into the second neural network model under the target domain to obtain a second output result; If the difference between the first output result and the second output result is within a preset error range, it is determined that the target neural network model under the target domain is obtained.
8. An optimization device for a neural network model, characterized in that: The method is applied to electronic equipment, and the optimization device of the neural network model includes: A compression unit, used for compressing the first neural network model to obtain a second neural network model; An acquisition unit, configured to acquire first historical data in a source domain and second historical data in a target domain; A training unit, configured to train the second neural network model by using the first historical data to obtain the second neural network model in the source domain; A processing unit, configured to obtain a target weight matrix of the target domain according to the second neural network model and the second historical data in the source domain through a model conversion device; The processing unit is also used by the model conversion device to perform model conversion processing on the trained second neural network model in the source domain according to the target weight matrix of the target domain to obtain the target neural network model in the target domain.
9. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores a program or instruction running on the processor, and when the program or instruction is executed by the processor, the steps of the optimization method of the neural network model as described in any one of claims 1 to 8 are implemented.
10. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps of the optimization method of the neural network model as described in any one of claims 1 to 8 are implemented.