Model Pruning Method, Apparatus and Electronic Device
By using the second neural network to calculate and update the structural information to prune the neural network, the problem of complex and inefficient pruning process in the prior art is solved, efficient model pruning is achieved and the application scope is expanded.
Patent Information
- Application Number
- CN202010769414.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-08-03
AI Technical Summary
In the prior art, the process of pruning neural network models is more complicated, requires a lot of time and is less efficient.
By using the second neural network to calculate the structural information based on the current neural network and setting index values, the structural information is calculated using the second neural network, and the structure updates the first neural network based on this to finally determine the model pruning result.
The model pruning process is simplified, the model pruning efficiency is improved, and no layer-by-layer pruning is required, and no derivatable functions are required during the pruning process, which expands the scope of application.
Smart Images

Figure CN111931930B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model pruning method, apparatus, and electronic device. Background Art
[0002] Since a neural network model requires a large amount of computing resources and storage resources for support, while the computing resources and storage resources of a mobile terminal are limited, this restricts the application of the neural network model in the mobile terminal. In related technologies, pruning processing is performed on the neural network model to reduce the amount of computation of the neural network, so as to reduce the consumption of computing resources and storage resources by the neural network. However, the process of pruning the neural network model is relatively complex, consuming a lot of time and having low efficiency. Summary of the Invention
[0003] In view of this, embodiments of this application are expected to provide a model pruning method, apparatus, and electronic device to solve the technical problem that the process of pruning a neural network model in related technologies is relatively complex and consumes a lot of time.
[0004] To achieve the above object, the technical solution of this application is implemented as follows:
[0005] Embodiments of this application provide a model pruning method, including:
[0006] Based on the first structure information of the current first neural network and the corresponding set index value, calculate the second structure information through a second neural network; wherein, the current first neural network is trained based on the set training samples; the initial first neural network is constructed based on the third structure information of a third neural network; the corresponding set index value is used to update the weight parameters of the second neural network, and the second neural network is used to output the second structure information based on the input first structure information after updating the weight parameters;
[0007] Update the structure of the current first neural network based on the second structure information;
[0008] In the case where the second neural network reaches the set convergence condition, determine the first neural network with the updated structure as the model pruning result corresponding to the third neural network; wherein,
[0009] When the first neural network is initially constructed or the structure of the first neural network is updated, train the first neural network based on the set training samples.
[0010] In the above solution, when constructing the initial first neural network, the method includes:
[0011] Input the third structural information of the third neural network into the second neural network to obtain the fourth structural information output by the second neural network;
[0012] Construct an initial first neural network based on the fourth structural information.
[0013] In the above solution, the calculation of the second structural information by the second neural network includes:
[0014] Use at least one set test sample to test the current first neural network to obtain the test result corresponding to each test sample in the at least one set test sample; the test result represents the set index value corresponding to the corresponding test sample;
[0015] Based on the test result corresponding to each test sample in the at least one set test sample, use a set loss function to calculate the loss value corresponding to the second neural network;
[0016] Update the weight parameters of the second neural network according to the calculated loss value;
[0017] In the case of updating the weight parameters, input the first structural information of the current first neural network into the second neural network to obtain the second structural information output by the second neural network.
[0018] In the above solution, the structural update of the current first neural network based on the second structural information includes at least one of the following:
[0019] Update the topological structure of the current first neural network based on the topological structure included in the second structural information;
[0020] Update the weight channels of the corresponding layer of the current first neural network based on the number of weight channels of the corresponding layer included in the second structural information; the number of weight channels represents the number of input channels and output channels of the corresponding layer;
[0021] Update the weight values of the corresponding layer of the current first neural network based on the weight precision of the corresponding layer included in the second structural information; the weight precision represents the number of bits occupied by the weight values of the corresponding layer;
[0022] Update the output precision of the activation function of the corresponding layer of the current first neural network based on the output precision of the activation function of the corresponding layer included in the second structural information; the output precision of the activation function represents the precision of the output result of the corresponding activation function;
[0023] Update the weight value of the corresponding layer of the current first neural network based on the pruning threshold of the weight of the corresponding layer included in the second structure information; wherein, when the absolute value of the weight value of the corresponding layer is less than the corresponding pruning threshold, set the corresponding weight value to zero; when the absolute value of the weight value of the corresponding layer is greater than or equal to the corresponding pruning threshold, keep the weight value of the corresponding layer unchanged.
[0024] In the above solution, the set index value includes at least one of the following:
[0025] The first index value; the first index value characterizes the performance parameter of the current first neural network;
[0026] At least one second index value; the second index value characterizes the cost parameter when the current first neural network processes each of the at least one set test samples.
[0027] In the above solution, when training the first neural network based on the set training samples, the method includes:
[0028] Determine the set training samples; wherein,
[0029] The set training samples include at least one sample pair; each sample pair in the at least one sample pair includes an input sample and a corresponding first calibration sample; wherein, the first calibration sample is used to compare with the output result obtained by the first neural network when inputting the corresponding input sample; the first calibration sample is determined by the corresponding second calibration sample and the corresponding reference sample of the corresponding input sample; the second calibration sample characterizes the calibration sample corresponding to the input sample of the third neural network; the reference sample characterizes the output result obtained by the third neural network when inputting the corresponding input sample.
[0030] In the above solution, when determining the set training samples, the method includes:
[0031] Perform a fusion process on the corresponding second calibration sample and the corresponding reference sample of the input sample based on the set weight value to obtain the corresponding first calibration sample; the weight value of the second calibration sample corresponding to the input sample is greater than the weight value of the corresponding reference sample.
[0032] An embodiment of the present application further provides a model pruning device, including:
[0033] A calculation unit, configured to calculate second structure information through a second neural network based on first structure information of a current first neural network and a corresponding set index value; wherein, the initial first neural network is constructed based on third structure information of a third neural network; the corresponding set index value is used to update weight parameters of the second neural network, and the second neural network is configured to output the second structure information based on the input first structure information after updating the weight parameters;
[0034] An update unit, configured to perform structure update on the current first neural network based on the second structure information;
[0035] A determination unit, configured to, when the second neural network reaches a set convergence condition, determine the first neural network with updated structure as a model pruning result corresponding to the third neural network; wherein,
[0036] When the first neural network is initially constructed or the structure of the first neural network is updated, the first neural network is trained based on set training samples.
[0037] An embodiment of the present application further provides an electronic device, including: a processor and a memory for storing a computer program that can run on the processor,
[0038] wherein, when the processor is configured to run the computer program, the above-mentioned any model pruning method is implemented.
[0039] An embodiment of the present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned any model pruning method is implemented.
[0040] In an embodiment of the present application, an initial first neural network is constructed based on structure information of a third neural network to be pruned, and based on first structure information of the first neural network and a corresponding set index value, a second neural network for pruning is used to calculate third structure information, and the second neural network is structurally updated based on the calculated third structure information to prune the second neural network. When the second neural network reaches a set convergence condition, the first neural network with updated structure is determined as a model pruning result corresponding to the third neural network. In this way, the structure of the entire first neural network can be updated based on the structure information calculated by the second neural network, without the need to perform layer-by-layer pruning on the first neural network, simplifying the model pruning process and improving the model pruning efficiency; in addition, during the model pruning process, there is no need to design a differentiable function for pruning, and the application scope of this model pruning method is wider. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic implementation flowchart of the model pruning method provided by an embodiment of the present application;
[0042] Figure 2 Schematic diagram of the implementation process of the model pruning method provided in another embodiment of the present application;
[0043] Figure 3 Schematic diagram for determining the second structure information in the model pruning method provided in the embodiment of the present application;
[0044] Figure 4 Schematic diagram of the implementation process of constructing an initial first neural network in the model pruning method provided in the embodiment of the present application;
[0045] Figure 5 Schematic diagram of the implementation process for determining the second structure information in the model pruning method provided in another embodiment of the present application;
[0046] Figure 6 Schematic diagram of the structure of a neural network provided in the embodiment of the present application;
[0047] Figure 7 Schematic diagram of the structure of the model pruning device provided in the embodiment of the present application;
[0048] Figure 8 Schematic diagram of the hardware composition structure of the electronic device provided in the embodiment of the present application. Detailed implementation manners
[0049] Before introducing the technical solution of the present application, the model pruning method in the related art is introduced first.
[0050] A model pruning method is provided in the related art, which mainly prunes the neural network layer by layer by reducing the filter dimension of the corresponding layer. In the process of layer-by-layer pruning, it takes a long time to finally determine the filter dimensions of each layer of the neural network. Among them, the filter dimension represents the size information of the filter, and this size information can represent the length and width of the filter, and can also represent the depth of the filter.
[0051] Another model pruning method is also provided in the related art. By embedding a pruning module (such as a threshold function for making the weights sparse) in the neural network, during the training of the neural network, the embedded pruning module is updated together with the weight parameters of the neural network, so as to achieve model pruning. During the training of the neural network, in order to ensure that the backpropagation can proceed normally, the embedded pruning module must be a differentiable function. When the embedded pruning module is not differentiable, this method cannot be used for model pruning, which limits the application scope of this model pruning method.
[0052] To solve the above technical problems, the present application provides a model pruning method. An initial first neural network is constructed based on the structural information of a third neural network to be pruned. Based on the first structural information of the first neural network and the corresponding set index value, a second neural network for pruning is used to calculate third structural information. Based on the calculated third structural information, the structure of the second neural network is updated to prune the second neural network. When the second neural network reaches the set convergence condition, the first neural network with the updated structure is determined as the model pruning result corresponding to the third neural network. The technical solution provided by the present application can update the structure of the entire first neural network based on the structural information calculated by the second neural network, without the need to prune the first neural network layer by layer, simplifying the model pruning process and improving the model pruning efficiency. In addition, during the model pruning process, there is no need to design a differentiable function for pruning, and the model pruning method provided by the present application has a wider application range.
[0053] The technical solution of the present application will be further elaborated in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0054] Figure 1 The figure shows a schematic implementation flow diagram of the model pruning method provided by an embodiment of the present application. In the embodiment of the present application, the execution subject of the model pruning method can be an electronic device such as a terminal or a server.
[0055] Refer to Figure 1 , the model pruning method provided by an embodiment of the present application includes:
[0056] S101: Based on the first structural information of the current first neural network and the corresponding set index value, calculate second structural information through a second neural network; wherein, the current first neural network is trained based on the set training samples; the initial first neural network is constructed based on the third structural information of the third neural network; the corresponding set index value is used to update the weight parameters of the second neural network, and the second neural network is used to output the second structural information based on the input first structural information after updating the weight parameters.
[0057] Here, the current first neural network refers to the first neural network that has been trained. When the electronic device first executes S101, the current first neural network is obtained by training the initial first neural network; when executing S101 for the Nth time, the current first neural network is obtained by training the first neural network with the updated structure obtained by executing S102 for the (N - 1)th time. N is an integer greater than or equal to 2.
[0058] Among them, the initial first neural network is a neural network constructed based on the third structural information of the trained third neural network. The third neural network is the neural network to be pruned, or the original network. The second neural network is used to prune the current second neural network. The second neural network is different from the first neural network.
[0059] The corresponding set index value characterizes the performance of the current first neural network. The corresponding set index value can be obtained when the electronic device tests the current first neural network using at least one set test sample. The corresponding set index value is used to calculate the loss value of the second neural network, so that the electronic device can update the weight parameters of the second neural network based on the loss value of the second neural network, and when the weight parameters are updated, output the second structural information based on the first structural information of the current first neural network.
[0060] It should be noted that the training samples used to train the third neural network and the training samples used to train the first neural network can be the same or different.
[0061] In practical applications, the structural information of the initial first neural network and the structural information of the third neural network can be the same or different. That is to say, the initial first neural network and the third neural network can be the same or different. When the initial first neural network and the third neural network are different, it can be manifested as any of the following:
[0062] The structure of the initial first neural network is simpler than the structure of the third neural network; for example, the number of layers of the first neural network is less than the number of layers of the third neural network, and the number of neurons in the corresponding layer of the first neural network is less than the number of neurons in the corresponding layer of the third neural network.
[0063] The precision of the weight values of the initial first neural network is less than the precision of the corresponding weight values of the third neural network.
[0064] Please refer to Figure 2 and Figure 3 , Figure 2 shows a schematic flowchart of the implementation process of the model pruning method provided by another embodiment of the present application; Figure 3 shows a schematic diagram for determining the second structural information in the model pruning method provided by the embodiment of the present application. The following will introduce the implementation process of the model pruning method in combination with Figure 2 and Figure 3 to introduce the implementation process of the model pruning method.
[0065] When the electronic device completes the training of the third neural network using at least one set first training sample, an initial first neural network is constructed based on the third structural information of the third neural network.
[0066] The electronic device trains the initial first neural network using at least one set second training sample to obtain the trained first neural network, and determines the trained first neural network as the current first neural network.
[0067] In practical applications, the first training sample and the second training sample may be the same or different. Each training sample includes an input sample and a corresponding calibration sample, and the calibration sample is used to compare with the result output by the neural network when the corresponding input sample is input.
[0068] Here, during the process of training the third neural network or the first neural network, the electronic device uses the corresponding loss function to calculate the loss value of the neural network based on the output result corresponding to the input sample in the corresponding training sample and the corresponding calibration sample. The weight parameters of the neural network are updated based on the calculated loss value. Among them, the electronic device backpropagates the calculated loss value in the neural network. During the process of backpropagating the calculated loss value to each layer of the neural network, the gradient of the corresponding loss function is calculated according to the loss value, and the weight parameters backpropagated to the current layer are updated along the descending direction of the gradient.
[0069] When the neural network meets the set stop update condition, the update of the weight parameters of the neural network is stopped, and the weight parameters obtained from the last update are determined as the weight parameters used by the trained neural network. Among them, the set stop update condition may be the convergence of the loss function of the neural network, or the set number of training epochs (a training epoch is the process of training the neural network once according to the input samples and the corresponding calibration samples in at least one training sample). Of course, the update stop condition is not limited to this, for example, it can also be the set mean average precision (mAP). The mean average precision is calculated based on the output results corresponding to the input samples and the corresponding calibration samples in all the training samples participating in the training.
[0070] When the current first neural network is trained, the electronic device determines the set index value corresponding to the current first neural network, and calculates the loss value of the second neural network using the set first loss function based on the determined set index value. Among them, the test sample is different from the training sample, and the training sample here includes the above-mentioned first training sample and second training sample.
[0071] When the loss value of the second neural network is calculated, the electronic device detects whether the second neural network reaches a set convergence condition. Here, the set convergence condition can be the number of weight updates (or the number of backpropagations), or the convergence of the first loss function. The convergence of the first loss function indicates that the loss value of the first loss function tends to be stable, or the loss value of the first loss function approaches a certain constant.
[0072] When the detection result indicates that the second neural network does not reach the set convergence condition, the electronic device updates the weight parameters of the second neural network based on the calculated loss value. When the detection result indicates that the second neural network reaches the set convergence condition, the update of the weight parameters of the second neural network is stopped. Among them, the electronic device performs backpropagation of the calculated loss value in the second neural network. During the process of backpropagating the calculated loss value to each layer of the second neural network, the gradient of the first loss function is calculated according to the loss value, and the weight parameters backpropagated to the current layer are updated along the descending direction of the gradient.
[0073] After updating the weight parameters of the second neural network, or when stopping the update of the weight parameters of the second neural network, the electronic device inputs the first structure information of the current first neural network into the second neural network, and obtains the second structure information output by the second neural network.
[0074] It should be noted that the fact that the first training sample and the second training sample mentioned above are the same means that the input samples in the first training sample and the input samples in the second training sample are the same, and the corresponding calibration samples in the first training sample and the corresponding calibration samples in the second training sample are also the same.
[0075] When the first training sample and the second training sample are different, there are the following two cases:
[0076] The input samples in the first training sample and the input samples in the second training sample are the same, but the corresponding calibration samples are different; for example, the corresponding calibration sample in the first training sample is the calibration sample of the corresponding input sample, and the corresponding calibration sample in the second training sample is the output result obtained by the third neural network when inputting the corresponding input sample;
[0077] The input samples in the first training sample and the input samples in the second training sample are different, and the corresponding calibration samples are also different.
[0078] S102: Perform a structure update on the current first neural network based on the second structure information.
[0079] When the electronic device obtains the second structure information calculated by the second neural network, it performs pruning processing on the current first neural network based on the second structure information to perform a structure update on the current first neural network.
[0080] The second structure information may include at least one of the following:
[0081] Topological structure, the number of weight channels of the corresponding layer, the weight precision of the corresponding layer, the output precision of the activation function of the corresponding layer, the pruning threshold of the weights of the corresponding layer, etc.
[0082] Among them, the number of weight channels characterizes the number of input channels and the number of output channels of the corresponding layer. The weight precision of the corresponding layer characterizes the number of bits occupied by the weights of the corresponding layer. The output precision of the activation function characterizes the precision of the output result of the corresponding activation function. The pruning threshold of the weights of the corresponding layer is used to set the weight values that meet the set pruning conditions to zero. Here, the set pruning conditions include any one of the following: the weight value is less than the pruning threshold of the corresponding weight; the absolute value of the weight value is less than the pruning threshold of the corresponding weight.
[0083] After the current first neural network updates its structure, the electronic device trains the first neural network using at least one set second training sample to obtain a trained first neural network, and determines the trained first neural network as the current first neural network.
[0084] When the detection result indicates that the second neural network has not reached the set convergence condition and the training of the first neural network after updating the structure is completed, return to S101, so as to execute S101 - S102 again. Here, when the detection result indicates that the second neural network has not reached the set convergence condition, it is necessary to train again based on the set index value corresponding to the first neural network after updating the structure.
[0085] When the detection result indicates that the second neural network has reached the set convergence condition, stop updating the weight parameters of the second neural network. The electronic device determines the weight parameters obtained in the last update as the weight parameters used by the second neural network. After the electronic device updates the structure of the current first neural network based on the second structure information output by the second neural network, it executes S103.
[0086] S103: In the case where the second neural network reaches the set convergence condition, determine the first neural network after structure update as the model pruning result corresponding to the third neural network; among them, when the first neural network is constructed initially or the structure of the first neural network is updated, train the first neural network based on the set training samples.
[0087] Here, in the case where the second neural network reaches the first set convergence condition, the electronic device determines the first neural network after structure update in S102 as the model pruning result corresponding to the third neural network. The first neural network after structure update refers to the first neural network updated based on the second structure information output by the second neural network after the last update of the weight parameters.
[0088] In some embodiments, the electronic device may also train the first neural network after the last structural update using at least one set second training sample, and determine the trained first neural network as the model pruning result corresponding to the third neural network.
[0089] In the technical solution provided in this embodiment, an initial first neural network is constructed based on the structural information of the third neural network (i.e., the original network). Based on the first structural information of the first neural network and the corresponding set index value, the second neural network for pruning is used to calculate the second structural information. Based on the calculated second structural information, the structure of the first neural network is updated to prune the constructed first neural network. When the second neural network reaches the set convergence condition, the first neural network with the updated structure is determined as the model pruning result corresponding to the third neural network. In this way, the structure of the entire first neural network can be updated based on the structural information calculated by the second neural network, without the need to perform layer-by-layer pruning on the first neural network, simplifying the model pruning process and improving the model pruning efficiency. In addition, during the model pruning process, there is no need to design a differentiable function for pruning, and the application scope of this model pruning method is wider.
[0090] As another embodiment of the present application, when training the first neural network based on the set training sample, the method includes:
[0091] Determine the set training sample; where
[0092] The set training sample includes at least one sample pair; each sample pair in the at least one sample pair includes an input sample and a corresponding first calibration sample; where the first calibration sample is used to compare with the output result obtained by the first neural network when inputting the corresponding input sample; the first calibration sample is determined by the corresponding second calibration sample of the corresponding input sample and the corresponding reference sample; the second calibration sample represents the calibration sample corresponding to the input sample of the third neural network; the reference sample represents the output result obtained by the third neural network when inputting the corresponding input sample.
[0093] Here, the second calibration sample is used to compare with the output result obtained by the third neural network when inputting the corresponding input sample.
[0094] In the technical solution provided in this embodiment, the first calibration sample in the training sample of the first neural network is calculated by the corresponding second calibration sample of the corresponding input sample and the corresponding reference sample, which can make the performance of the first neural network closer to the performance of the third neural network.
[0095] As another embodiment of the present application, when determining the set training samples, the method includes:
[0096] Performing a fusion process on the second calibration sample corresponding to the input sample and the corresponding reference sample based on a set weight value to obtain a corresponding first calibration sample; the weight value of the second calibration sample corresponding to the input sample is greater than the weight value of the corresponding reference sample.
[0097] Here, the sum of the weight value of the second calibration sample corresponding to the input sample and the weight value of the corresponding reference sample is 1.
[0098] To improve the accuracy or precision of the first neural network, the weight value of the second calibration sample corresponding to the input sample is greater than the weight value of the corresponding reference sample. For example, the weight value of the second calibration sample corresponding to the input sample is 0.8, and the weight value of the reference sample corresponding to the input sample is 0.2.
[0099] As another embodiment of the present application, Figure 4 shows a schematic implementation flow diagram of constructing an initial first neural network in the model pruning method provided by the embodiments of the present application. Referring to Figure 4 When constructing the initial first neural network, the method includes:
[0100] S201: Input the third structure information of the third neural network into the second neural network to obtain the fourth structure information output by the second neural network.
[0101] The electronic device initializes the second neural network. When the third neural network is trained, the electronic device inputs the third structure information of the third neural network into the second neural network to obtain the fourth structure information output by the second neural network.
[0102] S202: Construct an initial first neural network based on the fourth structure information.
[0103] Here, when the third structure information and the fourth structure information are different, an initial first neural network is constructed based on the fourth structure information. In practical applications, an initial first neural network can be constructed based on information such as the topological structure and weight parameters in the fourth structure information; it is also possible to copy the third neural network and perform pruning processing on the copied neural network based on the fourth structure information to obtain the initial first neural network.
[0104] When the third structure information and the fourth structure information are the same, the third neural network is copied to obtain the initial first neural network.
[0105] In the technical solution provided by this embodiment, the second neural network can construct an initial first neural network with performance close to that of the third neural network based on the structural information of the third neural network.
[0106] As another embodiment of the present application, Figure 5 shows a schematic implementation flowchart of determining the second structural information in the model pruning method provided by another embodiment of the present application. Refer to Figure 5 The calculation of the second structural information by the second neural network includes:
[0107] S301: Test the current first neural network with at least one set test sample to obtain a test result corresponding to each test sample in the at least one set test sample; the test result represents a set index value corresponding to the corresponding test sample.
[0108] When the current first neural network is trained, the electronic device inputs at least one set test sample into the current first neural network to obtain a test result corresponding to each test sample in the at least one set test sample output by the current first neural network.
[0109] Here, the test result corresponding to the test sample can be the output result obtained by the first neural network when inputting the corresponding test sample, or the relevant data obtained during the test of the first neural network.
[0110] In practical applications, the training samples and test samples corresponding to the current first neural network are different. During the test of the current first neural network, each test sample is used to perform a test once.
[0111] In one embodiment, the set index value includes at least one of the following:
[0112] The first index value; the first index value represents the performance parameter of the current first neural network;
[0113] At least one second index value; the second index value represents the cost parameter of the current first neural network when processing each test sample in the at least one set test sample; where
[0114] The second index value includes at least one of the following:
[0115] The amount of computation of the current first neural network;
[0116] The bandwidth of the current first neural network; the bandwidth represents the total amount of data transmitted by the current first neural network per unit time;
[0117] The storage space occupied by the weight parameters of the current first neural network;
[0118] The execution time of the current first neural network; the execution time represents the duration corresponding to one round of testing of the current first neural network.
[0119] The execution power consumption of the current first neural network; the execution power consumption represents the power consumption corresponding to one round of testing of the current first neural network.
[0120] Among them, the amount of computation of the current first neural network can be measured by the number of multiply-accumulate operations (MAC, Multiply Accumulate) executed by the first neural network.
[0121] S302: Based on the test results corresponding to each test sample in the at least one set test sample, use a set loss function to calculate the loss value corresponding to the second neural network.
[0122] The electronic device can determine the set index value corresponding to the test sample based on the test result corresponding to the test sample, and use a set loss function to calculate the loss value corresponding to the second neural network.
[0123] In practical applications, when the set index value includes a first index value and at least one second index value, the expression of the set loss function can be: Loss = cost – λ × effect.
[0124] Among them, Loss represents the loss value; cost represents at least one second index value of the current first neural network, and the higher the value of cost, the greater the cost of the first neural network; effect represents the first index value of the current first neural network, and the higher the value of effect, the better the performance of the first neural network. λ is a set constant. Here, when the set index value includes at least two second index values, cost corresponds to the total second index value. The electronic device can perform a weighted process on at least two second index values to obtain the total second index value.
[0125] In one embodiment, the expression of the set loss function can also be:
[0126] Loss n =(cost n –λ×effect n )-(cost n-1 –λ×effect n-1 )。
[0127] Among them, Loss n represents the loss value corresponding to the nth test of the current first neural network; cost nCharacterize the second metric value corresponding to the first neural network at the current nth test; effect n Characterize the first metric value corresponding to the first neural network at the current nth test; cost n-1 Characterize the second metric value corresponding to the first neural network at the current (n - 1)th test; effect n-1 Characterize the first metric value corresponding to the first neural network at the current (n - 1)th test. n is an integer greater than or equal to 1.
[0128] S303: Update the weight parameters of the second neural network according to the calculated loss value.
[0129] When the loss value of the second neural network is calculated, the electronic device updates the weight parameters of the second neural network based on the calculated loss value. Among them, the electronic device performs backpropagation of the calculated loss value in the second neural network. During the process of backpropagating the calculated loss value to each layer of the second neural network, the gradient of the first loss function is calculated according to the loss value, and the weight parameters backpropagated to the current layer are updated along the descending direction of the gradient.
[0130] S304: When the weight parameters are updated, input the first structure information of the current first neural network into the second neural network to obtain the second structure information output by the second neural network.
[0131] In practical applications, the first neural network and the third neural network can both be image super-resolution networks. As Figure 6 shown, the first neural network includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. The first neural network is used to perform super-resolution processing on the first image and output a second image; among them, the resolution of the second image is greater than that of the first image. The number of convolution kernels (or output channels) of the first convolutional layer is m1, the number of convolution kernels of the second convolutional layer is m2, and the number of convolution kernels of the third convolutional layer is 1. The convolution kernels of all convolutional layers can be 3×3.
[0132] Since the output channels of the third convolutional layer in the first network are restricted by functions and must be 1, the second neural network is used to optimize the output channel numbers m1 of the first convolutional layer and m2 of the second convolutional layer, thereby optimizing the first metric value and the second metric value of the first neural network. In practical applications, m1 = 64 and m2 = 32 in the first structure information; m′1 and m′2 in the second structure information are both positive integers. m′1 and m′2 satisfy any one of the following: m′1 is less than m1; m′2 is less than m2.
[0133] The first metric value of the first neural network at least includes the peak signal-to-noise ratio (PSNR) of the second image; the second metric value of the first neural network at least includes the number of weights of the neural network, and the number of weights is calculated based on the size of the convolutional kernel and the number of weight channels. The size of the convolutional kernel characterizes the length and width of the convolutional kernel, and can also characterize the depth of the convolutional kernel.
[0134] Among them, the number of weights quantity of the first neural network = 3×3×(m1 + m1×m2 + m2). Among them, the number of weight channels of the first neural network corresponds to (m1 + m1×m2 + m2).
[0135] Correspondingly, the expression of the loss function of the second neural network can be:
[0136] Loss n =(PSNR n – λ×quantity n ) - (PSNR n-1 – λ×quantity n-1 ).
[0137] PSNR n represents the peak signal-to-noise ratio corresponding to the nth test of the first neural network, and quantity n represents the number of weights corresponding to the nth test of the first neural network; PSNR n-1 represents the peak signal-to-noise ratio corresponding to the (n - 1)th test of the first neural network, and quantity n-1 represents the number of weights corresponding to the (n - 1)th test of the first neural network. n is an integer greater than or equal to 1.
[0138] Among them, when n = 1, PSNR n-1 and quantity n-1 can both be zero.
[0139] The output channel numbers m′1 of the first convolutional layer and m′2 of the second convolutional layer included in the second structural information of the first output input by the second neural network are both positive integers.
[0140] In the technical solution provided in this embodiment, a loss function corresponding to the second neural network is calculated based on the first index value and the second index value of the first neural network, and the weight parameters of the second neural network are updated based on the calculated loss function. When the weight parameters of the second neural network are updated, the first structure information of the first neural network is input into the second neural network, and the second structure information output by the second neural network is obtained. Since the first index value and the second index value of the first neural network are considered during the training of the second neural network, the first neural network updated based on the second structure information can achieve a balance between performance and cost.
[0141] As another embodiment of the present application, in S102, the updating the structure of the current first neural network based on the second structure information includes at least one of the following:
[0142] Updating the topological structure of the current first neural network based on the topological structure included in the second structure information; wherein, the topological structure can represent information such as the number of layers included in the first neural network, the name of each layer, and the topological structure of each layer; the topological structure of the corresponding layer can represent the number of neurons included in the corresponding layer and the connection relationship between the neurons, etc.;
[0143] Updating the weight channels of the corresponding layer of the current first neural network based on the number of weight channels of the corresponding layer included in the second structure information; the number of weight channels represents the number of input channels and the number of output channels of the corresponding layer;
[0144] Updating the weight values of the corresponding layer of the current first neural network based on the weight precision of the corresponding layer included in the second structure information; the weight precision represents the number of bits occupied by the weight values of the corresponding layer;
[0145] Updating the output precision of the activation function of the corresponding layer of the current first neural network based on the output precision of the activation function of the corresponding layer included in the second structure information; the output precision of the activation function represents the precision of the output result of the corresponding activation function;
[0146] Updating the weight values of the corresponding layer of the current first neural network based on the pruning threshold of the weight of the corresponding layer included in the second structure information; wherein, when the absolute value of the weight value of the corresponding layer is less than the corresponding pruning threshold, the corresponding weight value is set to zero; when the absolute value of the weight value of the corresponding layer is greater than or equal to the corresponding pruning threshold, the weight value of the corresponding layer remains unchanged.
[0147] In practical applications, the second structural information can be represented in the form of an array, a matrix, etc. For example, a first array is used to represent the topological structure of each layer in the first neural network; a second array is used to represent the number of weight channels in each layer of the first neural network; a third array is used to represent the weight precision of each layer in the first neural network; a fourth array is used to represent the output precision of the activation function of each layer in the first neural network; a fifth array is used to represent the pruning threshold of the weights of each layer in the first neural network. Among them,
[0148] When the number of weight channels of the corresponding layer in the first neural network is less than or equal to zero, it indicates that the electronic device deletes this layer when updating the structural information of the first neural network. Among them, the number of weight channels being less than or equal to zero means that both the number of input channels and the number of output channels of the corresponding layer are less than or equal to zero.
[0149] When the weight precision included in the second structural information of the first neural network indicates that the number of bits occupied by the weight value of the corresponding layer is k, when the electronic device performs structural update on the current first neural network, it updates the weight value of the corresponding layer based on the absolute value of the weight value w of the corresponding layer. Among them, when the absolute value of the weight value of the corresponding layer of the current first neural network is less than 1, the weight value of the corresponding layer of the current first neural network is updated to int(w×2 n ) / 2 n , here, int(w×2 n ) represents truncating and rounding w×2 n . When the absolute value of the weight value of the corresponding layer of the current first neural network is greater than or equal to 1, the weight value of the corresponding layer of the current first neural network is updated to int(w / 2 n )×2 n , int(w / 2 n ) represents truncating and rounding w / 2 n .
[0150] It should be noted that the weight precision of the corresponding layer of the first neural network is used to reduce the number of bits occupied by the weight value of the corresponding layer, thereby reducing the storage capacity and computational amount of the first neural network. For example, the weight precision in the first structural information indicates that the weight value of the corresponding layer occupies 32 bits (bit), and the weight value in the second structural information indicates that the weight value of the corresponding layer occupies 8 bits (that is, the value corresponding to the high eight bits is retained). After performing structural update on the first neural network based on the second structural information, the storage space occupied by the first neural network is reduced by three - quarters, and the hardware multiplier resources consumed by the operation are also significantly reduced.
[0151] When the output precision of the activation function included in the second structure information of the first neural network represents the number of bits occupied by the output result of the activation function of the corresponding layer as j, after the electronic device updates the structure of the current first neural network, it updates the output precision of the activation function of the corresponding layer based on the absolute value of the output result of the activation function of the corresponding layer. Among them, when the absolute value of the output result q of the activation function of the corresponding layer of the current first neural network is less than 1, the output result of the activation function of the corresponding layer of the current first neural network is updated to int(q×2^j) / 2^j, where int(q×2^j) represents truncating and taking the integer of q×2^j. When the absolute value of the output result q of the activation function of the corresponding layer of the current first neural network is greater than or equal to 1, the output result of the activation function of the corresponding layer of the current first neural network is updated to int(q / 2^j)×2^j; here, int(q / 2^j) represents truncating and taking the integer of q / 2^j.
[0152] It should be noted that by reducing the output precision of the activation function of the first neural network, the bandwidth required by the first neural network can be reduced, and at the same time, the computational amount of the first neural network can also be reduced.
[0153] Since when the absolute value of the weight value of the corresponding layer in the first neural network is less than the corresponding pruning threshold, the corresponding weight value is set to zero, the pruning threshold of the weight can increase the proportion of zeros in the weight value, thereby increasing the compression rate of the weight.
[0154] To implement the method of the embodiments of the present application, the embodiments of the present application further provide a model pruning device, which is set on an electronic device such as a terminal or a server, as Figure 7 shown, the model pruning device includes:
[0155] A calculation unit 71, configured to calculate second structure information through a second neural network based on the first structure information of the current first neural network and the corresponding set index value; wherein, the current first neural network is trained based on a set training sample; the initial first neural network is constructed based on the third structure information of the third neural network; the corresponding set index value is used to update the weight parameters of the second neural network, and the second neural network is configured to output the second structure information based on the input first structure information after updating the weight parameters;
[0156] An update unit 72, configured to update the structure of the current first neural network based on the second structure information;
[0157] A determination unit 73, configured to determine the first neural network with the updated structure as the model pruning result corresponding to the third neural network when the second neural network reaches the set convergence condition; wherein,
[0158] When constructing the initial first neural network or updating the structure of the first neural network, the first neural network is trained based on the set training samples.
[0159] In one embodiment, when constructing the initial first neural network, the computing unit 71 is further configured to:
[0160] Input the third structure information of the third neural network into the second neural network to obtain the fourth structure information output by the second neural network;
[0161] Construct the initial first neural network based on the fourth structure information.
[0162] In one embodiment, when the computing unit 71 calculates the second structure information through the second neural network, it is configured to:
[0163] Test the current first neural network with at least one set test sample to obtain the test result corresponding to each test sample in the at least one set test sample; the test result represents the set index value corresponding to the corresponding test sample;
[0164] Based on the test results corresponding to each test sample in the at least one set test sample, calculate the loss value corresponding to the second neural network by using a set loss function;
[0165] Update the weight parameters of the second neural network according to the calculated loss value;
[0166] When the weight parameters of the second neural network are updated, input the first structure information of the current first neural network into the second neural network to obtain the second structure information output by the second neural network.
[0167] In one embodiment, the structure update of the current first neural network based on the second structure information includes at least one of the following:
[0168] Update the topological structure of the current first neural network based on the topological structure included in the second structure information;
[0169] Update the weight channels of the corresponding layer of the current first neural network based on the number of weight channels of the corresponding layer included in the second structure information; the number of weight channels represents the number of input channels and output channels of the corresponding layer;
[0170] Update the weight values of the corresponding layer of the current first neural network based on the weight precision of the corresponding layer included in the second structure information; the weight precision represents the number of bits occupied by the weight values of the corresponding layer;
[0171] Update the output precision of the activation function of the corresponding layer of the current first neural network based on the output precision of the activation function of the corresponding layer included in the second structure information; the output precision of the activation function characterizes the precision of the output result of the corresponding activation function.
[0172] Update the weight value of the corresponding layer of the current first neural network based on the pruning threshold of the weight of the corresponding layer included in the second structure information; wherein, when the absolute value of the weight value of the corresponding layer is less than the corresponding pruning threshold, set the corresponding weight value to zero; when the absolute value of the weight value of the corresponding layer is greater than or equal to the corresponding pruning threshold, keep the weight value of the corresponding layer unchanged.
[0173] In one embodiment, the set index value includes at least one of the following:
[0174] The first index value; the first index value characterizes the performance parameter of the current first neural network.
[0175] The second index value; the second index value characterizes the cost parameter when the current first neural network processes each of the at least one set test samples; wherein,
[0176] The second index value includes at least one of the following:
[0177] The amount of computation of the current first neural network;
[0178] The bandwidth of the current first neural network; the bandwidth characterizes the total amount of data transmitted by the current first neural network per unit time.
[0179] The storage space occupied by the weight parameters of the current first neural network;
[0180] The execution time of the current first neural network; the execution time characterizes the duration corresponding to one round of testing performed by the current first neural network.
[0181] The execution power consumption of the current first neural network; the execution power consumption characterizes the power consumption corresponding to one round of testing performed by the current first neural network.
[0182] In one embodiment, when the first neural network is initially constructed or the structure of the first neural network is updated, when training the first neural network based on the set training samples, the determination unit 73 is used to determine the set training samples; wherein,
[0183] The set training samples include at least one sample pair; each sample pair in the at least one sample pair includes an input sample and a corresponding first calibration sample; wherein, the first calibration sample is used to compare with the output result obtained by the first neural network when inputting the corresponding input sample; the first calibration sample is determined by the corresponding second calibration sample and the corresponding reference sample of the corresponding input sample; the second calibration sample represents the calibration sample corresponding to the input sample of the third neural network; the reference sample represents the output result obtained by the third neural network when inputting the corresponding input sample.
[0184] In an embodiment, the determining unit 73 is configured to: perform a fusion process on the second calibration sample corresponding to the input sample and the corresponding reference sample based on the set weight value to obtain the corresponding first calibration sample; the weight value of the second calibration sample corresponding to the input sample is greater than the weight value of the corresponding reference sample.
[0185] In practical applications, each unit included in the model pruning device can be implemented by a processor in the model pruning device. Of course, the processor needs to run the program stored in the memory to implement the functions of the above-mentioned program modules.
[0186] It should be noted that: when the model pruning device provided in the above embodiment performs model pruning, only the above-mentioned division of each program module is used as an example. In practical applications, the above-mentioned processing can be allocated to different program modules according to needs, that is, the internal structure of the model pruning device is divided into different program modules to complete all or part of the above-described processing. In addition, the model pruning device provided in the above embodiment and the model pruning method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0187] Based on the hardware implementation of the above program module, and in order to implement the method of the embodiments of the present application, the embodiments of the present application further provide an electronic device. Figure 8 It is a schematic diagram of the hardware composition structure of the electronic device of the embodiments of the present application, as Figure 8 shown, the electronic device includes:
[0188] A communication interface 1, capable of interacting with other devices such as network devices for information.
[0189] A processor 2, connected to the communication interface 1 to achieve information interaction with other devices, and is used to execute the model pruning method provided by the above one or more technical solutions when running a computer program. And the computer program is stored on the memory 3.
[0190] Of course, in practical applications, the various components in the electronic device are coupled together through the bus system 4. It can be understood that the bus system 4 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 8 all kinds of buses are labeled as the bus system 4.
[0191] The memory 3 in the embodiments of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device.
[0192] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, Sync Link Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 3 described in the embodiments of the present application is intended to include, but is not limited to, these and any other suitable types of memories.
[0193] The method disclosed in the embodiments of the present application above can be applied to the processor 2 or implemented by the processor 2. The processor 2 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 2 or by instructions in the form of software. The above-mentioned processor 2 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, which is located in the memory 3. The processor 2 reads the program in the memory 3 and combines its hardware to complete the steps of the foregoing method.
[0194] When the processor 2 executes the program, it implements the corresponding processes in the various methods of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0195] In an exemplary embodiment, the embodiments of the present application also provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 3 including a stored computer program. The above computer program can be executed by the processor 2 to complete the steps of the foregoing method. The computer-readable storage medium may be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0196] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces. The indirect coupling or communication connection of the devices or units may be electrical, mechanical, or other forms.
[0197] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0198] In addition, each functional unit in the embodiments of the present application may all be integrated into one processing module, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0199] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0200] It should be noted that the technical solutions described in the embodiments of the present application can be arbitrarily combined without conflict.
[0201] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A model pruning method, characterized in that, it includes: Based on the first structure information of the current first neural network and the corresponding set index value, calculate the second structure information through the second neural network; wherein, the initial first neural network is constructed based on the third structure information of the third neural network; the corresponding set index value is used to update the weight parameters of the second neural network, and the second neural network is used to output the second structure information based on the input first structure information after updating the weight parameters; Update the structure of the current first neural network based on the second structure information; When the second neural network reaches the set convergence condition, determine the first neural network with the updated structure as the model pruning result corresponding to the third neural network; wherein, When the first neural network is initially constructed or the structure of the first neural network is updated, train the first neural network based on the set training samples; The first neural network and the third neural network are applied to super-resolution processing of images; The updating the structure of the current first neural network based on the second structure information includes at least one of the following: Update the topological structure of the current first neural network based on the topological structure included in the second structure information; Update the weight channels of the corresponding layer of the current first neural network based on the number of weight channels of the corresponding layer included in the second structure information; the number of weight channels represents the number of input channels and output channels of the corresponding layer; Update the weight values of the corresponding layer of the current first neural network based on the weight precision of the corresponding layer included in the second structure information; the weight precision represents the number of bits occupied by the weight values of the corresponding layer; Update the output precision of the activation function of the corresponding layer of the current first neural network based on the output precision of the activation function of the corresponding layer included in the second structure information; the output precision of the activation function represents the precision of the output result of the corresponding activation function; Update the weight values of the corresponding layer of the current first neural network based on the pruning threshold of the weight of the corresponding layer included in the second structure information; wherein, when the absolute value of the weight value of the corresponding layer is less than the corresponding pruning threshold, set the corresponding weight value to zero; when the absolute value of the weight value of the corresponding layer is greater than or equal to the corresponding pruning threshold, keep the weight value of the corresponding layer unchanged.
2. The method according to claim 1, characterized in that, When constructing the initial first neural network, the method includes: Input the third structure information of the third neural network into the second neural network to obtain the fourth structure information output by the second neural network; Construct an initial first neural network based on the fourth structure information.
3. The method according to claim 1, characterized in that, The calculating the second structure information through the second neural network includes: Test the current first neural network with at least one set test sample to obtain the test result corresponding to each test sample in the at least one set test sample; the test result represents the set index value corresponding to the corresponding test sample; Based on the test results corresponding to each test sample in the at least one set test sample, calculate the loss value corresponding to the second neural network by using a set loss function; Update the weight parameters of the second neural network according to the calculated loss value; When the weight parameters of the second neural network are updated, input the first structure information of the current first neural network into the second neural network to obtain the second structure information output by the second neural network.
4. The method according to claim 1 or 3, wherein, the set index value includes at least one of the following: a first index value; the first index value characterizes the performance parameter of the current first neural network; at least one second index value; the second index value characterizes the cost parameter when the current first neural network processes each test sample in the at least one set test sample.
5. The method according to claim 1, wherein, when training the first neural network based on the set training samples, the method includes: determine the set training samples; wherein, the set training samples include at least one sample pair; each sample pair in the at least one sample pair includes an input sample and a corresponding first calibration sample; wherein, the first calibration sample is used to compare with the output result obtained by the first neural network when inputting the corresponding input sample; the first calibration sample is determined by the corresponding second calibration sample and the corresponding reference sample of the corresponding input sample; the second calibration sample characterizes the calibration sample corresponding to the input sample of the third neural network; the reference sample characterizes the output result obtained by the third neural network when inputting the corresponding input sample.
6. The method according to claim 5, wherein, when determining the set training samples, the method includes: perform a fusion process on the second calibration sample and the corresponding reference sample corresponding to the input sample based on a set weight value to obtain the corresponding first calibration sample; the weight value of the second calibration sample corresponding to the input sample is greater than the weight value of the corresponding reference sample.
7. A model pruning device, wherein, includes: a calculation unit, configured to calculate second structure information through a second neural network based on the first structure information of the current first neural network and the corresponding set index value; wherein, the initial first neural network is constructed based on the third structure information of the third neural network; the corresponding set index value is used to update the weight parameters of the second neural network, and the second neural network is configured to output the second structure information based on the input first structure information after updating the weight parameters; an update unit, configured to perform a structure update on the current first neural network based on the second structure information; a determination unit, configured to determine the first neural network with the structure updated as the model pruning result corresponding to the third neural network when the second neural network reaches a set convergence condition; wherein, When constructing the initial first neural network or updating the structure of the first neural network, train the first neural network based on the set training samples; the first neural network and the third neural network are applied to perform super-resolution processing on images. The updating unit is configured to update the structure of the current first neural network based on the second structure information, including at least one of the following: Update the topological structure of the current first neural network based on the topological structure included in the second structure information; Update the weight channels of the corresponding layer of the current first neural network based on the number of weight channels of the corresponding layer included in the second structure information; the number of weight channels represents the number of input channels and the number of output channels of the corresponding layer; Update the weight values of the corresponding layer of the current first neural network based on the weight precision of the corresponding layer included in the second structure information; the weight precision represents the number of bits occupied by the weight values of the corresponding layer; Update the output precision of the activation function of the corresponding layer of the current first neural network based on the output precision of the activation function of the corresponding layer included in the second structure information; the output precision of the activation function represents the precision of the output result of the corresponding activation function; Update the weight values of the corresponding layer of the current first neural network based on the pruning threshold of the weight of the corresponding layer included in the second structure information; wherein, when the absolute value of the weight value of the corresponding layer is less than the corresponding pruning threshold, set the corresponding weight value to zero; when the absolute value of the weight value of the corresponding layer is greater than or equal to the corresponding pruning threshold, keep the weight value of the corresponding layer unchanged.
8. An electronic device Characterized in that It includes: A processor and a memory for storing a computer program that can run on the processor, wherein, when the processor is used to run the computer program, it executes the steps of the method according to any one of claims 1 to 6.
9. A storage medium, on which a computer program is stored, Characterized in that When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Convolutional neural network training method and device
CN109635927A
Pruning method based on double-flow network
CN111079691A