Convolutional Neural Network Compression Method, Device and Electronic Device
By introducing sensitivity as the criterion for judging the importance of connection in the convolutional neural network and pruning according to the preset sparseness, the problem of low compression efficiency of convolutional neural networks in the prior art is solved, and an efficient model compression and training process is realized.
Patent Information
- Application Number
- CN202111645899.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-30
AI Technical Summary
Existing convolutional neural networks are hindered by high computational volume in practical applications, and need to reduce the model size and calculation operations through pruning strategies, but this process is time-consuming and costly.
By obtaining the target training sample set of the target application scenario, the neural network model weight is initialized using the variance scaling method, the connection sensitivity is calculated, and the model is pruned according to the preset sparseness, and finally the sparse neural network model is trained using the target training sample set.
While ensuring the accuracy of the model after compression, the efficiency of model compression is improved, and the iterative optimization process of repeated training-pruning-fine-tuning is avoided.
Smart Images

Figure CN114330690B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of convolutional neural network compression, and in particular, to a method, device, and electronic device for compressing a convolutional neural network. Background Art
[0002] Currently, existing convolutional neural networks are greatly hindered by high computational complexity in practical applications. Different pruning strategies need to be adopted to reduce the model size, reduce the memory occupation during model operation, and at the same time reduce the number of computational operations without affecting the accuracy. Model pruning is usually an iterative optimization process of repeated training - pruning - fine-tuning. Although a compressed model with high accuracy can be obtained, this process takes a lot of time and has a high time cost. Summary of the Invention
[0003] In view of this, an object of the present invention is to provide a method, device, and electronic device for compressing a convolutional neural network to improve the efficiency of model compression while ensuring the accuracy of the compressed model.
[0004] In a first aspect, an embodiment of the present invention provides a method for compressing a convolutional neural network. The method includes: obtaining a target training sample set for a target application scenario; wherein the target training sample set is determined based on an initial training sample set of a neural network model to be compressed corresponding to the target application scenario; initializing weights of the neural network model to be compressed by using a variance scaling method to obtain an initial weight vector of the neural network model to be compressed; using a preset sparsity as a constraint condition to determine a weight optimization vector corresponding to the initial weight vector, and calculating sensitivities of all connections in the neural network model to be compressed according to the weight optimization vector; wherein the sensitivity is used to represent the importance degree of connections of each network layer in the neural network model to be compressed;
[0005] Pruning the neural network model to be compressed according to the preset sparsity and the sensitivity to obtain a sparse neural network model; wherein weights of the sparse neural network model are determined according to the preset sparsity and the sensitivity;
[0006] Training the sparse neural network model by using the target training sample set until a trained target neural network model is obtained; wherein the target neural network model is used to process data corresponding to the target application scenario.
[0007] Combined with the first aspect, an embodiment of the present invention provides a first possible implementation manner of the first aspect. The step of obtaining the target sample set includes: sampling the initial training sample set to obtain the target sample set where D represents the target sample set, xi represents the i-th sample, y i represents the label corresponding to the i-th sample, i represents the current batch, and n represents the number of samples in the target sample set.
[0008] Combined with the first aspect, the embodiment of the present invention provides a second possible implementation manner of the first aspect. Among them, the method further includes: defining the pruning of the neural network model to be compressed as a constrained optimization problem of the following formula:
[0009]
[0010] s, t.w ∈ R m , c ∈ {0, 1} m , ||c|| 0 ≤ k
[0011] Among them, L(·) represents the overall loss function, l(·) represents the partial loss function, ⊙ represents the Hadamard product, c represents the weight optimization vector, w represents the initial weight vector, ∥·∥ 0 represents the standard L 0 norm, m represents the total number of parameters of the neural network model to be compressed, {0, 1} m represents an m-dimensional vector with elements only 0 and 1, and k represents the preset sparsity.
[0012] Combined with the first aspect, the embodiment of the present invention provides a third possible implementation manner of the first aspect. Among them, taking the preset sparsity as a constraint condition, determining the weight optimization vector corresponding to the initial weight vector, and calculating the sensitivity of all connections in the neural network model to be compressed according to the weight optimization vector includes:
[0013] For each connection in the neural network model to be compressed, use the following formula to calculate the derivative of the overall loss function with respect to the weight optimization vector to approximately represent the impact of removing the connection on the loss of the neural network model to be compressed:
[0014]
[0015] s.t.w ∈ R m , c ∈ {0, 1} m , Hc|| 0 ≤ k
[0016] Among them, g j (w; D) represents the derivative value of the overall loss function corresponding to connection j with respect to the weight optimization vector, and e j represents the indicator vector of connection j;
[0017] According to the derivative values calculated for each connection, the sensitivity of each connection in the neural network model to be compressed is calculated using the following formula:
[0018]
[0019] where s j represents the sensitivity of connection j, |g j (w; D)| represents the absolute value of the derivative value corresponding to connection j, and N represents the number of connections in the neural network model to be compressed.
[0020] Combined with the first aspect, an embodiment of the present invention provides a fourth possible implementation manner of the first aspect. Among them, the step of pruning the neural network model to be compressed according to the preset sparsity and the sensitivity to obtain a sparse neural network model includes: sorting all the connections in the neural network model to be compressed in descending order of the sensitivity, and retaining the first k connections in the sorting result to obtain a first sparse neural network model; according to the preset sparsity and the sorting result, performing weighted processing on the connections of each network layer in the first sparse neural network model to obtain the sparse neural network model.
[0021] Combined with the first aspect, an embodiment of the present invention provides a fifth possible implementation manner of the first aspect. Among them, the step of performing weighted processing on the connections of each network layer in the first sparse neural network model according to the preset sparsity and the sorting result to obtain the sparse neural network model includes:
[0022] The following formula is used to calculate the optimized weight value corresponding to each of the first k connections:
[0023]
[0024] where w i ’ represents the optimized weight value of the i-th connection;
[0025] Assign the optimized weight value to each of the first k connections to obtain the sparse neural network model.
[0026] In combination with the first aspect, an embodiment of the present invention provides a sixth possible implementation manner of the first aspect. Among them, the step of training the sparse neural network model using the target training sample set until the trained target neural network model is obtained includes: inputting all samples in the target training sample set into the sparse neural network model; calculating the loss function value according to the prediction result of the sparse neural network model and the labels corresponding to all samples in the target training sample set, and stopping training until the number of iterations exceeds a preset number or the loss function value is less than a preset value, so as to obtain the trained target neural network model.
[0027] In a second aspect, an embodiment of the present invention further provides a convolutional neural network compression device. The device includes: a sample acquisition module, configured to acquire a target training sample set for a target application scenario; wherein, the target training sample set is determined based on an initial training sample set of a neural network model to be compressed corresponding to the target application scenario; an initialization module, configured to initialize the weights of the neural network model to be compressed by using a variance scaling method to obtain an initial weight vector of the neural network model to be compressed; a sensitivity calculation module, configured to determine a weight optimization vector corresponding to the initial weight vector with a preset sparsity as a constraint condition, and calculate the sensitivity of all connections in the neural network model to be compressed according to the weight optimization vector; wherein, the sensitivity is used to characterize the importance degree of the connections of each network layer in the neural network model to be compressed; a pruning module, configured to prune the neural network model to be compressed according to the preset sparsity and the sensitivity to obtain a sparse neural network model; wherein, the weights of the sparse neural network model are determined according to the preset sparsity and the sensitivity; a training module, configured to train the sparse neural network model using the target training sample set until the trained target neural network model is obtained; wherein, the target neural network model is used to process data corresponding to the target application scenario.
[0028] In a third aspect, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the above-mentioned convolutional neural network compression method.
[0029] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processing device, the steps of the above-mentioned convolutional neural network compression method are executed.
[0030] The convolutional neural network compression method, device, and electronic device provided by the embodiments of the present invention. The method includes: obtaining a target training sample set for a target application scenario; initializing the weights of the neural network model to be compressed using the variance scaling method to obtain an initial weight vector of the neural network model to be compressed; using a preset sparsity as a constraint condition, determining a weight optimization vector corresponding to the initial weight vector, and calculating the sensitivity of all connections in the neural network model to be compressed according to the weight optimization vector; pruning the neural network model to be compressed according to the preset sparsity and sensitivity to obtain a sparse neural network model; training the sparse neural network model using the target training sample set until a trained target neural network model is obtained. By adopting the above technology, after initializing the weights of the neural network model, sensitivity is introduced as a criterion for evaluating the importance of connections, so that redundant connections in the neural network model to be compressed can be identified according to sensitivity before training, and the redundant connections can be pruned according to the required sparsity level. Then, the standard model training process is used to train the pruned sparse neural network model. This operation method can avoid the iterative optimization process of repeated training-pruning-fine-tuning, and thus improve the compression efficiency of the neural network model while ensuring the performance of the pruned neural network model.
[0031] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.
[0032] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically provides preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 It is a schematic flowchart of a convolutional neural network compression method provided by an embodiment of the present invention;
[0035] Figure 2 It is a schematic structural diagram of a convolutional neural network compression device provided by an embodiment of the present invention;
[0036] Figure 3 It is a schematic structural diagram of another convolutional neural network compression device provided by an embodiment of the present invention;
[0037] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] Currently, existing convolutional neural networks are greatly hindered by high computational complexity in practical applications. Different pruning strategies need to be adopted to reduce the model size, reduce the memory occupancy during model operation, and at the same time reduce the number of computational operations without affecting the accuracy. Model pruning is usually an iterative optimization process of repeated training - pruning - fine-tuning. Although a compressed model with high accuracy can be obtained, this process takes a lot of time and has a high time cost.
[0040] Based on this, a convolutional neural network compression method, device, and electronic device provided by an embodiment of the present invention can improve the efficiency of model compression while ensuring the accuracy of the compressed model.
[0041] To facilitate the understanding of this embodiment, a convolutional neural network compression method disclosed in an embodiment of the present invention will be introduced in detail first. Refer to Figure 1 The flowchart of a convolutional neural network compression method shown, and this method may include the following steps:
[0042] Step S102, obtaining a target training sample set for a target application scenario; wherein, the target training sample set is determined based on an initial training sample set of a neural network model to be compressed corresponding to the target application scenario.
[0043] The above-mentioned target application scenarios mainly include target recognition (such as face recognition, fingerprint recognition), target tracking (such as vehicle tracking), etc. application scenarios that need to use a neural network model to complete corresponding tasks. Based on this, the neural network model to be compressed can be selected according to the target application scenario. For example, the neural network model to be compressed can adopt an R-CNN model, a Faster-RCNN model, a YOLO model, etc., and this is not limited.
[0044] In the actual application process, for the convenience of operation, the above-mentioned target training sample set can be determined according to the training sample set (i.e., the above-mentioned initial training sample set) used for previously training the neural network model to be compressed. Considering factors such as the model's computing power, the number of model parameters, and the timeliness of the model's processing tasks, the above-mentioned step S102 can specifically adopt one of the following two operation methods: (1) directly determine the above-mentioned initial training sample set as the above-mentioned target training sample set; (2) sample the above-mentioned initial training sample set, and determine the set composed of the sampled samples as the above-mentioned target training sample set.
[0045] Specifically, for the case where the neural network model to be compressed is a lightweight model (such as yolov5s, etc.), when the model's computing power permits, the entire coco128 sample set (including 128 pictures) can be directly selected as the above-mentioned target training sample set; the voc2007 sample set (including 9963 pictures) can also be sampled and the resulting mini-batch sample set can be determined as the above-mentioned target training sample set. In addition, considering the accuracy of model training, when sampling the above-mentioned initial training sample set, uniform sampling can be adopted to ensure that the class distribution of the resulting mini-batch sample set is consistent with the class distribution of the initial training sample set, so that the pruned neural network model can learn more comprehensive features.
[0046] As an example, the above-mentioned step S102 can specifically adopt the following operation method:
[0047] Sample the initial training sample set to obtain the target sample set where D represents the target sample set, x i represents the i-th sample, y i represents the label corresponding to the i-th sample, i represents the current batch, and n represents the number of samples in the target sample set.
[0048] Step S104, initialize the weights of the neural network model to be compressed using the variance scaling method to obtain the initial weight vector of the neural network model to be compressed.
[0049] Since a neural network model usually has multiple neurons, to avoid the gradient explosion or disappearance caused by the excessive output of neurons, it is necessary to make the input variance and output variance of each neuron in the neural network model to be compressed as consistent as possible. Based on this, the variance scaling method can be used to initialize the weights of the neural network model to be compressed. The variance scaling method can specifically use methods such as Xavier initialization and He initialization according to actual needs, and this is not limited. After the above-mentioned step S104, the input variance and output variance of each neuron in the entire neural network model to be compressed will be made inconsistent, thereby ensuring the robustness of the neural network model to be compressed during initialization.
[0050] In step S106, taking the preset sparsity as a constraint condition, determine the weight optimization vector corresponding to the initial weight vector, and calculate the sensitivity of all connections in the neural network model to be compressed according to the weight optimization vector; wherein, the sensitivity is used to characterize the importance of the connections of each network layer in the neural network model to be compressed.
[0051] First, preset the desired sparsity (i.e., the above-mentioned preset sparsity) k. k refers to the k weights retained in the model obtained after pruning the neural network model to be compressed (i.e., the number of non-zero elements in the weight matrix). For the target sample set For the neural network model to be compressed, the total number of parameters is m, the initial weight vector of the neural network model to be compressed is w, and a weight optimization vector c with elements of 0 or 1 (which can also be regarded as an indicator variable) is introduced to characterize whether the connections corresponding to each element in w are connected. c and w are matrices with the same number of rows and columns; use the element 1 in c to represent that the connection corresponding to the element in the corresponding position of w is in a connected state (i.e., not pruned), and use the element 0 in c to represent that the connection corresponding to the element in the corresponding position of w is in a disconnected state (i.e., pruned).
[0052] Based on this, the pruning of the above neural network model to be compressed can be defined as the constrained optimization problem in the following formula (1):
[0053]
[0054] s, t.w ∈ R m , c ∈ {0, 1} m , ||c|| 0 ≤ k
[0055] Wherein, L(·) represents the overall loss function, l(·) represents the partial loss function, ⊙ represents the Hadamard product, c represents the weight optimization vector, w represents the initial weight vector, ∥·∥ 0 represents the standard L 0 norm, m represents the total number of parameters of the neural network model to be compressed, {0, 1} m represents an m-dimensional vector with only 0 and 1 elements, and k represents the above-mentioned preset sparsity (i.e., the desired sparsity).
[0056] In the above formula (1), the elements in c can be set to 0 or 1 according to actual needs, and then the elements in the corresponding positions of w (i.e., the weight values in the neural network) are processed through the Hadamard product operation to achieve the purpose of weight sparsification. For example, if the neural network model has only one layer, c = [0, 0, 1, 0, 1], w = [0.13, 0.98, 0.56, 0.42, 1.23], then c⊙w = [0, 0, 0.56, 0, 1.23].
[0057] For simplicity of description, the "loss function" mentioned below refers to the above overall loss function L(·).
[0058] For connection j, the element in c corresponding to connection j can be defined as c j , c j = 1 indicates that the connection in the network is active, and c j = 0 indicates that the connection in the network has been pruned. To measure the impact of connection j on the loss of the model to be compressed, the following formula (2) can be used to calculate the difference between the loss of the model to be compressed when c j = 1 and the loss when c j = 0 while keeping other parameters unchanged:
[0059] ΔL j (w; D) = L(1⊙w; D) - L((1 - e j )⊙w; D) (2)
[0060] where 1 is an m-dimensional vector with all elements being 1, and e j represents the indicator vector of connection j, and e j is an m-dimensional vector with only the element at index j being 1 and the elements at other positions being 0;
[0061] In the above formula (2), 1⊙w can indicate that the neural network model to be compressed is in the fully connected state, and (1 - e j )⊙w can indicate that the neural network model to be compressed is in the state where connection j is pruned. Therefore, ΔL j (w; D) in the above formula (2) indicates the impact of removing connection j on the neural network model to be compressed.
[0062] In the actual application process, considering that the neural network model to be compressed needs to perform m + 1 forward propagations on the input samples, the computational cost involved in performing the corresponding calculations for each connection in the neural network model to be compressed according to the above formula (2) is very large. In addition, since c is an m-dimensional vector with elements that can be represented in binary and only have 0 and 1, the loss function L is not differentiable with respect to c. Based on this, to simplify the calculation and improve the efficiency of the entire model compression process, considering that ΔL j indicates the difference between the loss of the neural network model to be compressed when c j = 1 and the loss when c j = 0, when ΔL j is infinitesimal, the derivative value g j (w; D) of the loss function L with respect to the weight optimization vector c j can be used to replace ΔL jTo approximately characterize the impact of removing connection j on the loss of the neural network model to be compressed, and then optimize the above formula (2) into the following formula (3):
[0063]
[0064] s·t.w∈R m , c∈{0, 1} m , ||c|| 0 ≤k
[0065] where g j (w; D) represents the derivative value of the loss function corresponding to connection j with respect to the weight optimization vector, and e j represents the indicator vector of connection j;
[0066] ΔL j represents the change in the loss of the neural network model to be compressed caused by c j changing from 1 to 0, which can reflect the importance of connection j. For the convenience of forward propagation calculation, ΔL j can be further equivalent to the derivative value of the loss function L with respect to c j . According to the definition of the derivative, the derivative formula is In the above formula (3), ΔL j represents the rate of change of the loss function when c changes from 1 - δ to 1, and c changes to c - δe j .
[0067] Based on the above formula (3), the above step S106 can specifically adopt the following operation method:
[0068] Step 11, for each connection in the above neural network model to be compressed, use the simplified formula (3.1) corresponding to the following formula (3) to calculate the derivative of the loss function with respect to the weight optimization vector to approximately characterize the impact of removing the connection on the loss of the neural network model to be compressed:
[0069]
[0070] s.t.w∈R m , c∈{0, 1} m , ||c|| 0 ≤k
[0071] where g j (w; D) represents the derivative value of the loss function corresponding to connection j with respect to the weight optimization vector, and e j represents the indicator vector of connection j.
[0072] Step 12. According to the derivative value calculated for each connection, calculate the sensitivity (which can also be called the initial sensitivity) of each connection in the neural network model to be compressed using the following formula (4):
[0073]
[0074] where s j represents the sensitivity of connection j, |g j (w; D)| represents the absolute value of the derivative value corresponding to connection j, and N represents the number of connections in the neural network model to be compressed.
[0075] After the above Step 12, the magnitude of the above sensitivity can be used as the criterion for significance, that is: the greater the sensitivity of a connection, essentially it means that the connection has a greater positive or negative impact on the loss of the neural network model to be compressed. Based on this, when pruning the neural network model to be compressed, connections with relatively high sensitivity (i.e., important connections) can be retained while connections with relatively low sensitivity (i.e., redundant connections) can be removed.
[0076] Step S108. Prune the neural network model to be compressed according to the preset sparsity and sensitivity to obtain a sparse neural network model; where the weights of the sparse neural network model are determined according to the preset sparsity and sensitivity.
[0077] For ease of operation, the above Step S108 can specifically adopt the following operation method:
[0078] Step 21. Sort all the connections in the neural network model to be compressed in descending order of the sensitivity, and retain the first k connections in the sorting result to obtain a first sparse neural network model.
[0079] Step 22. According to the preset sparsity and the sorting result, perform weighted processing on the connections of each network layer in the first sparse neural network model to obtain the sparse neural network model.
[0080] Specifically, assign weight values within the range of [0, 1] to the above first k connections in sequence. The following formula can be specifically used to calculate the optimized weight value corresponding to each of the above first k connections:
[0081]
[0082] where w i ’ represents the optimized weight value of the i-th connection.
[0083] Step 23. Assign the above optimized weight value to each of the above first k connections to obtain the sparse neural network model.
[0084] After the above steps 21 to 23, the relatively important connections in the above sparse neural network model are activated, and the less important connections in the above sparse neural network model are suppressed, thereby amplifying the weight differences between the connections in the sparse neural network model, so as to better distinguish different connections according to the magnitude of the optimized weight values.
[0085] Step S110, training the sparse neural network model using the target training sample set until the trained target neural network model is obtained; wherein, the target neural network model is used to process the data corresponding to the target application scenario.
[0086] As a possible implementation manner, all samples in the above target training sample set can be input into the above sparse neural network model; the loss function value is calculated according to the prediction result of the above sparse neural network model and the labels corresponding to all samples in the above target training sample set, and the training is stopped until the number of iterations exceeds the preset number or the loss function value is less than the preset value, and the trained target neural network model is obtained.
[0087] The core idea of the convolutional neural network compression method provided by the embodiments of the present invention is: introducing sensitivity as a significance measure, and then determining the weights or neurons with lower sensitivity as the connections with the least performance degradation after removal, removing the connections with higher sensitivity, and then performing weighted processing on the remaining connections to amplify the differences between different connections. Since the calculation method of sensitivity can avoid the dependence of the significance measure on the loss value, the present invention eliminates the necessity of pre-training the neural network model.
[0088] The convolutional neural network compression method provided by the embodiment of the present invention includes: obtaining a target training sample set of a target application scenario; initializing the weights of the neural network model to be compressed by using the variance scaling method to obtain an initial weight vector of the neural network model to be compressed; using a preset sparsity as a constraint condition, determining a weight optimization vector corresponding to the initial weight vector, and calculating the sensitivity of all connections in the neural network model to be compressed according to the weight optimization vector; pruning the neural network model to be compressed according to the preset sparsity and sensitivity to obtain a sparse neural network model; training the sparse neural network model by using the target training sample set until a trained target neural network model is obtained. By adopting the above technology, after initializing the weights of the neural network model, sensitivity is introduced as a criterion for evaluating the importance of connections, so that redundant connections in the neural network model to be compressed can be identified according to the sensitivity before training, and the redundant connections can be pruned according to the required sparsity level. Then, the standard model training process is used to train the pruned sparse neural network model. This operation method can avoid the iterative optimization process of repeated training-pruning-fine-tuning, thereby improving the compression efficiency of the neural network model while ensuring the performance of the pruned neural network model.
[0089] Based on the above convolutional neural network compression method, the embodiment of the present invention also provides a convolutional neural network compression device, as shown in Figure 2 The device includes:
[0090] A sample acquisition module 21, configured to obtain a target training sample set of a target application scenario; wherein, the target training sample set is determined based on an initial training sample set of the neural network model to be compressed corresponding to the target application scenario.
[0091] An initialization module 22, configured to initialize the weights of the neural network model to be compressed by using the variance scaling method to obtain an initial weight vector of the neural network model to be compressed.
[0092] A sensitivity calculation module 23, configured to use a preset sparsity as a constraint condition, determine a weight optimization vector corresponding to the initial weight vector, and calculate the sensitivity of all connections in the neural network model to be compressed according to the weight optimization vector; wherein, the sensitivity is used to characterize the importance of connections of each network layer in the neural network model to be compressed.
[0093] A pruning module 24, configured to prune the neural network model to be compressed according to the preset sparsity and the sensitivity to obtain a sparse neural network model; wherein, the weights of the sparse neural network model are determined according to the preset sparsity and the sensitivity.
[0094] A training module 25, configured to train the sparse neural network model by using the target training sample set until a trained target neural network model is obtained; wherein, the target neural network model is configured to process data corresponding to the target application scenario.
[0095] The convolutional neural network compression device provided by the embodiment of the present invention, after initializing the weights of the neural network model, introduces sensitivity as a criterion for evaluating the importance of connections, so as to identify redundant connections in the neural network model to be compressed according to sensitivity before training and prune the redundant connections according to the required sparsity level, and then use the standard model training process to train the pruned sparse neural network model. This operation method can avoid the iterative optimization process of training-pruning-fine-tuning, and thus improve the compression efficiency of the neural network model while ensuring the performance of the pruned neural network model.
[0096] Based on the above convolutional neural network compression device, the embodiment of the present invention further provides another convolutional neural network compression device, as shown in Figure 3 shown, the device further includes:
[0097] A definition module 26, configured to define the pruning of the neural network model to be compressed as a constrained optimization problem of the following formula:
[0098]
[0099] s.t.w∈R m , c∈{0, 1} m , ||c|| 0 ≤k
[0100] wherein, L(·) represents the overall loss function, l(·) represents the partial loss function, ⊙ represents the Hadamard product, c represents the weight optimization vector, w represents the initial weight vector, ∥·∥ 0 represents the standard L 0 norm, m represents the total number of parameters of the neural network model to be compressed, {0, 1} m represents an m-dimensional vector with elements only 0 and 1, and k represents the preset sparsity.
[0101] The above sample acquisition module 21 is further configured to:
[0102] Sample the initial training sample set to obtain the target sample set wherein, D represents the target sample set, x i represents the i-th sample, y i represents the label corresponding to the i-th sample, i represents the current batch, and n represents the number of samples in the target sample set.
[0103] The above-mentioned sensitivity calculation module 23 is further configured to:
[0104] For each connection in the neural network model to be compressed, use the following formula to calculate the derivative of the overall loss function with respect to the weight optimization vector to approximately characterize the impact of removing the connection on the loss of the neural network model to be compressed:
[0105]
[0106] s.t.w∈R m , C∈{0, 1} m , ||c|| 0 ≤k
[0107] where, g j (w; D) represents the derivative value of the overall loss function corresponding to connection j with respect to the weight optimization vector at, e j represents the indicator vector of connection j;
[0108] According to the derivative values calculated for each connection, use the following formula to calculate the sensitivity of each connection in the neural network model to be compressed:
[0109]
[0110] where, s j represents the sensitivity of connection j, |g j (w; D)| represents the absolute value of the derivative value corresponding to connection j, and N represents the number of connections in the neural network model to be compressed.
[0111] The above-mentioned pruning module 24 is further configured to: sort all the connections in the neural network model to be compressed in descending order of the sensitivity, and retain the top k connections in the sorting result to obtain the first sparse neural network model; according to the preset sparsity and the sorting result, perform weighted processing on the connections of each network layer in the first sparse neural network model to obtain the sparse neural network model.
[0112] The above-mentioned pruning module 24 is further configured to:
[0113] Use the following formula to calculate the optimized weight value corresponding to each connection in the top k connections:
[0114]
[0115] where, w i ’ represents the optimized weight value of the i-th connection;
[0116] Assign the optimized weight value to each connection in the top k connections to obtain the sparse neural network model.
[0117] The above training module 25 is further configured to: input all samples in the target training sample set into the sparse neural network model; calculate the loss function value according to the prediction result of the sparse neural network model and the labels corresponding to all samples in the target training sample set, and stop training until the number of iterations exceeds a preset number or the loss function value is less than a preset value, so as to obtain a trained target neural network model.
[0118] The device provided by the embodiment of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.
[0119] The embodiment of the present invention provides an electronic device. Refer to Figure 4 As shown, the electronic device 100 includes: a processor 40, a memory 41, a bus 42, and a communication interface 43. The processor 40, the communication interface 43, and the memory 41 are connected through the bus 42. The processor 40 is configured to execute an executable module stored in the memory 41, such as a computer program.
[0120] Among them, the memory 41 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 43 (which can be wired or wireless), a communication connection between the system network element and at least one other network element can be realized, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0121] The bus 42 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 4 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0122] Among them, the memory 41 is used to store a program. After receiving an execution instruction, the processor 40 executes the program. The method executed by the device defined by the flow process disclosed in any embodiment of the foregoing embodiments of the present invention may be applied to the processor 40 or implemented by the processor 40.
[0123] The processor 40 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above convolutional neural network compression method can be completed by the integrated logic circuit of the hardware in the processor 40 or the instructions in the form of software. The above-mentioned processor 40 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the convolutional neural network compression method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 41, and the processor 40 reads the information in the memory 41 and combines its hardware to complete the steps of the above convolutional neural network compression method.
[0124] The computer program product of the convolutional neural network compression method, device, and electronic device provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the convolutional neural network compression method described in the foregoing method embodiments. For the specific implementation, reference can be made to the foregoing method embodiments and will not be elaborated herein.
[0125] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0126] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any technician familiar with the technical field of the present invention can still modify the technical solutions described in the foregoing embodiments, or easily conceive of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for compressing a convolutional neural network, characterized in that, the method comprises: obtaining a target training sample set for a target application scenario; wherein, the target application scenario includes target recognition or target tracking, the target training sample set is determined based on an initial training sample set of a neural network model to be compressed corresponding to the target application scenario, and the target training sample set is pictures; initializing weights of the neural network model to be compressed by using a variance scaling method to obtain an initial weight vector of the neural network model to be compressed; taking a preset sparsity as a constraint condition, determining a weight optimization vector corresponding to the initial weight vector, and calculating sensitivities of all connections in the neural network model to be compressed according to the weight optimization vector; wherein, the sensitivity is used to characterize the importance degree of connections of each network layer in the neural network model to be compressed; pruning the neural network model to be compressed according to the preset sparsity and the sensitivity to obtain a sparse neural network model; wherein, weights of the sparse neural network model are determined according to the preset sparsity and the sensitivity; training the sparse neural network model by using the target training sample set until a trained target neural network model is obtained; wherein, the target neural network model is used to process data corresponding to the target application scenario; the step of pruning the neural network model to be compressed according to the preset sparsity and the sensitivity to obtain a sparse neural network model includes: sorting all connections in the neural network model to be compressed in descending order of the sensitivity, and retaining the first k connections in the sorting result to obtain a first sparse neural network model; calculating an optimized weight value corresponding to each of the first k connections by using the following formula: where, w i ’ represents the optimized weight value of the i-th connection, and k represents the preset sparsity; allocating the optimized weight value to each of the first k connections to obtain the sparse neural network model.
2. The method according to claim 1, characterized in that, the step of obtaining a target sample set includes: Sample the initial training sample set to obtain the target sample set where D represents the target sample set, x i represents the i-th sample, y i represents the label corresponding to the i-th sample, i represents the current batch, and n represents the number of samples in the target sample set.
3. The method according to claim 2, characterized in that, the method further comprises: defining pruning of the neural network model to be compressed as a constrained optimization problem of the following formula: Among them, L(·) represents the overall loss function, l(·) represents the partial loss function, ⊙ represents the Hadamard product, c represents the weight optimization vector, w represents the initial weight vector, and ||·|| 0 represents the standard L 0 norm, m represents the total number of parameters of the neural network model to be compressed, {0, 1} m represents an m-dimensional vector with elements only 0 and 1, and k represents the preset sparsity.
4. The method according to claim 3, characterized in that, the step of taking a preset sparsity as a constraint condition, determining a weight optimization vector corresponding to the initial weight vector, and calculating sensitivities of all connections in the neural network model to be compressed according to the weight optimization vector includes: for each connection in the neural network model to be compressed, calculating a derivative of the overall loss function with respect to the weight optimization vector by using the following formula to approximately characterize the influence of removing the connection on the loss of the neural network model to be compressed: Among them, g j (w; D) represents the derivative value of the overall loss function corresponding to j with respect to the weight optimization vector, and e j represents the indication vector connecting j; according to the calculated derivative value corresponding to each connection, calculating the sensitivity of each connection in the neural network model to be compressed by using the following formula: where s j represents the sensitivity of connection j, and |g j (w; D)| represents the absolute value of the derivative value corresponding to connection j, and N represents the number of connections of the neural network model to be compressed.
5. The method according to any one of claims 1-4, characterized in that, The step of training the sparse neural network model using the target training sample set until the trained target neural network model is obtained includes: Inputting all samples in the target training sample set into the sparse neural network model; Calculating the loss function value according to the prediction result of the sparse neural network model and the labels corresponding to all samples in the target training sample set, and stopping training until the number of iterations exceeds a preset number or the loss function value is less than a preset value, so as to obtain the trained target neural network model.
6. A convolutional neural network compression device Characterized in that The device includes: A sample acquisition module, configured to acquire a target training sample set for a target application scenario; wherein, the target application scenario includes target recognition or target tracking, the target training sample set is determined based on an initial training sample set of a neural network model to be compressed corresponding to the target application scenario, and the target training sample set is pictures; An initialization module, configured to initialize the weights of the neural network model to be compressed by using a variance scaling method to obtain an initial weight vector of the neural network model to be compressed; A sensitivity calculation module, configured to determine a weight optimization vector corresponding to the initial weight vector with a preset sparsity as a constraint condition, and calculate the sensitivity of all connections in the neural network model to be compressed according to the weight optimization vector; wherein, the sensitivity is used to characterize the importance degree of connections of each network layer in the neural network model to be compressed; A pruning module, configured to prune the neural network model to be compressed according to the preset sparsity and the sensitivity to obtain a sparse neural network model; wherein, the weights of the sparse neural network model are determined according to the preset sparsity and the sensitivity; A training module, configured to train the sparse neural network model using the target training sample set until the trained target neural network model is obtained; wherein, the target neural network model is used to process data corresponding to the target application scenario; The pruning module is further configured to: Sort all connections in the neural network model to be compressed in descending order of the sensitivity, and retain the first k connections in the sorting result to obtain a first sparse neural network model; Calculate the optimized weight value corresponding to each of the first k connections by using the following formula: where w i ’ represents the optimized weight value of the i-th connection, and k represents the preset sparsity; Assign the optimized weight value to each of the first k connections to obtain the sparse neural network model.
7. An electronic device Characterized in that It includes a processor and a memory, the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the convolutional neural network compression method according to any one of claims 1 to 6.
8. A computer-readable storage medium, on which a computer program is stored Characterized in that When the computer program is run by a processing device, it executes the steps of the convolutional neural network compression method according to any one of claims 1 to 6.