Convolutional neural network model lightweight method and related device

By combining BN layers in the convolutional neural network model, transforming asymmetric convolutional layers, removing identity map branches, modifying convolution kernels and channels, replacing activation functions, and adding convolutional layers of non-identity map branches, the limitations on the symmetry and activation functions of the convolutional layer in the existing technology are solved, and a more efficient convolutional neural network model is achieved, which improves performance and application scope.

CN120012845APending Publication Date: 2025-05-16GUANGDONG ELECTRIC POWER SCI RES INST ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137283.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing convolutional neural network model structure reparameterization method requires that the convolution layer must be a conventional symmetric convolution, and when there are identity map branches, the activation function cannot be included in other branches, which limits the performance limit and application scope of the convolutional neural network.

Method used

A convolutional neural network model lightweight method is provided, suitable for models composed of one identity map branch and several non-identity map branches. The method includes combining the BN layer into the convolution layer, converting the asymmetric convolution layer into a 3×3 convolution layer, removing identity map branches, modifying the number of convolution kernels and channels, and replacing the ReLU activation layer as the PReLU activation layer, and finally adding the convolution layers of the non-identity map branches to obtain a convolution block with a single branch structure.

Benefits of technology

This method is not only suitable for symmetric convolution, but also for asymmetric convolution, and allows the convolution neural network branches to include activation layers, which improves the network's feature extraction ability and nonlinearity, thereby improving the performance limit during inference and expanding the usage scenarios of structural reparameterization method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012845A_ABST
    Figure CN120012845A_ABST
Patent Text Reader

Abstract

The invention discloses a convolutional neural network model lightweight method and a related device, and the method comprises the steps: merging each BN layer of a non-identical mapping branch into a previous closest convolutional layer, converting a 1 * 1 convolutional layer into a 3 * 3 convolutional layer, converting an asymmetric convolutional layer of the non-identical mapping branch into the 3 * 3 convolutional layer, and then removing the identical mapping branch, thereby obtaining the lightweight convolutional neural network model. Modifying the number of convolution kernels and the number of channels of each non-identical mapping branch, replacing the ReLU activation layer of each non-identical mapping branch with a PReLU activation layer, and finally adding the convolution layers at the corresponding positions of each non-identical mapping branch to obtain a convolution block of a single-branch structure. The technical problems that according to an existing convolutional neural network model structure re-parameterization method, a convolutional layer must be conventional symmetric convolution, when an identical mapping branch exists, other branches cannot contain an activation function, the performance upper limit of a convolutional neural network is limited, and the application range is limited are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network technology, and in particular to a convolutional neural network model lightweight method and related devices. Background Art

[0002] For neural networks, a structure corresponds to a set of parameters. During training, in order to obtain a larger receptive field and stronger feature fusion effect, a complex multi-branch residual structure is often designed. However, this structure has many parameters and a large amount of computation, which is not conducive to reasoning. Therefore, it is necessary to reparameterize it into a single-branch structure that is friendly to reasoning to save time and computing resource costs during reasoning. Structural reparameterization is the equivalent conversion of one network structure into another network structure, which is regarded as an effective method for nearly lossless model lightweighting. The object of structural reparameterization is a multi-branch convolutional block. For each branch, it may be an identity mapping (i.e., no convolutional block), a 1×1 convolutional layer plus a BN layer, or a 3×3 convolutional layer plus a BN layer. Existing methods complete structural reparameterization through the following two steps: (1) merging the batch normalization (BN) layer of each branch into the convolutional layer, and (2) merging the multi-branch structure into a single-branch structure. Specifically, the identity mapping layer and the 1×1 convolutional layer are first converted into 3×3 convolutional layers respectively, and then multiple 3×3 convolutional layers are merged. like Figure 2 As shown in the figure, it is assumed that the convolution block during training has three branches, including a 1×1 convolution layer plus a BN layer, a 3×3 convolution layer plus a BN layer, and an identity mapping layer. During inference, the BN layer is merged into the convolution layer through the structural reparameterization technology, and then multiple convolution layers are merged into a 3×3 convolution layer. However, the existing structural reparameterization method of convolutional neural network models requires that the convolution layer must be a conventional symmetric convolution, and when there is an identity mapping branch, the other branches cannot contain activation functions, which limits the performance upper limit of the convolutional neural network and restricts its application scope. Summary of the invention

[0003] The present invention provides a convolutional neural network model lightweight method and related devices, which are used to solve the technical problems that the existing convolutional neural network model structure re-parameterization method requires that the convolution layer must be a conventional symmetric convolution, and when there is an identity mapping branch, other branches cannot contain activation functions, which limits the performance upper limit of the convolutional neural network and restricts the scope of application.

[0004] In view of this, the present invention provides a convolutional neural network model lightweight method, which is applied to a convolutional neural network model composed of an identity mapping branch and a plurality of non-identity mapping branches, the non-identity mapping branch includes 2 or 3 convolutional layers, each convolutional layer is followed by a BN layer, and there is a ReLU activation layer between every two convolutional layers. The convolutional neural network model lightweight method includes:

[0005] Merge each BN layer of the non-identity mapping branch into the previous closest convolutional layer, and convert the 1×1 convolutional layer into a 3×3 convolutional layer;

[0006] Convert the asymmetric convolutional layer of the non-identity mapping branch to a 3×3 convolutional layer;

[0007] Remove the identity mapping branch;

[0008] Modify the number of convolution kernels and channels of each non-identity mapping branch, and replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0009] The convolutional layers at the corresponding positions of each non-identity mapping branch are added together to obtain a convolutional block with a single-branch structure.

[0010] Optionally, when the non-identity mapping branch in the convolutional neural network model includes two convolutional layers, the number of convolution kernels and the number of channels of each non-identity mapping branch are modified, and the ReLU activation layer of each non-identity mapping branch is replaced with a PReLU activation layer, including:

[0011] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels;

[0012] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ;

[0013] Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0014] Among them, the first convolution layer and the second convolution layer are set successively from the input to the output direction of the convolutional neural network model.

[0015] Optionally, when the non-identity mapping branch in the convolutional neural network model includes three convolutional layers, the number of convolution kernels and the number of channels of each non-identity mapping branch are modified, and the ReLU activation layer of each non-identity mapping branch is replaced with a PReLU activation layer, including:

[0016] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels;

[0017] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ;

[0018] For the third convolution layer of each non-identity mapping branch, c channels are added to each convolution kernel, where the c+kth channel added by the kth convolution kernel is a Dirac matrix, and the remaining channels are all-0 matrices. ;

[0019] Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0020] Among them, the first convolution layer, the second convolution layer and the third convolution layer are set successively from the input to the output direction of the convolutional neural network model.

[0021] Optionally, the asymmetric convolutional layer of the non-identity mapping branch is converted to a 3×3 convolutional layer, including:

[0022] The asymmetric convolutional layer of the non-identity mapping branch is converted into a 3×3 convolutional layer by patching.

[0023] Optionally, the asymmetric convolution of the non-identity mapping branch is converted to a 3×3 convolution by patching, including:

[0024] For the 1×3 convolutional layer of the non-identity mapping branch, a row of all-0 elements is padded in the upper and lower rows. For the 3×1 convolutional layer of the non-identity mapping branch, a column of all-0 elements is padded in the left and right columns.

[0025] The second aspect of the present invention provides a convolutional neural network model lightweight device, which is applied to a convolutional neural network model composed of an identity mapping branch and a plurality of non-identity mapping branches, wherein the non-identity mapping branch includes 2 or 3 convolutional layers, each convolutional layer is followed by a BN layer, and there is a ReLU activation layer between every two convolutional layers, including:

[0026] The merging unit is used to merge each BN layer of the non-identity mapping branch into the previous closest convolutional layer, converting the 1×1 convolutional layer into a 3×3 convolutional layer;

[0027] A conversion unit, used to convert the asymmetric convolutional layer of the non-identity mapping branch into a 3×3 convolutional layer;

[0028] An identity mapping removal unit, used for removing identity mapping branches;

[0029] The identity mapping branch modification unit is used to modify the number of convolution kernels and the number of channels of each non-identity mapping branch, and replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0030] The superposition unit is used to add the convolutional layers at the corresponding positions of each non-identical mapping branch to obtain a convolutional block with a single branch structure.

[0031] Optionally, when the non-identity mapping branch in the convolutional neural network model includes two convolutional layers, the identity mapping branch modification unit is specifically used to:

[0032] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels;

[0033] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ;

[0034] Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0035] Among them, the first convolution layer and the second convolution layer are set successively from the input to the output direction of the convolutional neural network model.

[0036] Optionally, when the non-identity mapping branch in the convolutional neural network model includes three convolutional layers, the identity mapping branch modification unit is specifically used to:

[0037] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels;

[0038] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ;

[0039] For the third convolution layer of each non-identity mapping branch, c channels are added to each convolution kernel, where the c+kth channel added by the kth convolution kernel is a Dirac matrix, and the remaining channels are all-0 matrices. ;

[0040] Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0041] Among them, the first convolution layer, the second convolution layer and the third convolution layer are set successively from the input to the output direction of the convolutional neural network model.

[0042] Optionally, the conversion unit is specifically used for:

[0043] The asymmetric convolution of the non-identity mapping branch is converted into a 3×3 convolution by patching.

[0044] Optionally, the conversion unit is specifically used for:

[0045] For the 1×3 convolutional layer of the non-identity mapping branch, a row of all-0 elements is padded in the upper and lower rows. For the 3×1 convolutional layer of the non-identity mapping branch, a column of all-0 elements is padded in the left and right columns.

[0046] A third aspect of the present invention provides a convolutional neural network model lightweight device, the device comprising a processor and a memory:

[0047] The memory is used to store program code and transmit the program code to the processor;

[0048] The processor is used to execute the convolutional neural network model lightweight method described in any one of the first aspects according to the instructions in the program code.

[0049] A fourth aspect of the present invention provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the convolutional neural network model lightweight method described in any one of the first aspects.

[0050] It can be seen from the above technical solutions that the convolutional neural network model lightweight method provided by the present invention has the following advantages:

[0051] The convolutional neural network model lightweight method provided by the present invention respectively merges each BN layer of the non-identity mapping branch into the previous closest convolution layer, converts the 1×1 convolution layer into a 3×3 convolution layer, converts the asymmetric convolution layer of the non-identity mapping branch into a 3×3 convolution layer, then removes the identity mapping branch, modifies the number of convolution kernels and channels of each non-identity mapping branch, replaces the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer, and finally adds the convolution layers at the corresponding positions of each non-identity mapping branch to obtain a convolution block with a single branch structure. The model lightweight method is not only applicable to symmetric convolution, but also to asymmetric convolution, and allows the convolutional neural network branch to contain an activation layer. Asymmetric convolution can increase the richness of features extracted by the network, and the activation layer can improve the nonlinearity of the network. Both can improve the performance of the network during training, thereby improving the performance upper limit of the network during inference, expanding the use scenarios of the structural reparameterization method, and solving the technical problem that the existing structural reparameterization method of the convolutional neural network model requires that the convolution layer must be a conventional symmetric convolution, and when there is an identity mapping branch, other branches cannot contain activation functions, which limits the performance upper limit of the convolutional neural network and restricts the scope of application. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0053] Figure 1 A schematic diagram of a flow chart of a convolutional neural network model lightweight method provided in the present invention;

[0054] Figure 2 Schematic diagram for existing reparameterization methods;

[0055] Figure 3 A logical schematic diagram of a re-parameterization method for a lightweight convolutional neural network model provided in the present invention;

[0056] Figure 4 A schematic diagram of the structure of a lightweight device for a convolutional neural network model provided in the present invention;

[0057] Figure 5 A schematic diagram of the structure of a lightweight device for a convolutional neural network model provided in the present invention. DETAILED DESCRIPTION

[0058] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] For easier understanding, see Figure 1 and Figure 3 The present invention provides a convolutional neural network model lightweight method, which is applied to a convolutional neural network model composed of an identity mapping branch and a plurality of non-identity mapping branches, wherein the non-identity mapping branch includes 2 or 3 convolutional layers, each convolutional layer is followed by a BN (BatchNormalization, BN) layer, and there is a ReLU activation layer between every two convolutional layers. The convolutional neural network model lightweight method includes:

[0060] Step 101: Merge each BN layer of the non-identity mapping branch into the previous closest convolutional layer, and convert the 1×1 convolutional layer into a 3×3 convolutional layer.

[0061] It should be noted that, in the embodiment of the present invention, first, according to the existing reparameterization method, each BN layer of each non-identity mapping branch is merged into the convolution layer closest to the front of the BN layer, that is, for each BN layer, it is merged into the convolution layer closest to the front of it, and the symmetric 1×1 convolution is converted to a 3×3 convolution. In the present invention, the convolution layer of the non-identity mapping branch is a 1×1 convolution layer, a 1×3 convolution layer, a 3×1 convolution layer, or a 3×3 convolution layer.

[0062] Step 102: Convert the asymmetric convolutional layer of the non-identity mapping branch to a 3×3 convolutional layer.

[0063] It should be noted that the asymmetric convolution layer, i.e., 1×3 convolution layer and 3×1 convolution layer, is converted to a 3×3 convolution layer by patching. Specifically, for the 1×3 convolution layer of the non-identity mapping branch, the convolution kernel is 1 row and 3 columns, and one row of all-0 elements is padded in the row direction, so the convolution kernel is 3 rows and 3 columns. For the 3×1 convolution layer of the non-identity mapping branch, one column of all-0 elements is padded in the column direction, so the convolution kernel is 3 rows and 3 columns.

[0064] Step 103: Remove the identity mapping branch.

[0065] It should be noted that the identity mapping branch is removed.

[0066] Step 104: modify the number of convolution kernels and the number of channels of each non-identity mapping branch, and replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer.

[0067] It should be noted that in the present invention, the identity mapping branch of the convolutional neural network model includes 2 or 3 convolutional layers. When the non-identity mapping branch in the convolutional neural network model includes 2 convolutional layers, assuming that the number of input channels of the non-identity mapping branch is c, the number of convolution kernels and channels of convolutional layers 1 to 2 are both c, and the first convolutional layer and the second convolutional layer are arranged successively from the input to the output direction of the convolutional neural network model.

[0068] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the remaining channels are all-0 matrices. The Dirac matrix has all elements 0 except the middle element which is 1, which can be expressed as:

[0069]

[0070] Where I is a matrix, m and n are the number of rows and columns of the matrix respectively.

[0071] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. .

[0072] At this time, the number of convolution kernels and the number of channels of convolution layers 1 and 2 are 2c and c, 2c respectively.

[0073] Then replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer. The formula for replacing the ReLU activation layer with the PReLU activation layer is:

[0074]

[0075] Among them, y is a nonlinear activation function For the input, q is the negative semi-axis slope. When q=0, PReLU degenerates into ReLU. For the original convolution kernel in the convolution layer, let q=0, and for the convolution kernel added in the convolution layer, let q=1.

[0076] The above adjustments are made to all non-identity mapping branches that have merged BN layers, converted 1×1 convolutions, and asymmetric convolutions to 3×3 convolutions.

[0077] When the non-identical mapping branch in the convolutional neural network model includes 3 convolutional layers, assuming that the number of input channels of the non-identical mapping branch is c, the number of convolution kernels and channels of convolutional layers 1 to 3 are both c, and the first convolutional layer, the second convolutional layer, and the third convolutional layer are set sequentially from the input to the output direction of the convolutional neural network model.

[0078] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the remaining channels are all-0 matrices. The Dirac matrix has all elements 0 except the middle element which is 1, which can be expressed as:

[0079]

[0080] Where I is a matrix, m and n are the number of rows and columns of the matrix respectively.

[0081] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. .

[0082] For the third convolution layer of each non-identity mapping branch, c channels are added to each convolution kernel, where the c+kth channel added by the kth convolution kernel is a Dirac matrix, and the remaining channels are all-0 matrices. .

[0083] At this time, the number of convolution kernels and the number of channels of convolution layers 1, 2, and 3 are 2c and c, 2c and 2c, and c and 2c, respectively.

[0084] Then replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer. The formula for replacing the ReLU activation layer with the PReLU activation layer is:

[0085]

[0086] Among them, y is a nonlinear activation function For the input, q is the negative semi-axis slope. When q=0, PReLU degenerates into ReLU. For the original convolution kernel in the convolution layer, let q=0, and for the convolution kernel added in the convolution layer, let q=1.

[0087] The above adjustments are made to all non-identity mapping branches that have merged BN layers, converted 1×1 convolutions, and asymmetric convolutions to 3×3 convolutions.

[0088] Step 105: Add the convolutional layers at the corresponding positions of each non-identical mapping branch to obtain a convolutional block with a single branch structure.

[0089] It should be noted that, after all non-identity mapping branches have completed the adjustment in step 104, the convolutional layers at the corresponding positions of each non-identity mapping branch are added together to obtain a convolutional block with a single branch structure, such as Figure 3 shown.

[0090] The convolutional neural network model lightweight method provided by the present invention is not only applicable to symmetric convolution, but also to asymmetric convolution, and allows the convolutional neural network branch to include an activation layer. Asymmetric convolution can increase the richness of the features extracted by the network, and the activation layer can improve the nonlinearity of the network. Both can improve the performance of the network during training, thereby improving the performance upper limit of the network during reasoning, expanding the use scenario of the structural reparameterization method, and solving the technical problem that the existing existing convolutional neural network model structural reparameterization method requires that the convolution layer must be a conventional symmetric convolution, and when there is an identity mapping branch, the activation function cannot be included in other branches, which limits the performance upper limit of the convolutional neural network and the scope of application is restricted.

[0091] For easier understanding, see Figure 4 The present invention provides an embodiment of a convolutional neural network model lightweight device, which is applied to a convolutional neural network model composed of an identity mapping branch and a plurality of non-identity mapping branches, wherein the non-identity mapping branch includes 2 or 3 convolutional layers, each convolutional layer is followed by a BN layer, and there is a ReLU activation layer between every two convolutional layers. The device includes:

[0092] The merging unit is used to merge each BN layer of the non-identity mapping branch into the previous closest convolutional layer, converting the 1×1 convolutional layer into a 3×3 convolutional layer;

[0093] A conversion unit for converting the asymmetric convolution of the non-identity mapping branch into a 3×3 convolution;

[0094] An identity mapping removal unit, used for removing identity mapping branches;

[0095] The identity mapping branch modification unit is used to modify the number of convolution kernels and the number of channels of each non-identity mapping branch, and replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0096] The superposition unit is used to add the convolutional layers at the corresponding positions of each non-identical mapping branch to obtain a convolutional block with a single branch structure.

[0097] When the non-identity mapping branch in the convolutional neural network model includes two convolutional layers, the identity mapping branch modification unit is specifically used for:

[0098] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels;

[0099] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ;

[0100] Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0101] Among them, the first convolution layer and the second convolution layer are set successively from the input to the output direction of the convolutional neural network model.

[0102] When the non-identity mapping branch in the convolutional neural network model includes three convolutional layers, the identity mapping branch modification unit is specifically used for:

[0103] For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels;

[0104] For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ;

[0105] For the third convolution layer of each non-identity mapping branch, c channels are added to each convolution kernel, where the c+kth channel added by the kth convolution kernel is a Dirac matrix, and the remaining channels are all-0 matrices. ;

[0106] Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer;

[0107] Among them, the first convolution layer, the second convolution layer and the third convolution layer are set successively from the input to the output direction of the convolutional neural network model.

[0108] The conversion unit is specifically used for:

[0109] The asymmetric convolution of the non-identity mapping branch is converted into a 3×3 convolution by patching.

[0110] The conversion unit is specifically used for:

[0111] For the 1×3 convolutional layer of the non-identity mapping branch, a row of all-0 elements is padded in the upper and lower rows. For the 3×1 convolutional layer of the non-identity mapping branch, a column of all-0 elements is padded in the left and right columns.

[0112] For easier understanding, see Figure 5 The present invention provides an embodiment of a lightweight device for a convolutional neural network model, the device comprising a processor and a memory:

[0113] The memory is used to store program code and transmit the program code to the processor;

[0114] The processor is used to execute any one of the convolutional neural network model lightweight methods provided in the present invention according to the instructions in the program code.

[0115] The present invention also provides an embodiment of a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute any one of the convolutional neural network model lightweight methods provided in the present invention.

[0116] The convolutional neural network model lightweight device, equipment and computer-readable storage medium provided in the present invention are used to execute the convolutional neural network model lightweight method provided in the present invention. The principles and technical effects achieved are the same as those of the convolutional neural network model lightweight method provided in the present invention, and will not be repeated here.

[0117] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A convolutional neural network model lightweight method, characterized in that: Applied to a convolutional neural network model consisting of an identity mapping branch and several non-identity mapping branches. The non-identity mapping branch includes 2 or 3 convolutional layers, each convolutional layer is followed by a BN layer, and there is a ReLU activation layer between every two convolutional layers. The lightweight method of the convolutional neural network model includes: Merge each BN layer of the non-identity mapping branch into the previous closest convolutional layer, and convert the 1×1 convolutional layer into a 3×3 convolutional layer; Convert the asymmetric convolutional layer of the non-identity mapping branch to a 3×3 convolutional layer; Remove the identity mapping branch; Modify the number of convolution kernels and channels of each non-identity mapping branch, and replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer; The convolutional layers at the corresponding positions of each non-identity mapping branch are added together to obtain a convolutional block with a single-branch structure.

2. The convolutional neural network model lightweight method according to claim 1, characterized in that: When the non-identity mapping branch in the convolutional neural network model includes two convolutional layers, the number of convolution kernels and the number of channels of each non-identity mapping branch are modified, and the ReLU activation layer of each non-identity mapping branch is replaced with a PReLU activation layer, including: For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels; For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ; Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer; Among them, the first convolution layer and the second convolution layer are set successively from the input to the output direction of the convolutional neural network model.

3. The convolutional neural network model lightweight method according to claim 1, characterized in that: When the non-identity mapping branch in the convolutional neural network model includes three convolutional layers, the number of convolution kernels and the number of channels of each non-identity mapping branch are modified, and the ReLU activation layer of each non-identity mapping branch is replaced with a PReLU activation layer, including: For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels; For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ; For the third convolution layer of each non-identity mapping branch, c channels are added to each convolution kernel, where the c+kth channel added by the kth convolution kernel is a Dirac matrix, and the remaining channels are all-0 matrices. ; Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer; Among them, the first convolution layer, the second convolution layer and the third convolution layer are set successively from the input to the output direction of the convolutional neural network model.

4. The convolutional neural network model lightweight method according to claim 1, characterized in that: Convert the asymmetric convolutional layer of the non-identity mapping branch to a 3×3 convolutional layer, including: The asymmetric convolutional layer of the non-identity mapping branch is converted into a 3×3 convolutional layer by patching.

5. The convolutional neural network model lightweight method according to claim 4, characterized in that: The asymmetric convolutional layer of the non-identity mapping branch is converted to a 3×3 convolutional layer by patching, including: For the 1×3 convolutional layer of the non-identity mapping branch, a row of all-0 elements is padded in the upper and lower rows. For the 3×1 convolutional layer of the non-identity mapping branch, a column of all-0 elements is padded in the left and right columns.

6. A convolutional neural network model lightweight device, characterized in that: Applied to a convolutional neural network model consisting of an identity mapping branch and several non-identity mapping branches. The non-identity mapping branch includes 2 or 3 convolutional layers, each followed by a BN layer, and a ReLU activation layer between every two convolutional layers, including: The merging unit is used to merge each BN layer of the non-identity mapping branch into the previous closest convolutional layer, converting the 1×1 convolutional layer into a 3×3 convolutional layer; A conversion unit, used to convert the asymmetric convolutional layer of the non-identity mapping branch into a 3×3 convolutional layer; An identity mapping removal unit, used for removing identity mapping branches; The identity mapping branch modification unit is used to modify the number of convolution kernels and the number of channels of each non-identity mapping branch, and replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer; The superposition unit is used to add the convolutional layers at the corresponding positions of each non-identical mapping branch to obtain a convolutional block with a single branch structure.

7. The convolutional neural network model lightweight device according to claim 6, characterized in that: When the non-identity mapping branch in the convolutional neural network model includes two convolutional layers, the identity mapping branch modification unit is specifically used for: For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels; For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ; Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer; Among them, the first convolution layer and the second convolution layer are set successively from the input to the output direction of the convolutional neural network model.

8. The convolutional neural network model lightweight device according to claim 6, characterized in that: When the non-identity mapping branch in the convolutional neural network model includes three convolutional layers, the identity mapping branch modification unit is specifically used for: For the first convolution layer of each non-identity mapping branch, add convolution kernels of the number of branch input channels, where the i-th channel of the i-th convolution kernel added is the Dirac matrix, and the other channels are all-0 matrices. , c is the number of branch input channels; For the second convolution layer of each non-identity mapping branch, first add c all-zero value channels to each convolution kernel, then add c convolution kernels with 2c channels. The c+jth channel of the jth convolution kernel added is the Dirac matrix, and the remaining channels are all-zero matrices. ; For the third convolution layer of each non-identity mapping branch, c channels are added to each convolution kernel, where the c+kth channel added by the kth convolution kernel is a Dirac matrix, and the remaining channels are all-0 matrices. ; Replace the ReLU activation layer of each non-identity mapping branch with a PReLU activation layer; Among them, the first convolution layer, the second convolution layer and the third convolution layer are set successively from the input to the output direction of the convolutional neural network model.

9. A convolutional neural network model lightweight device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the convolutional neural network model lightweight method described in any one of claims 1-5 according to the instructions in the program code.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program code, and the program code is used to execute the convolutional neural network model lightweight method described in any one of claims 1-5.