A Dual-Path Architecture Model Compression Method for Image Recognition Datasets
By designing a dual-path architecture and feature separation module in a deep convolutional neural network, the efficient compression of the model is achieved, the problem of difficult deployment of deep convolutional neural networks on mobile devices is solved, the computing speed is improved and the computing power requirements is reduced, and excellent performance is achieved in image recognition tasks.
Patent Information
- Application Number
- CN202111402292.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The existing deep convolutional neural networks are difficult to deploy effectively on mobile devices, mainly due to the large amount of calculation and parameters of the model, which leads to slow real-time computing speed and high requirements for chip computing power.
A dual-path architecture model compression method is designed. By constructing a dual-path structure for each convolution stage, the deep path is stacked by several convolution blocks, the shallow path contains only one layer of dimensional mapping, and the feature channel is separated through the channel attention mechanism in the feature separation module, and the channel proportion is adjusted to achieve model compression.
The image recognition data set is used to significantly compress the model calculation amount and parameter amount, which improves real-time computing speed, reduces the requirements for chip computing power, and achieves excellent test results on the ImageNet data set.
Smart Images

Figure CN114154409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, deep learning, and network compression technology, and specifically to a method for compressing a dual-path architecture model for an image recognition dataset. Background Art
[0002] Image recognition technology refers to the technology of recognizing targets in images. In the field of autonomous driving, how to accurately and efficiently recognize targets in road scenes while ensuring the speed of real-time computing and the low requirement for chip computing power is a major challenge in the fields of machine learning and artificial intelligence.
[0003] The exploration process of image recognition has gone through the following stages: early through template matching; using pattern recognition to complete recognition and evaluation; using deep neural networks for feature extraction; directly training an end-to-end network for recognition.
[0004] Pattern Recognition in image recognition is a process of automatically completing recognition and evaluation of shapes, patterns, curves, numbers, character formats, and graphics from a large amount of information and data, based on expert experience and existing knowledge, using computer and mathematical reasoning methods. Pattern recognition includes two stages, namely feature extraction and classification. The former is to select features for samples, and the latter is to classify and recognize an unknown sample set.
[0005] With the progress of machine learning algorithms and computer hardware, it has become possible to build large-scale deep convolutional neural networks (CNNs) containing multiple layers of convolutions. On a dataset of a large number of pictures, CNNs even outperform traditional image feature extraction methods in terms of model performance in image recognition. However, deep convolutional neural networks that can be effectively trained with the support of high-performance GPU cards are difficult to be effectively deployed on mobile devices represented by mobile phones. Therefore, model compression and lightweight model design have been the research hotspots of convolutional neural networks in recent years.
[0006] The channel attention mechanism models the relationship between channels from the input feature map through two fully connected layers to obtain a more effective feature representation.
[0007] The channel attention mechanism obtains a set of weights in the channel dimension by global average pooling of the input feature map, and captures the relationship between channels through these two fully connected layers. There is a dimensionality reduction between the two fully connected layers to reduce the number of parameters required for this module. Finally, the obtained weights are mapped to the interval from 0 to 1 through the Sigmoid activation, and then multiplied by the original feature map to achieve the activation and suppression of channels. Summary of the Invention
[0008] The object of the present invention is to establish a dual-path model taking the ResNet50 and MobileNetV1 architectures as examples, and then achieve efficient and accurate channel feature extraction through a feature separation module, and provide a method for compressing a dual-path architecture model for an image recognition dataset, which method achieves good recognition effects on the image recognition dataset.
[0009] In order to achieve the above object, a method for compressing a dual-path architecture model for an image recognition dataset is designed, which is characterized in that the method comprises the following steps:
[0010] Step 1: Construct the convolutional blocks in each stage of a certain convolutional neural network into a dual-path structure. The convolutional neural network includes a deep path and a shallow path. The deep path is stacked by a plurality of convolutional blocks, and the shallow path only contains one layer of dimension mapping;
[0011] Step 2: Pass the input feature map through a feature separation module;
[0012] Step 3: Map the feature map with positive weights to a specified dimension and then input it into the deep path, and input the channels with negative weights into the shallow path;
[0013] Step 4: Concatenate the feature maps that have undergone convolutional operations in the two paths back to the original dimension for output;
[0014] Step 5: Adjust the channel ratio between the two paths in each convolutional stage to achieve different degrees of compression of the model and then perform training.
[0015] The present invention also has the following preferred technical solutions:
[0016] 1. The dual-path architecture in Step 1 means that each convolutional stage includes two paths, where the deep path is stacked by a plurality of convolutional blocks, and the shallow path only contains one layer of dimension mapping.
[0017] 2. The feature separation operation in Step 2 means that: first, a set of channel weights is learned through a channel attention mechanism and mapped to the interval from -1 to 1, then ReLU activation is respectively performed and ReLU activation is performed after taking the inverse, and the two sets of weights obtained are multiplied by the original feature map to separate the corresponding feature channels. The specific method is as follows:
[0018] Step a1: Use global average pooling to compress the size of the feature channels from H×W×C to a representation of 1×1×C:
[0019]
[0020] Step a2: Model the relationship between channels for the input feature map through two fully connected layers, and use the Tanh activation function to map the output to the interval from -1 to 1:
[0021] ω = σ(W2ReLU(W1g(x)))
[0022] Step a3: Execute the ReLU activation function on ω to retain the positive weights, and take the negative and do the ReLU activation to get the negative weights:
[0023] ω1 = ReLU(ω)
[0024] ω2 = ReLU(-ω)
[0025] Step a4: By multiplying ω1 and ω2 with the original feature map, two feature maps can be obtained, corresponding to the channels with positive weights and negative weights respectively.
[0026] Compared with the prior art, the advantages of the present invention are as follows: The present invention constructs a dual-path architecture taking the ResNet50 and MobileNetV1 architectures as examples. After model compression, this architecture still has good recognition effects on image recognition data sets, and different degrees of model compression can be achieved only by adjusting the channel hyperparameters. Specifically, it is reflected as follows:
[0027] 1. When the dual-path structure constructed by the present invention compresses the model computational amount to one-half and one-quarter, compared with other pruning methods and lightweight models, the current best test results are obtained on the ImageNet data set.
[0028] 2. When the number of parameters is compressed from 25.6M to 13.2M parameters, and the computational amount is compressed from 4.1G to 2.26G FLOPs, a top-1 accuracy of 76.6% is achieved on the ImageNet data set, slightly higher than 76.4% of ResNet-50.
[0029] 3. When the model number of parameters and computational amount are further compressed to 6.8M parameters and 1.12G FLOPs, a top-1 accuracy of 74.5% is achieved on ImageNet, and this result is better than the current best pruning methods and lightweight design models.
[0030] 4. On the lightweight model MobileNetV1, more competitive results can still be obtained by constructing a dual-path design compared with pruning methods.
[0031] 5. Improve the real-time computing speed and reduce the requirement for chip computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flow chart of the present invention;
[0033] Figure 2 Schematic diagram of the feature separation module;
[0034] Figure 3 Schematic diagram of the dual-path architecture. Detailed implementation manners
[0035] The present invention will be further described below with reference to the accompanying drawings. The structure and principle of the present invention are very clear to those skilled in the art. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0036] The present invention proposes a method for compressing a dual-path architecture model for an image recognition dataset. The method takes a deep convolutional neural network model as the target model, realizes the compression of the model by constructing a dual-path architecture, and ensures the effective feature extraction of channels through a feature separation module, obtaining excellent image recognition test results.
[0037] The present invention realizes channel separation by deploying a feature separation module. However, the two fully connected layers included in the feature separation module have a large number of parameters. Therefore, the present invention deploys a feature separation module at each stage of the deep convolutional neural network, and only four such modules need to be deployed.
[0038] The present invention includes the following steps:
[0039] 1. Construct a dual-path architecture:
[0040] Generally, the number of parameters and the amount of computation of a convolutional neural network are large, and the compression of the model is realized by constructing a dual-path structure.
[0041] 2. Deploy a feature separation module in each dual-path architecture to realize feature separation.
[0042] This module includes two fully connected layers, and maps the learned weights to the interval from -1 to 1 through the Tanh activation function, and then realizes channel separation through ReLU activation and inverse ReLU activation.
[0043] 3. Input the separated feature channels into two paths respectively.
[0044] The channels with positive weights are input into the deep path, and the channels with low weights are input into the shallow path, reducing the computational amount of the model.
[0045] 4. Concatenate and output the two parts of feature channels.
[0046] Concatenating the two feature channels back to the original dimension can avoid feature loss.
[0047] 5. Train the deep convolutional neural network.
[0048] 6. Perform testing after training is completed.
[0049] After the model training is completed, fix the model parameters and perform verification on the test set to obtain the test results.
[0050] The following are specific embodiments for training an image recognition model. In this embodiment, the image dataset used for training is ImageNet.
[0051] 1. Train a dual-path model based on ResNet50 and MobileNetV1 using cross-entropy loss on the ImageNet dataset containing 1000 classes of objects.
[0052] 2. After training is completed, perform performance testing on the test set of ImageNet. The experimental results are shown in Tables 1 and 2 below.
[0053]
[0054] Table 1 Performance comparison of the dual-path model based on ResNet50
[0055]
[0056] Table 2 Performance comparison of the dual-path model based on MobileNetV1.
Claims
1. A method for compressing a dual-path architecture model for an image recognition dataset, characterized in that, The method comprises the following steps: Step 1: Construct the convolutional blocks in each stage of a certain convolutional neural network into a dual-path structure. The convolutional neural network includes a deep path and a shallow path. The deep path is stacked by several convolutional blocks, and the shallow path only contains one layer of dimensional mapping; Step 2: Pass the input feature map through the feature separation module; Step 3: Map the feature map with positive weights to the specified dimension and input it into the deep path, and input the channels with negative weights into the shallow path; Step 4: Concatenate the feature maps that have undergone convolutional operations in the two paths back to the original dimension for output; Step 5: Adjust the channel ratio between the two paths in each convolutional stage to achieve different degrees of compression of the model and then perform training; The feature separation operation in Step 2 refers to: first, learn a set of channel weights through the channel attention mechanism and map them to the interval from -1 to 1, then perform ReLU activation and negation followed by ReLU activation respectively, and multiply the two sets of weights obtained with the original feature map to separate the corresponding feature channels. The specific method is as follows: Step a1: Use global average pooling to compress the size of the feature channels from H×W×C to a representation of 1×1×C: Step a2: Pass the input feature map through two fully connected layers to model the relationship between channels, and use the Tanh activation function to map the output to the interval from -1 to 1: ω = σ(W2ReLU(W1g(x))) Step a3: Perform the ReLU activation function on ω to retain the positive weights, and perform ReLU activation after negation to obtain negative weights: ω1 = ReLU(ω) ω2 = ReLU(-ω) Step a4: By multiplying ω1 and ω2 with the original feature map, two feature maps can be obtained, corresponding to the channels with positive weights and negative weights respectively.