An Image Classification Method Based on the Fusion of Redundancy and Diversity Features
By introducing an extended fusion convolution filter into the convolutional neural network model, redundancy and diversity characteristics are obtained and fusion is solved, the problem of the model lacking diversity information in object detection is achieved, the effect of reducing the amount of parameters and calculations is improved, and the classification effect is improved.
Patent Information
- Application Number
- CN202211676949.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-12-26
AI Technical Summary
The existing convolutional neural network models lack the ability to capture diversity information during object detection, resulting in multiple overlapping partial object proposals in the proposal pool rather than a single complete object proposal, while the redundant features are too similar and lack diversity.
A lightweight convolutional network model is proposed by replacing part of the filter in the convolutional neural network model with an extended fusion convolutional filter. This method acquires redundant and diverse features through feature extraction module and feature expansion module, and fuses through channel attention module and grouping convolution to reduce the amount of model parameters and calculations.
While maintaining a certain amount of redundant features, the number of model parameters and calculations is effectively reduced, the accuracy and feature diversity of the model are improved, and the classification effect is improved.
Smart Images

Figure CN115937600B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image classification, and particularly relates to an image classification method based on the fusion of redundant and diverse features. Background Art
[0002] In recent years, various methods for reducing model complexity and network parameters, such as model quantization and network pruning, have been successively proposed and achieved good results in their respective different fields. One of the main research directions is to optimize the CNN network structure with large redundancy by exploring the dependence relationships across spatial and channel dimensions. On the one hand, redundant convolutional neural networks lack the ability to capture diverse information. When performing object detection, they often cannot accurately locate the target area and aggregate useful information from the foreground, resulting in multiple overlapping partial object proposals rather than a single complete object proposal in the proposal pool.
[0003] On the other hand, it is necessary to have a certain amount of redundancy in the feature map. Rich or even redundant information can often ensure a comprehensive understanding of the input data. The GhostNet network introduces a new Ghost module. While maintaining a certain number of inherent features, depthwise convolutions (DWC) with fewer parameters are used to process these inherent features and generate redundant features. However, this approach generates overly similar features and lacks diverse information. Therefore, this paper proposes an image classification method based on the fusion of redundant and diverse features. Summary of the Invention
[0004] The purpose of the present invention is to address the above problems and propose an image classification method based on the fusion of redundant and diverse features. This method can effectively reduce model redundancy, significantly reduce the number of parameters and computational complexity, and obtain diverse features while ensuring a certain amount of redundant features to guarantee model accuracy.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] An image classification method based on the fusion of redundant and diverse features proposed by the present invention includes the following steps:
[0007] S1. Replace some filters in the convolutional neural network model with extended fusion convolutional filters to form a lightweight convolutional network model. The extended fusion convolutional filters are set as follows:
[0008] S11. According to the first proportional parameter α, split the input channel number C of the replaced filter into a main part and an extended part. The main part contains αC channels, and the extended part contains (1 - α)C channels;
[0009] S12. Obtain redundant features and diverse features by using a feature extraction module and a feature expansion module. The feature extraction module includes a parallel basic branch and a diversity branch. The basic branch is a standard convolutional layer, and the diversity branch is a grouped convolutional layer. The feature expansion module is a point convolutional layer. The acquisition of redundant features and diverse features is as follows:
[0010] Use the basic branch to extract features from the main part to form redundant features;
[0011] Use the feature expansion module to extract features from the expansion part and increase the number of feature channels according to the hyperparameter γ to obtain expansion features;
[0012] Concatenate the main part and the expansion features to form a first feature, and use the diversity branch to extract features from the first feature to form diverse features;
[0013] S13. Adjust the output channels of the redundant features and the diverse features according to the second proportional parameter β, that is, the output channels of the redundant features are βC′, and the output channels of the diverse features are (1 - β)C′. Concatenate the redundant features and the diverse features along the channel dimension to form a second feature, and the output channels of the second feature are the output channels C′ of the replaced filter;
[0014] S14. Input the second feature into the channel attention module to obtain a third feature, and the output channels of the third feature are C′;
[0015] S15. Match the dimensions of the expansion features with the third feature through grouped convolution, and fuse them with the third feature in an element-wise addition manner to obtain an output feature map. The number of groups of the grouped convolution is the output channels βC′ of the redundant features, and the convolution kernel size is 1*1;
[0016] S2. Use the dataset to train and validate the lightweight convolutional network model to obtain a trained lightweight convolutional network model;
[0017] S3. Input the image to be processed into the trained lightweight convolutional network model to obtain a classification result.
[0018] Preferably, the convolutional neural network model is a ResNet50 model.
[0019] Preferably, the Resnet50 network includes a convolutional layer, a pooling layer, a first residual block, a second residual block, a third residual block, a fourth residual block, and a fully connected layer connected in sequence. The first residual block includes three concatenated bottleneck layers, the second residual block includes four concatenated bottleneck layers, the third residual block includes six concatenated bottleneck layers, and the fourth residual block includes three concatenated bottleneck layers. Each bottleneck layer is composed of a filter with a kernel size of 1*1, a filter with a kernel size of 3*3, a filter with a kernel size of 1*1, a BN layer, and a RELU activation function connected in series, and the filter with a kernel size of 3*3 is replaced by an extended fusion convolutional filter.
[0020] Preferably, the convolutional kernel size of the convolutional layer is 7*7.
[0021] Preferably, the diversity branch is a progressive grouped group convolutional layer, that is, as the network depth increases, the number of groups gradually increases.
[0022] Preferably, the number of groups of the diversity branch is the greatest common divisor of the number of channels of the main body part and the number of channels of the extended features.
[0023] Preferably, the dataset is the ImageNet dataset.
[0024] Preferably, the training process of the lightweight convolutional network model is as follows:
[0025] Input images of size 224*224 into the lightweight convolutional network model with a batch size of 256 and train for 100 epochs.
[0026] Preferably, the channel attention module is the SE attention module.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0028] This method uses an extended fusion convolutional filter to replace some filters in the traditional convolutional neural network model to form a lightweight convolutional network model. Different from the traditional convolutional neural network model, this lightweight convolutional network model can not only obtain inherent features with a certain degree of redundancy, but also generate diversity features containing necessary detailed information, improving the accuracy and feature diversity of the model while effectively reducing the number of model parameters and floating-point operations, thereby improving the classification effect; in addition, the extended fusion convolutional filter can also be used as a plug-and-play component to upgrade the existing convolutional neural network. Global unified setting of hyperparameters can facilitate the rapid construction of the model, and by reasonably configuring the hyperparameters of each layer of the model, the model has better performance. Description of the Drawings
[0029] Figure 1 It is a flowchart of the image classification method based on the fusion of redundant and diversity features of the present invention;
[0030] Figure 2 Schematic diagram of the extended fusion convolution filter structure of the present invention;
[0031] Figure 3 Visual feature maps of the basic branch and the diversity branch in the feature extraction module of the present invention. Detailed implementation manners
[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0033] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the description of this application herein are only for the purpose of describing specific embodiments, and are not intended to limit this application.
[0034] In order to capture diverse features and improve the model accuracy while effectively reducing the number of model parameters, the present invention proposes an image classification method based on the fusion of redundant and diverse features. Specifically, the input features are split into a main part and an extended part according to a certain ratio. The main part extracts redundant features and diverse features containing certain redundant information through a feature extraction module with a parallel dual-branch structure, and a channel attention module is used to increase the weights of the feature channels containing necessary information in the redundant and diverse features. The extended part strengthens the extraction of diverse information through a feature extension module, i.e., point convolution, and supplements the finally output features. At the same time, the extended fusion convolution filter proposed by the present invention can effectively reduce the number of model parameters and the amount of computation, and improve the performance.
[0035] As Figures 1-3 shown, an image classification method based on the fusion of redundant and diverse features includes the following steps:
[0036] S1. Replace some filters in the convolutional neural network model with extended fusion convolution filters to form a lightweight convolutional network model. The extended fusion convolution filter is set as follows:
[0037] S11. Split the input channel number C of the replaced filter into a main part and an extended part according to the first proportional parameter α. The main part contains αC channels, and the extended part contains (1 - α)C channels;
[0038] S12. Obtain redundant features and diverse features using the feature extraction module and the feature expansion module. The feature extraction module includes a parallel basic branch and a diverse branch. The basic branch is a standard convolutional layer, and the diverse branch is a grouped convolutional layer. The feature expansion module is a point convolutional layer. The acquisition of redundant features and diverse features is as follows:
[0039] Use the basic branch to extract features from the main part to form redundant features;
[0040] Use the feature expansion module to extract features from the expansion part and increase the number of feature channels according to the hyperparameter γ to obtain expansion features;
[0041] Concatenate the main part and the expansion features to form the first feature, and use the diverse branch to extract features from the first feature to form diverse features;
[0042] S13. Adjust the output channels of the redundant features and the diverse features according to the second proportional parameter β, that is, the output channels of the redundant features are βC′, and the output channels of the diverse features are (1 - β)C′. Concatenate the redundant features and the diverse features along the channel dimension to form the second feature, and the output channels of the second feature are the output channels C′ of the replaced filter;
[0043] S14. Input the second feature into the channel attention module to obtain the third feature, and the output channels of the third feature are C′;
[0044] S15. Match the dimensions of the expansion features and the third feature through grouped convolution, and fuse them with the third feature in an element-wise addition manner to obtain the output feature map. The number of groups of the grouped convolution is the output channels βC′ of the redundant features, and the convolution kernel size is 1*1.
[0045] In one embodiment, the convolutional neural network model is a ResNet50 model.
[0046] In one embodiment, the Resnet50 network includes a convolutional layer, a pooling layer, a first residual block, a second residual block, a third residual block, a fourth residual block, and a fully connected layer connected in sequence. The first residual block includes three cascaded bottleneck layers, the second residual block includes four cascaded bottleneck layers, the third residual block includes six cascaded bottleneck layers, and the fourth residual block includes three cascaded bottleneck layers. Each bottleneck layer is composed of a filter with a convolution kernel size of 1*1, a filter with a convolution kernel size of 3*3, a filter with a convolution kernel size of 1*1, a BN layer, and a RELU activation function connected in series, and the filter with a convolution kernel size of 3*3 is replaced by an extended fusion convolution filter.
[0047] In one embodiment, the convolution kernel size of the convolutional layer is 7*7.
[0048] In one embodiment, the diversity branch is a grouped convolutional layer with progressive grouping, that is, as the number of network layers deepens, the number of groups gradually increases.
[0049] In one embodiment, the number of groups of the diversity branch is the greatest common divisor of the number of channels in the main part and the number of channels in the extended features. And as the number of network layers deepens and the number of feature channels increases, the number of groups of the diversity branch is the greatest common divisor of αC and γ(1 - β)C'.
[0050] In one embodiment, the channel attention module is the SE attention module. Or other channel attention modules in the prior art can also be used.
[0051] The ResNet50 model is used as the backbone network for image segmentation. For the ImageNet dataset, the input of the ResNet50 model is a natural image (Input image) with a size of 224×224 and 3 channels. The first layer of the ResNet50 model is a 7*7 convolutional layer, followed by a pooling layer, and then four residual blocks are connected in series. Each residual block contains 3, 4, 6, and 3 bottleneck layers respectively. Each bottleneck layer is composed of a group of filters with a convolutional kernel size of 1*1, a filter with a convolutional kernel size of 3*3, a filter with a convolutional kernel size of 1*1, a BN layer, and a RELU activation function connected in series. Finally, a fully connected layer is used to output the classification result. For the four residual blocks in the backbone network, keep the filters with a convolutional kernel size of 1*1 in all bottleneck layers unchanged, and only replace the filters with a convolutional kernel size of 3*3 with extended fusion convolutional filters. Set the default values of the hyperparameters α = 0.5, β = 0.5, and γ = 2. As a plug-and-play component, globally setting the hyperparameters uniformly can facilitate the construction of the model, such as reasonably configuring the hyperparameters of each layer of the model, and can also make the model have better performance. It should be noted that the convolutional neural network model can also be replaced with other models in the prior art, such as VGG-16, ResNet-56, MobileNet, ShuffleNet, etc.
[0052] As Figure 2 shown, when replacing the filter with a convolutional kernel size of 3*3 with an extended fusion convolutional filter, the input feature map X (Input feature maps X) of the extended fusion convolutional filter is the input feature map of the filter with a convolutional kernel size of 3*3, with a size of C×W×H, where W is the width of the input feature map and H is the height of the input feature map. According to the first ratio parameter (Ratio of splitting) α, split the input channel number C of the replaced filter into the main part X M (Mainpart X M ) and the extended part X E(Expansion part X E ), with dimensions corresponding to αC×W×H and (1-α)C×W×H in sequence. The redundant features and diverse features are obtained by using the Feature Extraction Block and the Feature Expansion Block. The base branch of the Feature Extraction Block is a standard convolutional layer, the diversity branch of the Feature Extraction Block is a grouped convolutional layer, and the Feature Expansion Block is a point convolutional layer. The base branch is used to extract features from the main part to form redundant features; the Feature Expansion Block is used to extract features from the expansion part and increase the number of feature channels according to the hyperparameter (Times of expansion) γ to obtain expansion features; then the main part and the expansion features are concatenated (Concat operation) to form the first feature, and the diversity branch is used to extract features from the first feature to form diverse features. Figure 3 are the visualization feature maps of the base branch and the diversity branch. According to the second ratio parameter (Ratio of output) β, the output channels of the redundant features and the diverse features are adjusted, that is, the output channels of the redundant features are βC′, and the output channels of the diverse features are (1-β)C′, and the redundant features and the diverse features are concatenated (Concat operation) along the channel dimension to form the second feature, and the output channels of the second feature are the output channels C′ of the replaced filter. The second feature is input into the Channel attention module to obtain the third feature. The expansion features are dimensionally matched with the third feature through grouped convolution, changing the number of channels to C′, and the third feature and the dimensionally matched expansion features are fused in an element-wise addition manner (Add operation) to obtain the output feature map Y (Output feature maps Y), with dimensions of C′×W′×H′, where W′ is the width of the output feature map and H′ is the height of the output feature map.
[0053] Among them, in standard convolution, although the receptive field of each filter is limited spatially, it can capture the channel information of the complete input features. In group convolution (GWC), in order to reduce the redundant connections between channels, a filter can only be convolved with a subset of the input features, which can reduce the number of parameters. Pointwise convolution (PWC) can integrate information across channels and can complement group convolution (GWC). The feature expansion module uses pointwise convolution to expand the number of feature channels, compensating for the shortage of the number of channels in the feature subset that each filter in the group convolution can extract due to the increase in the number of groups in the diversity branch, and enhancing the extraction of diverse features. By cleverly combining the above methods, the redundant connections between channels are reasonably reduced to achieve the effect of reducing the computational cost of the model.
[0054] S2. Use the dataset to train and validate the lightweight convolutional network model to obtain a trained lightweight convolutional network model.
[0055] In one embodiment, the dataset is the ImageNet dataset.
[0056] In one embodiment, the training process of the lightweight convolutional network model is as follows:
[0057] Input images of size 224*224 into the lightweight convolutional network model for training for 100 epochs with a batch size of 256.
[0058] Training and validating the model using the dataset is a well-known technique to those skilled in the art and will not be elaborated here. Among them, 1 epoch means training once with all the samples in the dataset, and batch size is the batch size, indicating the number of samples selected for a single training.
[0059] S3. Input the image to be processed into the trained lightweight convolutional network model to obtain the classification result.
[0060] This method uses an extended fusion convolutional filter to replace some of the filters in the traditional convolutional neural network model to form a lightweight convolutional network model. Different from the traditional convolutional neural network model, this lightweight convolutional network model can not only obtain the inherent features with a certain degree of redundancy, but also generate diverse features containing necessary detailed information, improving the accuracy and feature diversity of the model while effectively reducing the number of model parameters and floating-point operations, thus enhancing the classification effect; in addition, the extended fusion convolutional filter can also be used as a plug-and-play component to upgrade the existing convolutional neural network. Global unified setting of hyperparameters can facilitate the rapid construction of the model, and by reasonably configuring the hyperparameters of each layer of the model, the model has better performance.
[0061] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0062] The above-described embodiments only express the embodiments of the present application that are relatively specific and detailed in description, but should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An image classification method based on the fusion of redundancy and diversity features, characterized in that: The image classification method based on the fusion of redundancy and diversity features includes the following steps: S1. Replace some filters in the convolutional neural network model with extended fusion convolutional filters to form a lightweight convolutional network model. The extended fusion convolutional filters are set as follows: S11. Split the input channel number C of the replaced filter into a main part and an extended part according to the first proportional parameter α. The main part contains αC channels, and the extended part contains (1 - α)C channels; S12. Use a feature extraction module and a feature extension module to obtain redundant features and diversity features. The feature extraction module includes a parallel basic branch and a diversity branch. The basic branch is a standard convolutional layer, and the diversity branch is a grouped convolutional layer. The feature extension module is a point convolutional layer. The acquisition of the redundant features and diversity features is as follows: Use the basic branch to extract features from the main part to form redundant features; Use the feature extension module to extract features from the extended part and increase the number of feature channels according to the hyperparameter γ to obtain extended features; Concatenate the main part and the extended features to form a first feature, and use the diversity branch to extract features from the first feature to form diversity features; S13. Adjust the output channel numbers of the redundant features and the diversity features according to the second proportional parameter β, that is, the output channel number of the redundant features is βC′, and the output channel number of the diversity features is (1 - β)C′. Concatenate the redundant features and the diversity features in the channel dimension to form a second feature. The output channel number of the second feature is the output channel number C′ of the replaced filter; S14. Input the second feature into a channel attention module to obtain a third feature. The output channel number of the third feature is C′; S15. Match the dimensions of the extended features through grouped convolution with the third feature, and fuse them with the third feature in an element-wise addition manner to obtain an output feature map. The number of groups of the grouped convolution is the output channel number βC′ of the redundant features, and the convolution kernel size is 1*1; S2. Use the data set to train and validate the lightweight convolutional network model to obtain a trained lightweight convolutional network model; S3. Input the image to be processed into the trained lightweight convolutional network model to obtain a classification result.
2. The image classification method based on the fusion of redundancy and diversity features according to claim 1, characterized in that: The convolutional neural network model is a ResNet50 model.
3. The image classification method based on the fusion of redundancy and diversity features according to claim 2, characterized in that: The Resnet50 network includes a convolutional layer, a pooling layer, a first residual block, a second residual block, a third residual block, a fourth residual block, and a fully connected layer connected in sequence. The first residual block includes three serially connected bottleneck layers. The second residual block includes four serially connected bottleneck layers. The third residual block includes six serially connected bottleneck layers. The fourth residual block includes three serially connected bottleneck layers. Each bottleneck layer is composed of a filter with a convolution kernel size of 1*1, a filter with a convolution kernel size of 3*3, a filter with a convolution kernel size of 1*1, a BN layer, and a RELU activation function connected in series, and the filter with a convolution kernel size of 3*3 is replaced with an extended fusion convolutional filter.
4. The image classification method based on the fusion of redundancy and diversity features according to claim 3, characterized in that: The convolution kernel size of the convolutional layer is 7*7.
5. The image classification method based on the fusion of redundancy and diversity features according to claim 1, characterized in that: The diversity branch is a grouped convolutional layer with progressive grouping, that is, as the network depth increases, the number of groups gradually increases.
6. The image classification method based on the fusion of redundancy and diversity features according to claim 5, characterized in that: The number of groups in the diversity branch is the greatest common divisor of the number of channels in the main part and the number of channels in the extended features.
7. The image classification method based on the fusion of redundancy and diversity features according to claim 1, characterized in that: The dataset is the ImageNet dataset.
8. The image classification method based on the fusion of redundancy and diversity features according to claim 1, characterized in that: The training process of the lightweight convolutional network model is as follows: Images of size 224*224 are input into the lightweight convolutional network model for training with a batch size of 256 for 100 epochs.
9. The image classification method based on the fusion of redundancy and diversity features according to claim 1, characterized in that: The channel attention module is the SE attention module.