Fine Classification Method for Complex Objects Based on Dual-Channel Attention Mechanism

By adopting a dual-channel attention-intensive connection network in complex scenarios, combining spatial and channel attention mechanisms, and dynamic weighting processing features, the problem of difficulty in fully learning scene features and ignoring small-scale information in the existing technology is solved, and a high-precision complex scene classification is achieved.

CN116452874BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202310420559.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2025-07-01
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

The prior art is difficult to fully learn scene features in complex scenarios, and it is easy to ignore image small-scale information, resulting in a decrease in classification accuracy.

Method used

A dual-channel attention dense connection network is adopted to extract features through a single convolutional layer and a maximum pooling layer, and a transition layer with halved channels is added between dense blocks. Combining the spatial attention mechanism and channel attention mechanism, dynamic weighting is used to enhance the feature extraction of high-resolution remote sensing images.

Benefits of technology

It effectively solves the problem that the scene characteristics cannot be fully learned in complex scenarios, improves the model's feature expression ability, and completely extracts image information, improving classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452874B_ABST
    Figure CN116452874B_ABST
Patent Text Reader

Abstract

The present invention discloses a fine classification method for complex targets based on a dual-channel attention mechanism. The implementation steps are as follows: construct a dual-channel attention dense connection network, perform dynamic weighting processing on the features output by the feature extraction module through autonomous learning, and adopt multi-scale pooling operations in the process of obtaining the channel feature weights; generate a training set; train the dual-channel attention dense connection network; classify complex scene images. The present invention overcomes the deficiencies of the prior art in that it is prone to errors in classifying confusing complex scenes and will ignore the effective information of small scales in images. By constructing a dual-channel attention dense connection network, the present invention can still maintain a high classification accuracy for classifying confusing scenes and can completely extract image information in the face of complex images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and further relates to a fine classification method for complex objects based on a dual-channel attention mechanism in the field of image classification technology. The present invention is used for fine classification of complex scene objects and can be applied to working scenarios such as e-commerce, medical image processing, artificial intelligence, and high-resolution image processing. Background Art

[0002] Due to the diversity of real-world scenes, there is too much background information in the complex scene target area, and the existing technologies have certain limitations and cannot effectively solve the problem of complex object classification. For complex scenes, the current technologies face challenges in the construction of the sample library, the design of the network model, the training method, the network parameters, and the classification method. The attention mechanism conforms to the human recognition mechanism in various different types such as image classification and image recognition. The attention mechanism model can automatically focus on the salient region features of the classification target during the classification process, assign different degrees of attention to different features, enable the classification to focus on important features and suppress unnecessary features, and can effectively improve the classification effect.

[0003] Nanjing University of Posts and Telecommunications disclosed an improved algorithm for image classification based on a convolutional neural network in its patent document "Improved Method for Image Classification Based on Convolutional Neural Network" (application number CN 201910624323.0, publication number CN 110321967 B). It uses the AlexNet network model as the basic framework, first performs appropriate preprocessing and data augmentation on the input image, extracts features through the convolutional layer of the neural network, then retains the main features through the pooling layer while reducing the parameters and computational amount of the next layer, uses the method of multi-scale convolution to make the network model no longer limited by the size of the input image, and further reduces the dimension of the feature map using the LDA algorithm, and finally obtains the predicted classification of the image. However, the disadvantage of this method is still that for the classification of some easily confused complex scenes, such as grasslands, farmlands, forests, etc., when the sample size is small, the problem of reduced classification accuracy will occur due to the inability to fully learn the scene features.

[0004] Guilin University of Electronic Technology discloses an image classification method based on an attention mechanism in its patent document "An Image Classification Method Based on an Attention Mechanism" (application number CN 202110517855.1, publication number CN 113408577 A). The method includes performing frequency decomposition on each channel of the feature map based on the discrete cosine transform and jointly representing the global information of the channel with multiple frequency components, and then calculating the channel attention weight information; obtaining the channel attention mechanism by weighting each channel of the feature map based on the weight information, then calculating the spatial attention weight of each pixel of the feature map, and then weighting and summing each spatial pixel of the feature map to obtain the spatial attention mechanism; embedding the channel attention mechanism and the spatial attention mechanism into ResNet to obtain an image classification convolutional neural network and training it. However, the deficiencies of this method are still as follows: when classifying images with complex targets, the process of feature extraction only uses a single pooling, which will ignore the small-scale information of the image and the extraction of image information is not complete enough. In addition, its network structure is relatively simple, which affects the classification accuracy. Summary of the Invention

[0005] The object of the present invention is to aim at the deficiencies of the above-mentioned prior art and propose a fine classification method for complex targets based on a dual-channel attention mechanism, aiming to solve the problems that the prior art cannot fully learn scene features in the classification of confusing scenes and will ignore the small-scale information of the image, and the extraction of image information is not complete enough.

[0006] The idea for achieving the object of the present invention is as follows: The present invention constructs a dual-channel attention dense connection network, which uses a single convolutional layer and a max pooling layer, then adds four dense blocks, and adds a transition layer with the number of channels halved between the dense blocks to extract image features. Then, an attention mechanism module is added, and finally a global pooling layer and a fully connected layer are connected to output. The attention mechanism module dynamically weights the features through self-learning, and is formed by serially combining the spatial attention mechanism and the channel attention mechanism, so that more useful information can be obtained in terms of space and channels, enhancing the extraction of useful information of key features of high-resolution remote sensing images, thereby achieving adaptive feature refinement and improving the feature representation ability of the model, so as to solve the problem that the prior art cannot fully learn scene features in the classification of easily confused scenes. When the present invention dynamically weights the features at the channel level, multi-scale pooling operations are used to obtain channel attention weights. The present invention performs three max pooling operations on the image after feature extraction by the convolutional pooling layer and the dense block at different sizes based on the channels, and the pooled sizes are 1×1×C, 2×2×C, and 4×4×C respectively. The feature maps with sizes of 2×2×C and 4×4×C are then globally averaged pooled, and the three finally obtained 1×1×C features are added and input into a multi-layer perceptron, and are mapped through Sigmoid to generate channel feature weights, so that multi-scale image information can be obtained, thereby solving the problem that the prior art ignores the small-scale effective information of the image and the image information extraction is not complete enough.

[0007] The specific steps for achieving the object of the present invention are as follows:

[0008] Step 1, construct a dual-channel attention dense connection network:

[0009] Step 1.1, build a feature extraction module and set the parameters of the feature extraction module;

[0010] Step 1.2, build a channel attention module composed of a series connection of a pooling module and a multi-layer perceptron. Among them, the pooling module is formed by connecting the second max pooling layer and the third max pooling layer in series with the mean pooling layer respectively and then connecting them in parallel with the first max pooling layer. The multi-layer perceptron is formed by connecting the first linear layer, the first activation layer, the second linear layer, and the second activation layer in series. The first to third max pooling layers use multi-scale pooling operations, and the output sizes are set to 1×1, 2×2, and 4×4 respectively. The output size of the mean pooling layer after pooling is set to 1×1;

[0011] The input number of the first linear layer of the multi-layer perceptron is set to 248, and the output number is set to 496; the input number of the second linear layer is set to 496, and the output number is set to 248; the first activation layer is implemented using the ReLU activation function; the second activation layer uses the Sigmoid activation function to obtain the channel feature weights; the channel feature weights are multiplied by the input of the channel attention module to obtain the output of the channel attention module;

[0012] Step 1.3: Build a spatial attention module, the structure of which includes a max-pooling layer, an average-pooling layer, a convolutional layer, and an activation layer. Among them, after concatenating the max-pooling layer and the average-pooling layer on the channel dimension, they are successively connected in series with the convolutional layer and the activation layer; set the output size of the max-pooling layer after pooling to 1×1; set the output size of the average-pooling layer after pooling to 1×1; set the number of input channels of the convolutional layer to 2, the number of output channels to 1, and the convolutional kernel size to 7×7; use the Sigmoid activation function in the activation layer to obtain the spatial feature weights; multiply the spatial feature weights by the input of the spatial attention module to obtain the output of the spatial attention module.

[0013] Step 1.4: Connect the feature extraction module, the channel attention module, the spatial attention module, and the classification module in series successively to form a dual-channel attention dense connection network.

[0014] Step 2: Generate a training set:

[0015] Step 2.1: Select at least 2800 complex scene images to form a sample set, which includes at least 7 scene types.

[0016] Step 2.2: Reset the size of all images in the sample set to 400×400.

[0017] Step 2.3: Label the scene type corresponding to each image in the sample set; form the labeled images into a training set.

[0018] Step 3: Input the training set into the dual-channel attention dense connection network, use the gradient descent method to iteratively update the network parameters, and perform dynamic weighting processing on the features output by the feature extraction module through self-learning to obtain a trained dual-channel attention dense connection network.

[0019] Step 4: Classify complex scene images:

[0020] Input the complex scene image to be classified into the trained dual-channel attention dense connection network, and output the classification label of the image.

[0021] The present invention has the following advantages compared with the existing technology:

[0022] First, since the present invention constructs a dual-channel attention dense connection network, it enables the classification to focus on important features and suppress unnecessary features, refines the image features, improves the model's expressiveness, overcomes the deficiency of the existing technology in being unable to fully learn scene features when classifying confusing scenes, and enables the present invention to still maintain a high classification accuracy for classifying confusing scenes.

[0023] Second, when the present invention performs dynamic weighted processing on features at the channel level, a multi-scale pooling operation is used to obtain channel attention weights, and multi-scale image information can be obtained, overcoming the deficiency of the prior art that small-scale effective information in images will be ignored. When the present invention classifies high-resolution remote sensing images with complex and diverse targets, it can completely extract image information, thereby improving the accuracy of image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of the present invention;

[0025] Figure 2 is a schematic structural diagram of a dual-channel attention dense connection network in the present invention;

[0026] Figure 3 is a schematic structural diagram of a channel attention module in the present invention;

[0027] Figure 4 is a schematic structural diagram of a spatial attention module in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0028] The present invention will be further described below with reference to the drawings and embodiments.

[0029] Refer to Figure 1 , and the implementation steps of the embodiments of the present invention will be further described.

[0030] Step 1, construct a dual-channel attention dense connection network.

[0031] Refer to Figure 2 , and a further detailed description will be made on the dual-channel attention dense connection network constructed by the present invention.

[0032] Step 1.1, build a feature extraction module, and its structure is sequentially connected in series as: input layer, convolutional layer, activation layer, max pooling layer, first dense block, first transition layer, second dense block, second transition layer, third dense block, third transition layer, fourth dense block. Among them, the first to fourth dense blocks are respectively composed of four convolutional layers connected in series, and the first to third transition layers are composed of a convolutional layer and an average pooling layer connected in series. Set the number of channels of the input layer to 3 and the size to 400×400. Set the size of the convolutional kernel of the convolutional layer to 7×7, the number of channels to 64, and the stride to 2. The activation layer is implemented using the ReLU activation function. Set the size of the max pooling layer to 3×3 and the stride to 2. Set the number of convolutional layers in the first to fourth dense blocks to 4, and the number of input and output channels of each convolutional layer to 200, 100, 50, and 25 respectively. Set the size of the convolutional kernel of the convolutional layer in the first to third transition layers to 1×1, and the pooling stride of the average pooling layer to 2.

[0033] Step 1.2: Build a channel attention module composed of a pooling module and a multi-layer perceptron in series, as shown in Figure 3 . Among them, the pooling module is formed by connecting the second max-pooling layer, the third max-pooling layer, and the average pooling layer in series and then connecting them in parallel with the first max-pooling layer. The multi-layer perceptron is formed by connecting the first linear layer, the first activation layer, the second linear layer, and the second activation layer in series. The first to third max-pooling layers adopt multi-scale pooling operations, and the output sizes are set to 1×1, 2×2, and 4×4 respectively. The output size after pooling of the average pooling layer is set to 1×1. The number of input channels of the first linear layer of the multi-layer perceptron is set to 248, and the number of output channels is set to 496. The number of input channels of the second linear layer is set to 496, and the number of output channels is set to 248. The first activation layer is implemented using the ReLU activation function. The second activation layer uses the Sigmoid activation function to obtain the channel feature weights, and the channel feature weights are multiplied by the input of the channel attention module to obtain the output of the channel attention module.

[0034] Step 1.3: Build a spatial attention module, as shown in Figure 4 . Its structure includes a max-pooling layer, an average pooling layer, a convolutional layer, and an activation layer. Among them, the max-pooling layer and the average pooling layer are concatenated on the channel dimension and then connected in series with the convolutional layer and the activation layer in sequence. The output size after pooling of the max-pooling layer is set to 1×1, and the output size after pooling of the average pooling layer is set to 1×1. The number of input channels of the convolutional layer is set to 2, and the number of output channels is set to 1. The size of the convolutional kernel is set to 7×7. The activation layer uses the Sigmoid activation function to obtain the spatial feature weights, and the spatial feature weights are multiplied by the input of the spatial attention module to obtain the output of the spatial attention module.

[0035] Step 1.4: Connect the feature extraction module, the channel attention module, the spatial attention module, and the classification module in series to form a dual-channel attention dense connection network. The classification module is composed of a global pooling layer and a fully connected layer connected in series. Its parameter settings are as follows: the size after pooling of the global pooling layer is set to 1×1; the number of input channels of the fully connected layer is set to 25, and the number of output channels is set to 7.

[0036] Step 2: Generate a training set and a test set.

[0037] In the embodiment of the present invention, all samples of the RSSCN7 high-resolution remote sensing image dataset are selected to form a sample set, which contains a total of 2800 complex scene images. The size of each image is 400×400 pixels. The sample set contains 7 categories, namely grassland, forest, farmland, parking lot, residential area, industrial area, and river / lake. Each category contains 400 images. 320 images are selected from each category, a total of 2240 images are used to form the training set, and 80 images are selected from each category, a total of 560 images are used to form the test set.

[0038] Step 3: Train a dual-channel attention dense connection network.

[0039] Input the training set into the deep learning network, calculate the loss value between the predicted classification label output by the network and the actual classification label, and use the adaptive moment estimation algorithm to iteratively update the network parameters until the network loss function converges, obtaining a trained dual-channel attention dense connection network.

[0040] The loss function is as follows:

[0041]

[0042] Among them, S(θ) represents the loss value between the predicted label and the true label, n represents the total number of samples in the training set, q represents the total number of sample categories in the training set, represents the one-dimensional predicted probability distribution vector output by the network, represents the true one-dimensional label probability distribution vector.

[0043] Step 4: Classify the test set samples.

[0044] Input the samples of the test set into the trained network to output the classification labels of the samples.

[0045] The following further illustrates the effect of the present invention in combination with simulation experiments:

[0046] 1. Simulation experiment conditions:

[0047] The hardware platform for the simulation experiment of the present invention is: the processor is an Intel Core i5-8300H CPU with a main frequency of 2.3 GHz and 8 GB of memory.

[0048] The software platform for the simulation experiment of the present invention is: Windows 10 operating system and Pycharm2022.

[0049] The dataset for the simulation experiment of the present invention uses the RSSCN7 dataset, which is a dataset publicly disclosed by Qin Zou et al. from Wuhan University in their published paper "Deep learning based feature selection for remote sensing scene classification" ([J] / / Proceedings of the IEEE Geoscience and Remote Sensing Letters. 2015: 2321-2325).

[0050] 2. Simulation content and its result analysis:

[0051] The simulation experiment of the present invention uses the method of the present invention and three existing technologies (S-head-attention, M-head-attention, pre-trained Resnet50 features + SVM) to classify all samples in the RSSCN7 high-resolution remote sensing image dataset respectively.

[0052] In the simulation experiment of the present invention, S-head-attention and M-head-attention refer to:

[0053] The method based on self-attention convolutional neural network proposed by Li Yanfu, Fan Xijian, Yang Xubing et al. in their published paper "Remote sensing image classification framework based on self-attention convolutional neural network" ([J] / / Proceedings of the Journal of Beijing Forestry University. 2021, 43(10): 81-88).

[0054] In the simulation experiment of the present invention, pre-trained Resnet50 features + SVM refers to:

[0055] The method based on transfer learning proposed by Gong Xi, Chen Zhanlong, Wu Liang et al. in their published paper "Transfer learning based mixture of experts classification model for high-resolution remote sensing scene classification" ([J] / / Proceedings of the Acta Optica Sinica. 2021, 41(23): 2301003).

[0056] In order to verify the effect of the simulation experiment of the present invention, use the formula Calculate the classification accuracy of the method of the present invention and the three existing technologies respectively. As shown in Table 1, where Acc represents the classification accuracy, R represents the number of samples correctly classified in the experiment, and N represents the total number of experimental samples.

[0057] Table 1 Classification accuracy results of the present invention and existing technologies in the simulation experiment

[0058] Classification method Classification accuracy S-head-attention 88.14% M-head-attention 91.31% pre-trainedResnet50features+SVM 89.92% The method of the present invention 93.39%

[0059] As can be seen from Table 1, the classification accuracy of the method of the present invention has been significantly improved compared with the prior art.

Claims

1. A fine classification method for complex targets based on a dual-channel attention mechanism, characterized in that, Construct a dual-channel attention dense connection network that includes a channel attention module and a spatial attention module. Dynamically weight the features output by the feature extraction module through self-learning, and adopt multi-scale pooling operations in the process of obtaining channel feature weights. The steps of this classification method are as follows: Step 1, construct a dual-channel attention dense connection network: Step 1.1, build a feature extraction module and set the parameters of the feature extraction module; Step 1.2, build a channel attention module composed of a pooling module and a multi-layer perceptron in series. Among them, the pooling module is composed of a second max-pooling layer, a third max-pooling layer, and an average pooling layer in series and then in parallel with the first max-pooling layer. The multi-layer perceptron is composed of a first linear layer, a first activation layer, a second linear layer, and a second activation layer in series. The first to third max-pooling layers adopt multi-scale pooling operations, and the output sizes are set to 1×1, 2×2, and 4×4 respectively. The output size after pooling of the average pooling layer is set to 1×1. The input number of the first linear layer of the multi-layer perceptron is set to 248, and the output number is set to 496. The input number of the second linear layer is set to 496, and the output number is set to 248. The first activation layer is implemented using the ReLU activation function. The second activation layer uses the Sigmoid activation function to obtain the channel feature weights. The channel feature weights are multiplied by the input of the channel attention module to obtain the output of the channel attention module; Step 1.3, build a spatial attention module. Its structure includes a max-pooling layer, an average pooling layer, a convolutional layer, and an activation layer. Among them, the max-pooling layer and the average pooling layer are concatenated on the channel and then in series with the convolutional layer and the activation layer in turn. The output size after pooling of the max-pooling layer is set to 1×1. The output size after pooling of the average pooling layer is set to 1×1. The input channel number of the convolutional layer is set to 2, the output channel number is set to 1, and the convolutional kernel size is set to 7×7. The activation layer uses the Sigmoid activation function to obtain the spatial feature weights. The spatial feature weights are multiplied by the input of the spatial attention module to obtain the output of the spatial attention module; Step 1.4, sequentially connect the feature extraction module, the channel attention module, the spatial attention module, and the classification module in series to form a dual-channel attention dense connection network; Step 2, generate a training set: Step 2.1, select at least 2800 complex scene images to form a sample set, and this sample set includes at least 7 scene types; Step 2.2, reset the size of all images in the sample set to 400×400; Step 2.3, label the scene type corresponding to each image in the sample set; form the labeled images into a training set; Step 3, input the training set into the dual-channel attention dense connection network, use the gradient descent method to iteratively update the network parameters, and dynamically weight the features output by the feature extraction module through self-learning to obtain a trained dual-channel attention dense connection network; Step 4, classify complex scene images: Input the complex scene image to be classified into the trained dual-channel attention dense connection network, and output the classification label of the image.

2. The fine classification method for complex targets based on the dual-channel attention mechanism according to claim 1, wherein, The structures of the feature extraction module described in Step 1.1 are connected in series in the following order: input layer, convolutional layer, activation layer, max pooling layer, first dense block, first transition layer, second dense block, second transition layer, third dense block, third transition layer, fourth dense block. Among them, the first to fourth dense blocks are each composed of four convolutional layers connected in series, and the first to third transition layers are composed of a convolutional layer and an average pooling layer connected in series.

3. The fine classification method for complex objects based on the dual-channel attention mechanism according to claim 2, characterized in that The parameters of the feature extraction module described in Step 1.1 are as follows: set the number of channels of the input layer to 3 and the size to 400×400; set the size of the convolutional kernel of the convolutional layer to 7×7, the number of channels to 64, and the stride to 2; the activation layer is implemented using the ReLU activation function; set the size of the max pooling layer to 3×3 and the stride to 2; set the number of convolutional layers in the first to fourth dense blocks to 4, and the number of input and output channels of each convolutional layer to 200, 100, 50, and 25 respectively; set the size of the convolutional kernel of the convolutional layer in the first to third transition layers to 1×1 and the pooling stride of the average pooling layer to 2.

4. The fine classification method for complex targets based on a dual-channel attention mechanism according to claim 1, characterized in that The classification module described in Step 1.4 is composed of a global pooling layer and a fully connected layer connected in series in that order. Its parameter settings are as follows: set the size after pooling of the global pooling layer to 1×1; set the number of inputs of the fully connected layer to 25 and the number of outputs to 7.

5. The fine classification method for complex objects based on the dual-channel attention mechanism according to claim 1, characterized in that The 7 scenario types described in Step 2.1 refer to grassland, forest, farmland, parking lot, residential area, industrial area, and rivers and lakes.

6. The fine classification method for complex objects based on the dual-channel attention mechanism according to claim 1, wherein The gradient descent method described in Step 3 refers to: inputting the training set into the deep learning network, calculating the loss value between the predicted classification label output by the network and the actual classification label, and using the adaptive moment estimation algorithm to iteratively update the network parameters until the network loss function converges, obtaining a trained dual-channel attention dense connection network.

7. The fine classification method for complex targets based on the dual-channel attention mechanism according to claim 6, wherein The loss function is as follows: Among them, S(θ) represents the loss value between the predicted label and the true label, n represents the total number of samples in the training set, and q represents the total number of sample categories in the training set. represents the one-dimensional predicted probability distribution vector output by the network. represents the true one-dimensional label probability distribution vector.

Citation Information

Patent Citations

  • Improved algorithm for image classification based on convolutional neural network

    CN110321967A

  • An Improved Image Classification Method Based on Convolutional Neural Networks

    CN110321967B

  • Image classification method based on attention mechanism

    CN113408577A

  • A method and apparatus for expression recognition

    CN109002766A

  • Remote sensing image classification method based on attention mechanism deep Contourlet network

    CN110728224A