Hyperspectral object classification method based on large kernel attention mechanism and MLP hybrid

By using the large-core attention mechanism module and feature enhancement module with DenseNet structure in hyperspectral remote sensing image classification, the problem of difficulty in effectively mining spatial region and spectral band information in the prior art is solved, and a higher classification accuracy and a more efficient training process is achieved.

CN116630723BActive Publication Date: 2025-05-16XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310791850.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-05-16
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

The existing hyperspectral remote sensing image classification technology is difficult to effectively dig information on spatial regions and spectral bands, resulting in low classification accuracy and long training time.

Method used

The hyperspectral land object classification method of the large-nuclear attention mechanism module and the feature enhancement module based on the DenseNet structure is adopted to extract spatial and spectral information through the large-nuclear attention mechanism module, and information interaction and fusion are achieved through the feature enhancement module.

Benefits of technology

It improves the classification accuracy of hyperspectral remote sensing images, reduces the amount of parameters, simplifies the model training process, improves classification efficiency, and performs better than existing methods in simulation experiments on multiple data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630723B_ABST
    Figure CN116630723B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral feature classification method based on a mixture of a large kernel attention mechanism and MLP, which mainly solves the problem that the existing technology has weak spatial and spectral feature extraction capabilities and cannot fully utilize neighborhood information. Its implementation scheme is: obtain a hyperspectral data set and a data annotation set from a public website, and divide the training set and test set samples; construct a large kernel attention mechanism module and a feature enhancement module based on the DenseNet structure respectively, and connect the two in series to form a feature extraction network; use the training set to calculate the loss of the entire network through a multi-classification focal loss function, and use the stochastic gradient descent method to iteratively optimize the network parameters; input the test set into the trained extraction network to obtain the classification results of the hyperspectral image. The present invention can fully explore the complex content in remote sensing images, enhance the feature extraction capabilities of image space and spectrum, greatly improve the image classification effect, and can be used for urban development, resource exploration and environmental monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to a hyperspectral remote sensing image ground object classification method, which can be used for urban development, resource exploration and environmental monitoring. Background Art

[0002] Hyperspectral remote sensing image classification plays a vital role in the field of remote sensing and has attracted more and more attention due to its wide application. Existing hyperspectral classification techniques are mainly divided into two categories.

[0003] The first category is the hyperspectral classification method based on machine learning, that is, traditional machine learning algorithms such as support vector machine (SVM) and random forest (RF) have been widely used to accurately classify objects in remote sensing images. The features they extract mainly include shallow features, such as texture features, shape features, and color features. Due to the complexity and diversity of the scenes involved in hyperspectral remote sensing images, and the problems of "same object, different spectrum" and "same spectrum, different objects" in hyperspectral images, this type of method cannot well represent the complex content in hyperspectral remote sensing images.

[0004] The second category is the hyperspectral remote sensing image object classification method based on deep learning. This type of method is currently represented by graph convolutional neural network methods and Transformer-based methods. This type of method can extract high-level semantic features in remote sensing images and grasp the relationship between pixels. Therefore, the classification effect is good and it is currently a widely used hyperspectral remote sensing image classification method. Although the method based on deep convolutional neural network has achieved great success in the hyperspectral remote sensing image classification task, there are still some problems that need to be further improved. First, the method based on deep convolutional neural network is good at capturing global information from remote sensing images, but cannot thoroughly mine the local knowledge hidden in complex remote sensing images and the band information from the spectrum; secondly, the method based on graph convolutional neural network will be disturbed by the number of data division blocks due to the specific hyperspectral remote sensing data division, and the number of blocks for remote sensing data division is usually small. This leads to the misuse of a lot of noisy information in feature extraction. Although the head attention mechanism based on Transformer can capture different positional relationships. However, the information in the spectral channels of the hyperspectral image is not fully understood; at the same time, due to the extensiveness of the feature mapping in the Transformer, the multi-layer Transformer requires a large amount of memory and time during the training phase, which is not conducive to specific applications in actual scenarios. In addition, since most existing Transformer networks use flattening operations and linear projections to sequentially encode spatial information, they cannot effectively utilize local spatial spectral information and position information.

[0005] In view of the problems and defects in the existing hyperspectral remote sensing image classification technology, how to quickly mine the spatial regions and spectral bands that are useful for the classification task so that the model can achieve reasonable classification accuracy for each category is a difficult problem that technical personnel in this field need to solve.

[0006] Zhu Minghao et al. proposed a new hyperspectral classification method in the Institute of Electrical and Electronics Engineers (IEEE). This method organically combines the residual network ResNet with the convolutional attention module, and uses the spectral attention module and the spatial attention module to enhance the classification performance. Through the combination of the spectral attention module and the spatial attention module, useful bands and spatial information can be more effectively identified, thereby completing the classification faster. This method can effectively capture the key areas and spectral bands in the image, greatly improving the classification efficiency of the network. At the same time, since this method embeds the attention module into the residual block, it helps to effectively reduce the problem of overfitting. Its disadvantage is that the training model requires a large amount of data to support it, and the training time is too long.

[0007] Meng Fanbo and others proposed a hyperspectral image classification method 3DOC-SSAN based on 3D Octave convolution and spatial-spectral attention. This method first uses four 3D Octave convolutions to capture spatial-spectral features from the high and low frequency aspects of the image. Secondly, two attention models are introduced from the spatial and spectral dimensions to highlight the spatial region and specific spectral bands in the classification task. Finally, an information complementation model is designed to transfer important information between spatial and spectral attention features, and the spatial-spectral features that are beneficial to the hyperspectral classification task are integrated through the information complementation model. However, due to the use of 3D convolution operations in this paper, although this method improves the classification accuracy of hyperspectral images, it takes a lot of time to train and has high computational complexity.

[0008] Hong Danfeng et al. developed a new model called SpectralFormer, which designed two simple but effective modules, namely grouped spectral embedding GSE and cross-layer adaptive fusion CAF, to form a cross-layer encoder TE module. This module learns the local detailed spectral band representation on the pixel and fuses the shallow features to the deep features, improving the spectral information processing mode and learning the spectral representation information from the grouped adjacent bands. Although this method performs well in capturing spectral features, it does not have the ability to effectively capture the local semantic features of hyperspectral images, which leads to the failure to fully utilize the spatial information of hyperspectral remote sensing images, resulting in the classification efficiency can not be further improved. Summary of the invention

[0009] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a hyperspectral object classification method based on a large kernel attention mechanism and MLP hybrid to fully extract the spatial and spectral features of hyperspectral images, reduce the number of parameters, and improve classification efficiency.

[0010] To achieve the above object, the technical solution of the present invention comprises the following steps:

[0011] (1) Construct training sample set and test sample set:

[0012] 1a) Obtain the original data set and data annotation set of hyperspectral images from the public website:

[0013] 1b) Select proportional values ​​from each non-zero category in the data annotation set, save the corresponding position coordinates of these values ​​in the data annotation set, find the pixel points at the corresponding coordinate positions in the original data set, and perform mirror segmentation with each pixel point as the center and the set imgsize parameter as the diameter to generate a training set;

[0014] 1c) Generate a test set using the method in 1b) for all remaining samples;

[0015] (2) Build a feature extraction network:

[0016] 2a) A large-core attention mechanism module based on the DenseNet structure is established, which includes three DenseBlock sub-modules, to effectively extract information at different spatial positions and different spectral bands of the input image.

[0017] 2b) establishing a feature enhancement module consisting of a spatial mixing MLP layer and a channel mixing MLP layer to realize information interaction between the spectral dimension and the spatial dimension;

[0018] 2c) Connect the large core attention mechanism module based on the DenseNet structure and the feature enhancement module in series to form a feature extraction network, and use the multi-classification focus loss function as the loss function of the extraction network. FL ;

[0019] (3) Train the feature extraction network:

[0020] 3a) Input the training set into the feature extraction network and calculate its loss FL value;

[0021] 3b) Using the gradient descent method, gradually reduce the value of the loss function to update the network parameters until the set maximum number of iterations is completed to obtain a trained feature extraction network;

[0022] (4) Input the test set into the trained feature extraction network to obtain the output vector. Use the softmax function and argmax function in the output vector to obtain the position index of the maximum value. The position index is the final classification result of each test sample.

[0023] Compared with the prior art, the present invention has the following advantages:

[0024] 1) The present invention constructs a feature extraction network composed of a large core attention mechanism module based on the DenseNet structure and a feature enhancement module connected in series, which can mine different spatial information and spectral band information in hyperspectral remote sensing images and realize hyperspectral classification in an end-to-end manner. It is not only simple to operate, but also can fully interpret the complex content in hyperspectral remote sensing images.

[0025] 2) In the feature extraction network of the present invention, a large-core attention mechanism module based on the DenseNet structure is established. By jointly mining the long-range relationship and spectral information between pixels in a patch, an attention map is generated to measure the importance of all pixels in a patch. The DenseNet structure can not only effectively alleviate the problem of gradient disappearance, but also enhance feature reuse, and can fully tap the potential of the network model for feature extraction.

[0026] 3) In the feature extraction network of the present invention, since a feature enhancement module is established, the spatial features at different spatial positions and the spectral features between different channels can be communicated by alternately executing two different types of layers, TokenMixing and Channel Mixing, and information interaction between the two dimensions can be promoted, thereby realizing information fusion in the spatial domain and the spectral domain, enhancing more information features, and greatly improving the classification performance.

[0027] Simulation experiments show that the classification accuracy of the present invention is better than other existing methods, and the overall classification effect is more stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the implementation process of the present invention;

[0029] Figure 2 It is a schematic diagram of the structure of the large core attention mechanism module based on the DenseNet structure in the present invention;

[0030] Figure 3 yes Figure 2 Schematic diagram of the attention layer structure in the large core attention mechanism module based on the DenseNet structure;

[0031] Figure 4 It is a schematic diagram of the structure of the feature enhancement module in the present invention. DETAILED DESCRIPTION

[0032] The examples and effects of the present invention are further described in detail below with reference to the accompanying drawings.

[0033] Reference Figure 1 This example is based on a hyperspectral feature classification method based on a mixture of CNN and MLP, including constructing a data set, building a feature extraction network, and image classification. The specific implementation steps are as follows:

[0034] Step 1: Construct training set and test set.

[0035] 1.1) Obtain the original data set and data annotation set of the hyperspectral images on the public website, and select proportional values ​​from each category that is not 0 in the data annotation set, and save the position coordinates of these values ​​corresponding to the data annotation set.

[0036] 1.2) Find the pixel point with the corresponding coordinate position in the original data set, and perform mirror segmentation with each pixel P as the center and the parameter set by imgsize as the diameter to construct a 3D space block P i ∈R d×d×B As the final training set, d represents the width and length of the image, and B represents the number of channels. In this example, d is set to 23, B is set to 200, and the imgsize parameter is set to 23;

[0037] 1.3) Obtain the test set by following the pattern of step 1.2) for all remaining samples.

[0038] Step 2: Build a feature extraction network.

[0039] 2.1) Construct a large core attention mechanism module based on the DenseNet structure:

[0040] Reference Figure 2 , the implementation of this step is as follows:

[0041] 2.1.1) Create a DenseBlock submodule consisting of a transition layer, two convolutional integration layers, two attention layers, and two concatenation layer submodules:

[0042] 2.1.1.1) Set the parameters and structure of each layer:

[0043] The transition layer is composed of a convolution layer with a convolution kernel size of 1×1 and a pooling layer with a size of 2 connected in sequence;

[0044] Each convolutional integration layer consists of a convolutional layer with a convolution kernel size of 1×1, a convolutional layer with a convolution kernel size of 3×3, and a batch processing layer cascaded;

[0045] Reference Figure 3Each attention layer consists of a spatial local convolution block with a convolution kernel size of (2d-1)×(2d-1), a convolution kernel size of The spatial long-range convolution block and a channel convolution block with a convolution kernel size of 1×1 are connected in series, where d is the expansion rate and K represents the length and width of the input feature;

[0046] 2.1.1.2) Establish 4-way connection relationship between each layer:

[0047] The first path: transition layer → first convolutional integration layer → first attention layer → first splicing layer → second convolutional integration layer → second attention layer → second splicing layer;

[0048] The second path: transition layer → first splicing layer → second convolutional integration layer → second attention layer → second splicing layer;

[0049] The third route is: transition layer → second splicing layer;

[0050] The fourth path is: first attention layer → second splicing layer.

[0051] The image transmission relationship of the 4 channels is:

[0052] In the first path, the input image X outputs feature F after passing through the transition layer, and the feature F outputs feature F1 after passing through the first attention layer. Feature F and F1 pass through the first splicing layer and output feature F2. At the same time, F1 continues to enter the second convolutional integration layer and the second attention layer, and the output feature is F3, which is the feature obtained before the first path inputs the second splicing layer.

[0053] In the second path, the obtained feature F2 is input into the second convolutional integration layer, and after passing through the second attention layer, the feature F4 is output, that is, the feature obtained before the second path inputs the second concatenation layer.

[0054] In the third path, the output feature F of the transition layer is the feature obtained before the third path is input to the second concatenation layer.

[0055] In the fourth path, the first attention layer d outputs feature F1, which is the feature obtained before the fourth path is input to the second concatenation layer;

[0056] The four features F3, F4, F, F1 are concatenated in the second layer to finally obtain a feature Fout1 output in a Denseblock.

[0057] 2.1.2) Cascade three DenseBlock sub-modules with the same structure in sequence to form a large core attention mechanism module.

[0058] 2.2) Establish feature enhancement module:

[0059] 2.2.1) Create a spatial hybrid MLP layer consisting of a spatial transposition operation and a first fully connected layer, whose output is:

[0060] U *,i =X *,i +W2σ(W1LayerNorm(X) *,i ),

[0061] In the formula, X represents the transposed feature, U *,i represents the features after spatial hybrid MLP, LayerNorm represents layer normalization, σ represents the activation function, and X *,i Columns representing features;

[0062] 2.2.2) Create a channel-mixed MLP layer consisting of a channel transposition operation and a second fully connected layer, whose output is:

[0063] Y j,* =U j,* +W4σ(W3LayerNorm(X) j,* ),

[0064] Where U j,* Indicates U *,i The feature after transposition, X j,* Rows representing features

[0065] 2.2.3) Reference Figure 4 , the spatial mixed MLP layer and the channel mixed MLP layer are connected, that is, spatial transposition operation → first fully connected layer → channel transposition operation → second fully connected layer, to form a relational feature enhancement module, where the dimension of each fully connected layer is 128.

[0066] 2.3) Connect the large core attention mechanism module based on the DenseNet structure and the feature enhancement module in series to form the entire feature extraction network;

[0067] 2.4) Use the multi-classification focal loss function as the loss function of the extraction network Loss FL , which is expressed as:

[0068]

[0069] Among them, N is the number of samples, K is the total number of categories, and y i represents the true label category value of sample i, γ is a preset value, and P ik The probability value of predicting the Kth category for the i-th sample.

[0070] Step 3: Train the feature extraction network.

[0071] 3.1) Input the training set constructed in 1.1) into the feature extraction network and calculate the loss function Loss between the output of the feature extraction network and the original labels in the training set FL ;

[0072] 3.2) Use gradient descent method to update the feature network parameters θ:

[0073] 3.2.1) Set the initial training batch m = 0 and the maximum number of iterations T = 200;

[0074] 3.2.2) Calculate the feature extraction network parameters θ after the current training update m+1 :

[0075]

[0076] Among them, α is the learning rate during the training phase, θ m is the network parameter before the current training update, and L(·) is the derivative of the parameters in the current network;

[0077] 3.2.3) Repeat step 3.2.2) until the maximum number of iterations T in the training phase is reached, completing the training of the feature extraction network.

[0078] Step 4: Use the test set to obtain classification results.

[0079] 4.1) Input the test set into the trained feature extraction network model to output the feature vector F;

[0080] 4.2) Use the softmax function to convert the values ​​in the feature vector F to between [0-1], and then use the argsmax function to calculate the index of the maximum value in the vector F: F out =argmax(softmax(F)), where F out The final result, F∈1×K is the output feature vector, K is the total number of categories, and the index value is the classification category of each pixel in the image.

[0081] The effect of the present invention can be further illustrated by the following simulation:

[0082] 1. Simulation conditions

[0083] The simulation environment of the present invention selected the framework of python 3.8+pytorch 1.7 and was completed on a workstation with GeForce RTX2080Ti and 11G memory.

[0084] The four datasets used in the simulation are Indian Pines dataset, PaviaU dataset, Houston2013 dataset and Salinas dataset.

[0085] The Indian Pines dataset was collected by the airborne visible / infrared imaging spectrometer AVIRIS sensor in a farmland area in northwestern Indiana. The original hyperspectral image wavelength range is between 0.4-2.5 microns. The spatial resolution is 20m. The selected area size contains 145×145 pixels. After noise removal and atmospheric correction, 200 spectral bands are selected for experiments. The dataset contains a variety of objects and materials in real scenes such as coniferous forests and farmlands, such as corn, soybeans, and pine trees. There are 16 different ground object categories, including a total of 10,249 manually labeled data samples, and the remaining 10,776 pixels are background pixels.

[0086] The paivaU dataset is a hyperspectral remote sensing image dataset obtained in the area near the University of Pavia using the reflective optical system imaging spectrometer ROSIS-3HS sensor. The spectral coverage range is between 430-860nm, and the geometric resolution of the pixel is 1.3m. After removing the bands affected by noise, 103 spectral bands were selected for the experiment. The area size contains 610×340 pixels, of which there are a total of 42,776 labeled pixels, which are divided into 9 land categories. Mainly including asphalt, grass, gravel and other objects.

[0087] The Houston2013 dataset was taken in 2012 at the University of Houston and its vicinity using the airborne sensor ITRES-CASI1500. After calibration, 144 spectral bands were selected for the experiment. Its spatial resolution is 2.5m. The dataset was published in the IEEE Geoscience and Remote Sensing Society Data Fusion Competition in 2013. The area is very complex. For the consistency of the experiment, the training set and the test set of the provided dataset were fused, and the obtained dataset contained 349×1905 pixels. There are 15029 labeled pixels in total. It contains 15 categories in total.

[0088] The Salinas dataset was taken by the AVIRIS sensor in the Salinas Valley of California, with a spatial resolution of 3.7 meters and 224 continuous bands. After removing 20 water-absorbing bands (108-112, 154-167, 224), the actual band used for training is 204. The area contains 512×340 pixels, mainly including 16 ground object categories such as corn, soybeans, cucumbers, tomatoes, etc.

[0089] 2. Simulation content

[0090] Simulation 1. Under the above simulation conditions, the present invention and the existing seven methods SVM, 2D-CNN, 3D-CNN, SSRN, 3DOCM-SSAN, HyperX, and HSI-SSFTT are respectively used to perform classification on the Indian Pines dataset, and their respective overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 1.

[0091] Table 1 Classification performance of the present invention and seven existing methods on the Indian Pines dataset

[0092]

[0093] Simulation 2. Under the above simulation conditions, the present invention and the existing seven methods SVM, 2D-CNN, 3D-CNN, SSRN, 3DOCM-SSAN, HyperX, and HSI-SSFTT are used to perform classification on the paviaU dataset, and their respective overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 2.

[0094] Table 2 Classification performance of the present invention and the existing seven methods on the paviaU dataset

[0095]

[0096] Simulation 3. Under the above simulation conditions, the present invention and the existing seven methods SVM, 2D-CNN, 3D-CNN, SSRN, 3DOCM-SSAN, HyperX, and HSI-SSFTT are used to classify the Houston 2013 data set, and their overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 3.

[0097] Table 3 Classification performance of the present invention and the seven existing methods on the Houston 2013 dataset

[0098]

[0099] Simulation 4. Under the above simulation conditions, the present invention and the existing seven methods SVM, 2D-CNN, 3D-CNN, SSRN, 3DOCM-SSAN, HyperX, and HSI-SSFTT are respectively used to perform classification on the Salinas data set, and their respective overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 4.

[0100] Table 4 Classification performance of the present invention and the existing seven methods on the Salinas dataset

[0101]

[0102] The sources of the seven existing methods in the above table are as follows:

[0103] SVM is a method for hyperspectral data classification published in IEEE, namely: JAGualtieri andS.Chettri, "Support vector machines for classification ofhyperspectraldata," in Proc.IEEE Int.Geosci.Remote Sens.Symp.(IGARSS),vol.2,Jul.2000,pp.813–815.

[0104] 2D-CNN is a method for hyperspectral data classification published in IEEE, namely: G.Cheng, C.Yang, X.Yao, L.Guo, and J.Han, "When deep learning meets metric learning: Remote sensing image scene classifification via learning discriminative CNNs," IEEE Trans. Geosci. Remote Sens. vol. 56, no. 5, pp. 2811–2821, May 2018.

[0105] 3D-CNN is a method for hyperspectral data classification published in IEEE, namely: Y.Xu, L.Zhang, B.Du, and F.Zhang, "Spectral–spatial unifified networks for hyperspectral image classification," IEEE Trans.Geosci.Remote Sens., vol.56, no.10, pp.5893–5909, Oct.2018.

[0106] SSRN is a method for hyperspectral data classification published in IEEE, namely: Z. Zhong, J. Li, Z. Luo, and M. Chapman, "Spectral–spatial residual network for hyperspectral image classification: A3-D deep learning framework," IEEE Trans. Geosci. Remote Sens., vol. 56, no. 2, pp. 847–858, Feb. 2018.

[0107] 3DOCM-SSAN is a method for hyperspectral data classification published in IEEE, namely: Tang X, Meng F, Zhang X, et al. Hyperspectral image classification based on 3-D octave convolution with spatial–spectral attention network[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 59(3): 2430-2447.

[0108] HyperX is a method for hyperspectral data classification published in IEEE, namely: Yang X, Cao W, Lu Y, et al. Hyperspectral Image Transformer Classification Networks[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60:1-15.

[0109] HSI-SSFTT is a method for hyperspectral data classification published in IEEE, namely: Sun L, Zhao G, Zheng Y, et al. Spectral–Spatial Feature Tokenization Transformer for Hyperspectral Image Classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 1-14.

[0110] It is obvious from Tables 1 to 4 that the classification accuracy of the present invention is higher and the generalization ability is stronger in four commonly used hyperspectral remote sensing data sets compared with the seven existing methods, which further illustrates the superiority of the network model and method proposed in the present invention.

[0111] It can be seen from Table 1 to Table 4 that the overall error of the present invention is smaller and the classification accuracy is higher, indicating that the present invention has better performance than the existing method.

Claims

1. A hyperspectral object classification method based on a large kernel attention mechanism and MLP hybrid, characterized in that: The steps include: (1) Construct training sample set and test sample set: 1a) Obtain the original data set and data annotation set of hyperspectral images from the public website: 1b) Select proportional values ​​from each non-zero category in the data annotation set, save the corresponding position coordinates of these values ​​in the data annotation set, find the pixel points at the corresponding coordinate positions in the original data set, and perform mirror segmentation with each pixel point as the center and the set imgsize parameter as the diameter to generate a training set; 1c) Generate a test set using the method in 1b) for all remaining samples; (2) Build a feature extraction network: 2a) Establish a large-core attention mechanism module based on the DenseNet structure, which includes three DenseBlock submodules, to effectively extract information at different spatial positions and different spectral bands of the input image; 2b) establishing a feature enhancement module consisting of a spatial mixing MLP layer and a channel mixing MLP layer to realize information interaction between the spectral dimension and the spatial dimension; 2c) Connect the large core attention mechanism module based on the DenseNet structure and the feature enhancement module in series to form the entire feature extraction network, and use the multi-classification focus loss function as the loss function of the extraction network. FL ; (3) Train the feature extraction network: 3a) Input the training set into the feature extraction network and calculate its loss FL value; 3b) Using the gradient descent method, gradually reduce the value of the loss function to update the network parameters until the set maximum number of iterations is completed to obtain a trained feature extraction network; (4) Input the test set into the trained feature extraction network to obtain the output vector. Use the softmax function and argmax function in the output vector to obtain the position index of the maximum value. The position index is the final classification result of each test sample.

2. The method according to claim 1, characterized in that The imgsize parameter set in step 1b) refers to the pre-set length and width of the training set and test set.

3. The method according to claim 1, characterized in that In step 2a), a large core attention mechanism module based on the DenseNet structure including three DenseBlock submodules is established, and the implementation is as follows: 2a1) Construct a DenseBlock submodule consisting of a transition layer, two convolutional integration layers, two attention layers, and two concatenation layer submodules. Its transmission relationship is divided into four paths: The first path is: transition layer → first convolutional integration layer → first attention layer → first splicing layer → second convolutional integration layer → second attention layer → second splicing layer; The second path is: transition layer → first concatenation layer → second convolutional integration layer → second attention layer → second concatenation layer; The third route is: transition layer → second splicing layer; The fourth route is: first splicing layer → second splicing layer; The two convolution integration layers have the same internal structure, and each convolution integration layer is composed of a convolution layer with a convolution kernel size of 1×1, a convolution layer with a convolution kernel size of 3×3, and a batch processing layer cascaded; The transition layer is composed of a convolution layer with a convolution kernel size of 1×1 and a pooling layer with a size of 2 in cascade; The concatenation layer consists of a concatenation-addition operation; The attention layer consists of a spatial long-range convolution block, a spatial local convolution block, a channel convolution block string, a concatenation and an addition operation; The convolution kernel size of this spatial local convolution block is (2d-1)×(2d-1). The convolution kernel size of the spatial long-range convolution block is d is the expansion rate, K represents the length and width of the input feature; The convolution kernel size of this channel convolution block is 1×1; 2a2) Cascade three DenseBlock sub-modules with the same structure in sequence to form a large core attention mechanism module.

4. The method according to claim 1, characterized in that: The spatial mixing MLP layer and the channel mixing MLP layer constituting the feature enhancement module in step 2b) have the following structure: The spatial hybrid MLP layer contains a spatial transposition operation and the first fully connected layer The channel mixing MLP layer includes a channel transposition operation and a second fully connected layer The connection relationship between the two is: spatial transposition operation → first fully connected layer → channel transposition operation → second fully connected layer; The dimension of each fully connected layer is 128.

5. The method according to claim 1, characterized in that: The extraction network loss function loss set in step 2c) FL , which is expressed as follows: Among them, N is the number of samples, K is the total number of categories, and y i represents the true label category value of sample i, γ is a preset value, and P ik The probability value of predicting the Kth category for the i-th sample.

6. The method according to claim 1, characterized in that Step 3b) Use the gradient descent method to gradually reduce the value of the loss function to update the network parameters, which is implemented as follows: 3a1) Assume that the initial training batch m=0, the maximum number of iterations in the training phase T=200, and α is the learning rate value in the training phase; 3a2) Calculate the feature extraction network parameters θ after the current training update m+1 : Among them, θ m is the network parameter before the current training update, and L(·) is the derivative of the parameters in the current network; 3a3) Repeat 3a2) until the maximum number of iterations T is reached, completing the training of the feature extraction network.

7. The method according to claim 1, characterized in that In step (4), the softmax function and argmax function are used in the output vector to obtain the position index of the maximum value. The formula is as follows: F out =argmax(softmax(F)) Among them, F out The final result, F∈1×K is the output feature vector, and K is the total number of categories.

Citation Information

Patent Citations

  • Vehicle classification method based on multi-branch local attention network

    CN113610144A

  • Remote sensing image cloud detection method and device, computer device and storage medium

    CN115359370A