Three-stage morphological coordinate attention network model and method for image classification
By introducing a three-stage morphological coordinate attention network model in hyperspectral image classification, and using the morphological coordinate attention module to enhance feature representation, the problem of insufficient feature extraction and fusion in the existing technology is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202510286809.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing hyperspectral image classification methods are difficult to effectively capture long-distance and global features, and the feature extraction and fusion are insufficient, resulting in insufficient classification accuracy.
A three-stage morphological coordinate attention network model is proposed, and the feature representation is enhanced by three morphological coordinate attention modules (MCAC, MCAWH and MCAWHC). Combined with 3D convolution and transformer encoder, the full fusion of spectral and spatial features is achieved.
By introducing a morphological coordinate attention module, the network can learn object features more effectively, improve the classification accuracy of hyperspectral images, and show superiority on multiple public data sets.
Smart Images

Figure CN120219829A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of image processing, specifically a three-stage morphological coordinate attention network model and method for image classification. Background Art
[0002] In recent years, remote sensing technology has been widely applied in many application fields, such as urban land change monitoring, crop analysis, etc. Since hyperspectral images (HSIs) usually carry rich information about objects in their spectra, HSI classification aiming to determine the pixel classes in HSIs has been widely explored in many studies, including those based on convolutional neural networks (CNNs) and transformers.
[0003] Although current hyperspectral image classification methods have achieved improved results, few methods focus on jointly spectral-spatial attention to extract discriminative features of hyperspectral images. Currently, there are techniques that use three-dimensional convolution during preprocessing to extract discriminative features from hyperspectral images, and there are related techniques that use transformers to integrate features from three branches of height, width, and depth. However, due to the locality of the receptive field of convolution, these methods have limitations in capturing long-distance and global features, and these networks tend to perform spectral and spatial attention extraction independently, with insufficient feature extraction and fusion, making it difficult to accurately classify hyperspectral images. Summary of the Invention
[0004] To solve the deficiencies of current technologies, the present invention combines existing technologies and starts from practical applications to provide a three-stage morphological coordinate attention network model and method for image classification. By proposing three morphological coordinate attention modules (MCA), a three-stage morphological coordinate attention network is formed to accurately classify hyperspectral images.
[0005] The technical solution of the present invention is as follows:
[0006] According to one aspect of the present invention, a three-stage morphological coordinate attention network model for image classification is provided, including: modifying the coordinate attention mechanism CA using morphological convolution to learn the structural information of objects in a non-linear manner to obtain a position-aware channel attention module MCA C , performing channel-aware spatial attention operations in the height and width dimensions to obtain a plane-based position-aware channel attention module MCA WH , combining MCA WH and MCA C into a three-dimensional-based position-aware channel attention module MCA WHC , using MCA C , MCA WH , MCAWHC Construct a three-stage morphological coordinate attention network for hyperspectral image classification.
[0007] Furthermore, in the network model, first start from the principal component analysis module PCA, and extract patches from the image after PCA dimensionality reduction, including three stages;
[0008] Stage 1: Input the patch into the three-dimensional encoding module of the hyperspectral image to encode the data, and the three-dimensional encoding module simultaneously extracts the comprehensive features of the spectrum and space;
[0009] Stage 2: Input the output of Stage 1 into the spectral and spatial branches respectively to extract spectral features and spatial features;
[0010] Stage 3: Input the features of the two branches in Stage 2 into the feature fusion module S 2 F 2 and input the fused features into the fully connected layer FC to obtain the prediction map.
[0011] Furthermore, the spectral branch includes four convolutional-based position-aware channel attention modules C-MCA C and four transformer-based position-aware channel attention modules E-MCA C , C-MCA C is connected to E-MCA C in a skip connection manner;
[0012] The spatial branch includes four convolutional-based planar position-aware channel attention modules C-MCA WH and four transformer-based planar position-aware channel attention modules E-MCA WH , C-MCA WH is connected to E-MCA WH in a skip connection manner.
[0013] Furthermore, the three-dimensional encoding module includes 3D convolution and a three-dimensional position-aware channel attention module MCA WHC ;
[0014] The C-MCA C includes 2D convolution and a position-aware channel attention module MCA C , and MCA C is added to the transformer encoder module to form the E-MCA C ;
[0015] The C-MCA WH includes 2D convolution and a planar position-aware channel attention module MCAWH Add MCA to the transformer encoder module WH to form the E-MCA WH .
[0016] Furthermore, the position-aware channel attention module MCA C has the following specific structure:
[0017] 1) Perform morphological convolution operations along the width and height respectively, and then perform average pooling respectively to enhance the attention to the position-aware channels;
[0018] 2) Concatenate the feature maps obtained from the above steps, and then use a two-dimensional convolution to reduce the number of channels from C to C / r;
[0019] 3) Use another two-dimensional convolution to expand the number of channels from C / r to C;
[0020] 4) Input the joint feature map into the two-dimensional morphological convolution MC2d to enhance the spatial expression of the extracted feature map for the target;
[0021] 5) After dimensional transformation, obtain the attention map.
[0022] Furthermore, the plane-based position-aware channel attention module MCA WH has the following specific structure:
[0023] 1) Concatenate the features obtained from the channel-based morphological convolution operation MC C with the features obtained from the width-based morphological convolution operation MC W and the height-based morphological convolution operation MC H while keeping the height and width unchanged respectively;
[0024] 2) Through two two-dimensional convolutions, reduce the height from H to H / r and the width from W to W / r respectively;
[0025] 3) Through another two two-dimensional convolutions, increase the height from H / r to H and the width from W / r to W respectively;
[0026] 4) Before feature segmentation, embed two two-dimensional morphological convolutions MC2d to enhance the channel-aware spatial attention features.
[0027] Furthermore, the three-dimensional-based position-aware channel attention module MCA WHC has the following specific structure:
[0028] 1) One-dimensional convolutions MC W , MC H and MC CFirst, perform convolution on the input features along the width, height, and channel dimensions respectively;
[0029] 2) Keep the depth, width, and height unchanged respectively, and connect three combinations of every two result features separately;
[0030] 3) Enhance the three groups of combined features through compression and expansion of the channels, width, and height respectively; use three two-dimensional morphological convolutions MC2d to enhance three types of attention.
[0031] According to another aspect of the present invention, there is provided an image classification method based on the above model, including the following steps:
[0032] S1. Divide the data set into a training set and a test set, and preprocess these two data sets respectively;
[0033] S2. Set calculation parameters related to the number of epochs and the initial learning rate;
[0034] S3. Send the processed training set into the principal component analysis module PCA to reduce the dimension of the data;
[0035] S4. Send the data after dimension reduction into the three-dimensional encoding module in stage 1 to encode the data;
[0036] S5. Send the output of stage 1 into the spectral and spatial branches in stage 2 in parallel to extract spectral and spatial features;
[0037] S6. Send the spectral and spatial features into the spectral-spatial-feature fusion module S 2 F 2 in stage 3 to obtain the fused features;
[0038] S7. Send the fused features into the fully connected layer FC to predict the final pixel labels and obtain the prediction map;
[0039] S8. Send the prediction map and the label map into the loss function for gradient update;
[0040] S9. After each epoch ends, send the test set into the network and successively go through steps S3, S4, S5, S6, and S7 to obtain the prediction map.
[0041] Furthermore, use the mean absolute error as an index. If the index is less than the index of the previous epoch, save the model parameters.
[0042] Advantages of the present invention:
[0043] The present invention proposes three attention modules (MCA C , MCA WHand MCA WHC ) to enhance the feature representation of objects in HIS. A three-stage morphological coordinate attention network is introduced for hyperspectral image classification by using the proposed attention module. Due to the non-linearity of morphological convolution in the morphological coordinate attention module MCA and the correlation between the height, width, and channel dimensions of MCA, the three-stage morphological coordinate attention network helps to learn object features. Through a large number of experiments on multiple public datasets, the superiority of the network proposed in the present invention is verified from both qualitative and quantitative aspects, proving that the present invention can accurately classify hyperspectral images. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is the structural diagram of the network model of the present invention.
[0045] Figure 2 is the position-aware channel attention module MCA of the present invention C structural diagram.
[0046] Figure 3 is the position-aware channel attention module MCA based on the plane WH structural diagram.
[0047] Figure 4 is the position-aware channel attention module MCA based on three dimensions WHC structural diagram.
[0048] Figure 5 is the label map of the input image in the embodiment.
[0049] Figure 6 is Figure 5 corresponding prediction map. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In combination with the accompanying drawings and specific embodiments, the present invention will be further described. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0051] Embodiment 1
[0052] This embodiment provides a three-stage morphological coordinate attention network model for image classification. This network model mainly extracts spatial spectral features by fusing features in three dimensions.
[0053] First, the morphological convolution is used to modify the Coordinate Attention (CA) mechanism to learn the structural information of the object in a non-linear manner, resulting in an inherent position-aware channel attention module, denoted as MCA. C Then, considering that the spectral and spatial features are separable, this embodiment proposes a plane-based position-aware channel attention module MCA WH that performs channel-aware spatial attention operations in the height and width dimensions. Next, MCA WH and MCA C are combined into a three-dimensional-based position-aware channel attention module MCA WHC to consider the attention to the three dimensions of the initial encoded features. Finally, this embodiment uses the above three attention modules to construct a three-stage morphological coordinate attention network for hyperspectral image classification.
[0054] Refer to Figure 1 as shown, which is the structural diagram of the three-stage morphological coordinate attention network model.
[0055] As Figure 1 shown, it starts with the Principal Component Analysis (PCA) module, mainly including three stages: for the image after PCA dimensionality reduction, patches are extracted. Stage 1: The patches are input into the three-dimensional encoding (3D encoding) of the hyperspectral image to encode the data. The 3D encoding simultaneously extracts the comprehensive spectral and spatial features, enhancing the context representation ability of the features. The 3D encoding module is composed of a series connection of 3D convolution and the proposed MCA WHC . Stage 2: The output of Stage 1 is respectively input into the spectral and spatial branches. Stage 2-1: The spectral branch extracts spectral features, which are connected in a skip connection manner by four convolution-based position-aware channel attention modules C-MCA C and four transformer-based position-aware channel attention modules E-MCA C . The first four and the last four modules effectively fuse cross-layer features through skip connections from the first module to the third module and from the second module to the fourth module respectively. Each C-MCA C is composed of 2D convolution and MCA C . In this embodiment, MCA C is added to the traditional transformer encoder to form E-MCA C . Stage 2-2: The spatial branch extracts spatial features, which are composed of four convolution-based plane position-aware channel attention modules C-MCA WH and four transformer-based plane position-aware channel attention modules E-MCA WHConnected in a skip connection manner. Each C-MCA WH Consists of 2D convolution and MCA WH In this embodiment, MCA is added to the traditional transformer encoder WH To form E-MCA WH Stage 3: Input the features of the two branches in Stage 2 into the feature fusion module S 2 F 2 Among them, (d) in Figure 1 Is the schematic diagram of S 2 F 2 Finally, the fused features are input into the fully connected layer FC to obtain the prediction map.
[0056] The morphological convolution MC consists of dilation and erosion operations. MC is a generalization of ordinary convolution. However, in this embodiment, it is regarded as a non-linear learnable pooling, and they are used to construct our morphological coordinate attention module MCA.
[0057] For the position-aware channel attention module MCA C Of this embodiment, its structure is as shown in Figure 2 Shown.
[0058] Considering that hyperspectral images are vulnerable to noise and are inherently non-linear, this embodiment performs the following operations in MCA C Mainly:
[0059] The original coordinate attention mechanism CA is indeed a position-aware channel attention mechanism. The core operations are mainly (i) average pooling along the x-axis and y-axis, (ii) reducing the number of channels from C to C / r using a two-dimensional convolution, and (iii) expanding the number of channels from C / r to C using another two-dimensional convolution. Although the first operation enables the channel attention in (ii) and (iii) to focus on the target position, due to its linearity and fixed parameters, it cannot adapt to the data. Considering that hyperspectral images are vulnerable to noise and are inherently non-linear, this embodiment replaces the x and y Ave Pool in (i) with MC in the width and height dimensions (i.e., MC W And MC H ) to learn robust features. In addition, after (iii), this embodiment also uses channel-based two-dimensional MC (MC2d) to enhance the attention to the position-aware channels. The modified CA is called MCA C . Specifically, the operations inside the position-aware channel attention module MCA C Are as follows:
[0060] (1) Perform morphological convolutional operations (MC) along the width and height respectively, and then perform average pooling respectively to enhance the attention to the position-aware channels. The morphological convolutional operation based on the width and the morphological convolutional operation based on the height are represented by MC W and MC H respectively. (2) Concatenate the feature maps obtained in the above steps, and then use a two-dimensional convolution to reduce the number of channels from C to C / r, where r is a set value, mainly used to reduce the number of channels to reduce the computational load. (3) Use another two-dimensional convolution to expand the number of channels from C / r to C. (4) Input the joint feature map into the two-dimensional morphological convolution MC2d to enhance the spatial expression of the extracted feature map for the target. (5) Finally, through dimensional transformation, an attention map is obtained. It should be noted that MCA C helps to preserve the features of tiny edges because MCA C sequentially includes dilation and erosion operations.
[0061] For the plane-based position-aware channel attention module MCA WH of this embodiment, its structure is as Figure 3 shown.
[0062] The task of hyperspectral image classification aims to predict the label of each pixel in the hyperspectral image, mainly relying on the features extracted from the spectrum. Therefore, it is necessary to find an effective method to utilize the channel depth features to label each pixel. For this purpose, by introducing MC into MCA C , this embodiment proposes a channel-aware MCA WH . Specifically, first concatenate the features obtained by the channel-based morphological convolutional operation MC C with the features obtained by MC W and MC H , and the height and width remain unchanged respectively. Then, through two two-dimensional convolutions, the height is reduced from H to H / r and the width is reduced from W to W / r respectively. Next, through another two two-dimensional convolutions, the height is increased from H / r to H and the width is increased from W / r to W respectively.
[0063] Finally, before feature segmentation, two MC2d are embedded to enhance the channel-aware spatial attention features. It can be seen from the process after "Split" that MCA WH also includes spatial attention independent of channels, that is, it increases the width of height perception and the height of width perception attention. In contrast, MCA WH is position-aware spatial attention, while MCA C is position-aware channel attention.
[0064] For the three-dimensional-based position-aware channel attention module MCA of this embodimentWHC , and its structure is as Figure 4 shown.
[0065] Since 3D convolution has been proven effective in feature extraction for hyperspectral images, this embodiment proposes 3D MCA (i.e., MCA WHC ), which promotes the initial encoding of hyperspectral images in both spectral and spatial dimensions by integrating MCA C and MCA WHC . One-dimensional convolutional MC W , MC H and MC C first perform convolutions on the input features along the width, height, and channel dimensions respectively. Then, by keeping the depth, width, and height unchanged respectively, all three combinations of every two resulting features are concatenated separately. In addition, these three concatenated features are enhanced by compression and expansion along the channel, width, and height respectively. Therefore, such enhancements are width-height plane-aware channel attention, height-channel plane-aware width attention, and channel-width plane-aware height attention respectively. Finally, three MC2ds are used to enhance these three types of attention.
[0066] The position-aware channel attention module MCA C aims to make full use of the rich spectral information in the HSI. The plane-based position-aware channel attention module MCA WH aims to use channel features for pixel labeling, which is an inherent requirement of HSIC. Therefore, MCA C and MCA WH complement each other in feature enhancement for HSIC. The 3D-based position-aware channel attention module MCA WHC is responsible for the initial feature encoding of the HSI in both spectral and spatial dimensions.
[0067] Current hyperspectral image classification networks tend to perform spectral and spatial attention extraction independently, and the feature extraction fusion is not sufficient. This embodiment proposes three morphological coordinate attention modules (MCA) to form a three-stage morphological coordinate attention network for accurately classifying hyperspectral images. As shown in Table 1 below, the evaluation metrics of several currently popular methods on three datasets, namely IP (Indian Pine), PU (Pavia University Scene), and HT (Houston2013), are presented.
[0068] Table 1 Comparison Table of Evaluation Metrics
[0069]
[0070] As can be seen from the table, the network model of this embodiment has almost obtained the highest values of all evaluation indicators, indicating its effectiveness and accuracy in capturing spectral-spatial features.
[0071] Embodiment 2
[0072] This embodiment provides an image classification method based on a three-stage morphological coordinate attention network model. The specific steps are as follows:
[0073] 1. First, divide the dataset into a training set and a test set, and then preprocess these two datasets respectively, including normalization and flipping.
[0074] 2. Set the number of epochs to 200 and the initial learning rate to 1e -4 . The optimizer is Adam, and the learning rate decreases by 10 times every 50 epochs.
[0075] 3. Feed the processed training set into the principal component analysis module PCA to reduce the dimensionality of the data.
[0076] 4. Feed the data after dimensionality reduction into the three-dimensional encoding module (stage 1) to encode the data. The three-dimensional encoding module is composed of 3D convolution and the proposed MCA WHC connected in series. The three-dimensional encoding simultaneously extracts the comprehensive features of the spectrum and space, enhancing the context representation ability of the features.
[0077] 5. Feed the output of stage 1 into the spectral and spatial branches in parallel (stage 2) to extract spectral and spatial features. The spectral branch is composed of four C-MCA C and four E-MCA C connected in a skip connection manner. Each C-MCA C is composed of 2D convolution and MCA C , and MCA C is added to the traditional transformer encoder to form E-MCA C . The spatial branch is composed of four C-MCA WH and four E-MCA WH connected in a skip connection manner. Each C-MCA WH is composed of 2D convolution and MCA WH , and MCA WH is added to the traditional transformer encoder to form E-MCA WH .
[0078] 6. Feed the spectral and spatial features into the spectral-spatial-feature fusion module S 2 F 2 (stage 3) to obtain the fused features.
[0079] 7. Feed the fused features into the fully connected layer FC for predicting the final pixel labels to obtain a prediction map.
[0080] 8. Feed the prediction map and the label map into the loss function for gradient update.
[0081] 9. After each epoch ends, the test set is fed into the network and goes through steps 3, 4, 5, 6, and 7 in sequence to obtain a prediction map. The mean absolute error is used as an indicator (the smaller the better). If the indicator is smaller than that of the previous epoch, the model parameters are saved.
[0082] Reference Figure 5 As shown, it is the label of the input image, Figure 6 and it is the prediction map. Through visual verification, it can be observed that for the image classification method proposed in this embodiment, the output result contains less noise and can achieve accurate classification of images.
Claims
1. A three-stage morphological coordinate attention network model for image classification, characterized by: include: Use morphological convolution to modify the coordinate attention mechanism CA, learn the structural information of the object in a nonlinear way, and obtain the position-aware channel attention module MCA C , performing channel-aware spatial attention operations on the height and width dimensions to obtain the plane-based position-aware channel attention module MCA WH , MCA WH and MCA C Merged into a three-dimensional position-aware channel attention module MCA WHC , using MCA C 、MCA WH 、MCA WHC Constructing a three-stage morphological coordinate attention network for hyperspectral image classification.
2. The three-stage morphological coordinate attention network model for image classification according to claim 1, characterized in that: In the network model, we first start with the principal component analysis module PCA to extract patches from the image after PCA dimension reduction, which includes three stages; Phase 1: Input the patch into the 3D encoding module of the hyperspectral image to encode the data. The 3D encoding module extracts the comprehensive spectral and spatial features at the same time. Stage 2: Input the output of stage 1 into the spectral and spatial branches to extract spectral features and spatial features respectively; Stage 3: Input the two branch features of stage 2 into the feature fusion module S 2 F 2 In , the fused features are input into the fully connected layer FC to obtain the prediction graph.
3. The three-stage morphological coordinate attention network model for image classification according to claim 2, characterized in that: The spectral branch consists of four convolutional-based position-aware channel attention modules C-MCA C and four transformer-based position-aware channel attention modules E-MCA C , C-MCA C With E-MCA C Use jump connection method to connect; The spatial branch consists of four convolution-based planar position-aware channel attention modules C-MCA. WH and four transformer-based planar position-aware channel attention modules E-MCA WH , C-MCA WH With E-MCA WH Use jump connection method.
4. The three-stage morphological coordinate attention network model for image classification according to claim 3, characterized in that: The three-dimensional encoding module includes a 3D convolution and a three-dimensional position-aware channel attention module MCA. WHC ; The C-MCA C Includes 2D convolution and position-aware channel attention module MCA C , add MCA to the transformer encoder module C Composition of the E-MCA C ; The C-MCA WH Includes 2D convolution and plane-based position-aware channel attention module MCA WH , add MCA to the transformer encoder module WH Composition of the E-MCA WH .
5. The three-stage morphological coordinate attention network model for image classification according to claim 1, characterized in that: The position-aware channel attention module MCA C The specific structure is as follows: (1) Morphological convolution operations are performed along the width and height respectively, followed by average pooling to enhance the focus on position-aware channels; (2) Concatenate the feature maps obtained in the above steps, and then use two-dimensional convolution to reduce the number of channels from C to C / r; (3) Use another 2D convolution to expand the number of channels from C / r to C; (4) Input the joint feature map into the two-dimensional morphological convolution MC2d to enhance the spatial expression of the target by the extracted feature map; (5) After dimension transformation, the attention map is obtained.
6. The three-stage morphological coordinate attention network model for image classification according to claim 5, characterized in that: Plane-based Position-aware Channel Attention Module MCA WH The specific structure is as follows: (1) The channel-based morphological convolution operation MC C The obtained features are combined with the width-based morphological convolution operation MC W and height-based morphological convolution operation MC H The obtained features are concatenated, and the height and width remain unchanged; (2) Through two 2D convolutions, the height is reduced from H to H / r and the width is reduced from W to W / r; (3) Through two more 2D convolutions, the height is increased from H / r to H and the width is increased from W / r to W; (4) Before feature segmentation, two two-dimensional morphological convolutions MC2d are embedded to enhance the channel-aware spatial attention features.
7. The three-stage morphological coordinate attention network model for image classification according to claim 6, characterized in that: Three-dimensional position-aware channel attention module MCA WHC The specific structure is as follows: (1) One-dimensional convolution MC W , MC H and MC C First, the input features are convolved along the width, height, and channel dimensions respectively; (2) Keeping the depth, width and height unchanged, concatenate the three combinations of each two resulting features separately; (3) The features of the three groups are enhanced by compressing and expanding the channel, width and height respectively; (4) Use three two-dimensional morphological convolutions MC2d to enhance the three types of attention.
8. An image classification method based on the model according to any one of claims 1 to 7, characterized in that: The steps include: S1. Divide the data set into a training set and a test set, and preprocess the two data sets respectively; S2, set epoch, initial learning rate and related calculation parameters; S3, sending the processed training set to the principal component analysis module PCA to reduce the dimension of the data; S4, sending the reduced-dimensional data to the three-dimensional encoding module of stage 1 to encode the data; S5, sending the output of stage 1 to the spectral and spatial branches of stage 2 in parallel, respectively, to extract spectral and spatial features; S6, send the spectrum and spatial features to the spectrum-space-feature fusion module S in stage 3 2 F 2 In the above example, we obtain the fusion features. S7, sending the fused features to the fully connected layer FC to predict the final pixel label and obtain the predicted image; S8, send the prediction graph and label graph to the loss function for gradient update; S9. After each epoch, the test set is sent to the network, and then goes through steps S3, S4, S5, S6, and S7 in sequence to obtain the prediction graph.
9. The classification method according to claim 8, characterized in that: The mean absolute error is used as the indicator. If the indicator is less than the indicator of the previous epoch, the model parameters are saved.
Citation Information
Cited By
Agricultural greenhouse extraction method based on improved U-Net model
CN120495906A