Hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features

By introducing the alignment technology of spatial topology and inter-class features into the hyperspectral remote sensing image classification network, the shortcomings of the cross-domain hyperspectral remote sensing image classification method in the existing technology in reducing domain offsets are solved, and higher classification accuracy is achieved.

CN119992215AActive Publication Date: 2025-05-13QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Patent Information

Application Number
CN202510178192.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing cross-domain hyperspectral remote sensing image classification methods have limited performance in reducing domain offset, mainly ignoring the influence of spatial topology between the source domain and the target domain on domain offset, resulting in low classification accuracy.

Method used

A hyperspectral remote sensing image classification method based on spatial topology structure and inter-class features is proposed. By constructing a hyperspectral remote sensing image classification network, the network includes a feature extraction network, a domain alignment network and a classifier. The domain alignment network adopts distributed alignment and graph alignment, and enhances feature alignment and spatial topology through knowledge distillation of teacher networks and student networks, as well as the use of graph convolutional layers.

Benefits of technology

By enhancing the alignment of inter-class features and spatial topology, the classification accuracy of hyperspectral remote sensing images is significantly improved and higher classification accuracy is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992215A_ABST
    Figure CN119992215A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral remote sensing image classification method based on a spatial topological structure and inter-class features, and relates to the technical field of hyperspectral remote sensing image classification. According to the method, the feature map output by the feature enhancement network comprises enhanced local space information and global space information, and the hyperspectral remote sensing image classification network model with higher classification precision obtained by training is used for classification prediction, so that the method can obtain a better classification effect, and the classification accuracy of the hyperspectral remote sensing image is improved. Tests show that the hyperspectral remote sensing image classification method provided by the invention has relatively significant classification accuracy for the second type of ground feature types and the fourth type of ground feature types of the HYRANK data set, and particularly has more significant classification accuracy for the fourth type of ground feature types of the HYRANK data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of hyperspectral remote sensing image classification, and in particular to a hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features. Background Art

[0002] Domain shift refers to the difference between the source domain and the target domain. In cross-domain hyperspectral remote sensing image classification, domain shift is a major challenge to classification accuracy. This is mainly because the hyperspectral remote sensing image acquisition process is very sensitive to various factors, such as sensor nonlinearity, seasonal changes, and weather fluctuations. These factors may lead to significant differences in spectral reflectance between the source domain and the target domain for the same land cover category.

[0003] In order to reduce the domain shift between the source domain and the target domain, and thus improve the classification accuracy on the target domain, researchers have begun to introduce deep learning technology into cross-domain hyperspectral remote sensing image classification to solve the problem of domain shift. However, the ability of cross-domain hyperspectral remote sensing image classification methods based on deep learning in the prior art to reduce domain shift is limited. This is because these cross-domain hyperspectral remote sensing image classification methods mainly emphasize the overall distribution alignment between the source domain hyperspectral remote sensing images and the target domain hyperspectral remote sensing images, but ignore the spatial topological structure and inter-class features of the source domain hyperspectral remote sensing images and the target domain hyperspectral remote sensing images. The influence of the domain shift makes the classification accuracy of this type of cross-domain hyperspectral remote sensing image classification method low. To this end, this application proposes a hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features. Summary of the invention

[0004] In order to make up for the deficiencies of the prior art, the present invention provides a hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features.

[0005] The technical solution of the present invention is: The present application provides a hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features, comprising the following steps: A hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features specifically includes the following steps: S1: Obtain source domain sub-image pairs and target domain sub-image pairs for training; S2: Construct a hyperspectral remote sensing image classification network. The hyperspectral remote sensing image classification network includes a feature extraction network, a domain alignment network and a classifier connected in sequence. The feature extraction network is used to extract global features and local features, the domain alignment network is used to achieve distribution alignment and graph alignment, and the classifier is used to predict classification results. The domain alignment network includes a feature enhancement network and a distribution alignment module and a graph alignment module respectively connected to the output end of the feature enhancement network. The distribution alignment module includes a teacher network and a student network. The graph alignment module includes branches I and II. The input end of the teacher network and the input end of branch I are both connected to the output end of the feature enhancement network. The source domain output of the feature enhancement network The enhanced feature map is used as the input of the teacher network and branch I; the input of the student network and the input of branch II are connected to the output of the feature extraction network, and the target domain enhanced feature map output by the feature enhancement network is used as the input of the student network and branch II; the output of the teacher network and the output of branch I are connected to the input of the classifier; the knowledge learned by the teacher network is transferred to the student network by knowledge distillation, so that the feature map obtained by the student network can be aligned with the feature map obtained by the teacher network, and the mutual alignment is distribution alignment; through optimal transportation, the feature map obtained by branch II is aligned with the feature map obtained by branch I, and the mutual alignment is graph alignment; S3: Based on the total loss of the hyperspectral remote sensing image classification network, the hyperspectral remote sensing image classification network is trained using the source domain sub-image pairs and the target domain sub-image pairs to obtain the hyperspectral remote sensing image classification network model; S4. Fill pixels of the same size on each boundary of the hyperspectral remote sensing image to be classified and predicted, and the size of the filled pixels is 1 / 2 patch size plus 1, and then map to obtain image X, divide image X into two sub-images to be predicted of the same size, and the two sub-images to be predicted constitute a sub-image pair to be predicted, and then input the sub-image pair to be predicted into the hyperspectral remote sensing image classification network model, and process them through the feature extraction network and the feature enhancement network in turn to obtain an enhanced feature map, and then input the enhanced feature map into the classifier for classification prediction, and the image classification category prediction result and classification accuracy can be obtained.

[0006] Preferably, step S1 specifically includes the following steps: obtaining a hyperspectral remote sensing image dataset, dividing a source hyperspectral remote sensing image and a target hyperspectral remote sensing image according to the scenes in the hyperspectral remote sensing image dataset; and preprocessing the source hyperspectral remote sensing image and the target hyperspectral remote sensing image to obtain image S, image T and image U, and then obtaining a source domain sub-image pair and a target domain sub-image pair based on image S and image T, respectively.

[0007] Preferably, in step S1, the source hyperspectral remote sensing image and the target hyperspectral remote sensing image are preprocessed to obtain image S, image T and image U, and then based on image S and image T, a source domain sub-image pair and a target domain sub-image pair are obtained respectively, specifically comprising the following steps: S1-1: Fill pixels of the same size on each boundary of the source hyperspectral remote sensing image and the target hyperspectral remote sensing image. The pixel size filled is 1 / 2 of the patch size plus 1, and the primary processed source hyperspectral remote sensing image and the primary processed target hyperspectral remote sensing image are obtained. S1-2: 80% of the first-processed source hyperspectral remote sensing images with corresponding labels are randomly selected and then used together with all the first-processed target hyperspectral remote sensing images without corresponding labels as the training set, and all the first-processed target hyperspectral remote sensing images with corresponding labels are used as the test set; Then, 5% of the primary processing source hyperspectral remote sensing images are randomly selected from the training set for flipping and radiation noise processing to perform data enhancement, and the secondary processing source hyperspectral remote sensing images are obtained. The secondary processing source hyperspectral remote sensing images all have corresponding labels, and the secondary processing source hyperspectral remote sensing images are also added to the training set; among them, flipping includes left-right flipping and up-down flipping, and the probabilities of left-right flipping and up-down flipping are the same; radiation noise includes Gaussian noise, Poisson noise and salt and pepper noise, and the probabilities of radiation Gaussian noise, Poisson noise and salt and pepper noise are the same; S1-3: All images in the training set and the test set are input into the feature space for mapping. The primary processed source hyperspectral remote sensing images with corresponding labels and the secondary processed source hyperspectral remote sensing images with corresponding labels in the training set are mapped to obtain image S. The primary processed target hyperspectral remote sensing images without corresponding labels in the training set are mapped to obtain image T. The primary processed target hyperspectral remote sensing images with corresponding labels in the test set are mapped to obtain image U. S1-4: Each image S is divided into two source domain sub-images of the same dimensional size. The two source domain sub-images divided based on the same image S are a pair of source domain sub-images, called a source domain sub-image pair; each image T is divided into two target domain sub-images of the same dimensional size. The two target domain sub-images divided based on the same image T are a pair of target domain sub-images, called a target domain sub-image pair.

[0008] Preferably, in step S2, the feature extraction network includes a Mamba branch and a convolution branch; the input end of the Mamba branch is connected to the input end of the convolution branch, the output end of the Mamba branch and the output end of the convolution branch are both connected to the Concat layer, the Concat layer is sequentially connected to the first 2D convolution layer and the first Add layer, and the input end of the first Add layer is also connected to the input end of the convolution branch and the input end of the Mamba branch.

[0009] Preferably, in step S2, the Mamba branch includes a first normalization layer, a first linear layer, a depth-wise separable convolution layer, a first SiLU layer, a 2D selective scanning module, a second normalization layer, and an element-by-element multiplication module connected in sequence; the output end of the first normalization layer is also connected in sequence to the second linear layer and the second SiLU layer, the output end of the second SiLU layer is connected to the input end of the element-by-element multiplication module, and the element-by-element multiplication unit is also connected in sequence to the third linear layer and the second Add layer; the input end of the first normalization layer is also connected to the input end of the second Add layer.

[0010] Preferably, in step S2, the convolution branch includes a second 2D convolution layer, a first RELU layer, a third 2D convolution layer, a second RELU layer and a fourth 2D convolution layer connected in sequence; the input end of the second 2D convolution layer in the convolution branch is connected to the input end of the first normalization layer in the Mamba branch, and the output end of the fourth 2D convolution layer and the output end of the second Add layer are both connected to the Concat layer; the input end of the second 2D convolution layer is also connected to the input end of the first Add layer.

[0011] Preferably, in step S2, the teacher network includes four channel conversion units connected in sequence, each of which includes a convolutional layer, a batch normalization layer, a ReLU layer and a maximum pooling layer connected in sequence, wherein the convolution kernel size of the convolution layer in the first channel conversion unit is 7×7, the convolution kernel size of the convolution layer in the second channel conversion unit, the convolution kernel size of the convolution layer in the third channel conversion unit, and the convolution kernel size of the convolution layer in the fourth channel conversion unit are all 3×3.

[0012] Preferably, in step S2, the student network includes four dimensional conversion units connected in sequence, and the four dimensional conversion units each include a fully connected layer and a ReLU layer connected in sequence.

[0013] Preferably, in step S2, branch I and branch II have the same structure and function; branch I includes two graph convolution layers connected in sequence, wherein the first graph convolution layer is used to capture global spatial topological structure information, and the second graph convolution layer is used to capture more detailed local spatial topological structure information.

[0014] Preferably, in step S2, the feature enhancement network includes a fifth 2D convolutional layer, a GELU layer, a sixth 2D convolutional layer, and a channel segmentation layer connected in sequence, the output end of the channel segmentation layer is connected to four depth-separable convolutional layers, the output ends of the four depth-separable convolutional layers are all connected to the Concat layer, the Concat layer is sequentially connected to the seventh 2D convolutional layer and the Add layer, and the input end of the Add layer is also connected to the input end of the fifth 2D convolutional layer; wherein the convolution kernel size of the fifth 2D convolutional layer is 3×3, and the convolution kernel size of the sixth 2D convolutional layer and the seventh 2D convolutional layer is 1×1; in the present application, the feature enhancement network is used to enhance the local spatial information and global spatial information in the feature map output by the feature extraction network.

[0015] Preferably, step S3 specifically includes the following steps: inputting a pair of source domain sub-images and a pair of target domain sub-images into a hyperspectral remote sensing image classification network, calculating the total loss of the hyperspectral remote sensing image classification network, and then optimizing the gradient and back-propagating, updating the model parameters of the hyperspectral remote sensing image classification network, and completing the training of one epoch; repeating the training for 500 epochs to complete the training of one training segment, after one training segment ends, the classifier outputs the classification category prediction result and classification accuracy; repeating the training of the training segment until the classification accuracy output by the next training segment is greater than the classification accuracy of the current training segment, then saving the parameters of the hyperspectral remote sensing image classification network during the last epoch training of the next training segment as the final model parameters, and obtaining the hyperspectral remote sensing image classification network model.

[0016] Preferably, in step S3, a pair of source domain sub-images and a pair of target domain sub-images are input for each epoch of training, and during each epoch of training, the input source domain sub-image pair is a pair of source domain sub-images randomly selected from all source domain sub-image pairs, and the input target domain sub-image is also a pair of target domain sub-images randomly selected from all target domain sub-images.

[0017] Compared with the prior art, the present invention has the following beneficial effects: The feature extraction network of the present application can extract features from the input source domain sub-image pair, target domain sub-image and test sub-image pair respectively, and the source domain feature map M s , target domain feature map M t And the test feature map, source domain feature map M s , target domain feature map M t And the test feature map has rich supplementary information, global spatial information and local spatial information; the feature enhancement network in the domain alignment network has rich supplementary information, global spatial information and local spatial information; and the feature enhancement network in the domain alignment network has rich supplementary information, global spatial information and local spatial information. sAfter processing, the source domain enhanced feature map with enhanced local spatial information and global spatial information can be output, and the feature enhancement network can enhance the target domain feature map M t After processing, it is possible to output a target domain enhanced feature map with enhanced local spatial information and global spatial information. In the process of training a hyperspectral remote sensing image classification network, the present application uses distribution alignment and graph alignment to achieve domain alignment, thereby obtaining a hyperspectral remote sensing image classification network model with higher classification accuracy. In the process of distribution alignment, the present application transfers the knowledge learned by the teacher network in the process of processing the source domain enhanced feature map to the student network through knowledge distillation, so that the feature map obtained by the student network in the process of processing the target domain enhanced feature map can be aligned with the feature map obtained by the teacher network in processing the source domain enhanced feature map (this mutual alignment is distribution alignment), thereby better aligning the inter-class features of the source domain and the target domain, and reducing the impact of the difference in inter-class features on domain offset. The feature map output by the teacher network and the feature map output by the student network have a higher similarity, so as to achieve the purpose of improving the classification accuracy of the hyperspectral remote sensing image classification network; in the process of graph alignment, the present application adopts branch II to process the target domain enhanced feature map, and through optimal transportation, the feature map obtained by branch II in the process of processing the target domain enhanced feature map can be aligned with the feature map obtained by branch I in processing the source domain enhanced feature map (this mutual alignment is graph alignment), so that the spatial topological structure information of the source domain and the target domain can be better aligned, and the influence of the difference between the spatial topological structure information on the domain offset is reduced, so that the feature map output by branch I and the feature map output by branch II Figure 2 There is a higher similarity between them, thereby achieving the purpose of improving the classification accuracy of the hyperspectral remote sensing image classification network. In addition, since the local spatial information and global spatial information in the source domain enhanced feature map and the target domain enhanced feature map are enhanced during the training process of the hyperspectral remote sensing image classification network of the present application, the difference between the features of the classes and the difference between the spatial topological structure information will be smaller during the distribution alignment and graph alignment process, and the difference between the features of the classes and the difference between the spatial topological structure information will have a smaller impact on the domain offset, so that the present application can obtain a hyperspectral remote sensing image classification network model with better classification accuracy.

[0018] When the present application performs classification prediction on the hyperspectral remote sensing images to be classified and predicted, the classifier is used to perform classification prediction on the feature map output by the feature enhancement network and output the classification category prediction result. When the method described in the present application classifies the hyperspectral remote sensing images, since the feature map output by the feature enhancement network contains enhanced local spatial information and global spatial information, and the hyperspectral remote sensing image classification network model with higher classification accuracy trained by the present application is used for classification prediction, the method described in the present application can achieve better classification results. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic diagram of the network structure of the hyperspectral remote sensing image classification network in the present invention; Figure 2 It is a schematic diagram of the network structure of the feature extraction network. Figure 2 In the example, when the input is a source domain sub-image pair, the feature extraction network outputs the source domain feature map M s ; Figure 3 yes Figure 1 Schematic diagram of the network structure of the feature enhancement network; Figure 4 yes Figure 1 Schematic diagram of the network structure of the teacher network in the Chinese textbook; Figure 5 yes Figure 1 Schematic diagram of the network structure of the middle school student network; Figure 6 yes Figure 1 Schematic diagram of the network structure of branch I. DETAILED DESCRIPTION

[0020] In order to help those skilled in the art better understand the technical solution of the present invention, the technical solution in the embodiment will be described in detail below with reference to the accompanying drawings. It should be emphasized that the embodiment is only a part of the present invention, not all embodiments. According to the embodiments of the present invention, other embodiments that can be obtained by those skilled in the art without performing creative work should be included in the protection scope of the present invention.

[0021] This embodiment provides a hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features, which specifically includes the following steps: S1: Obtain source domain sub-image pairs and target domain sub-image pairs for training; specifically comprising the following steps: obtaining a hyperspectral remote sensing image dataset, dividing source hyperspectral remote sensing images and target hyperspectral remote sensing images according to scenes in the hyperspectral remote sensing image dataset; and preprocessing the source hyperspectral remote sensing images and the target hyperspectral remote sensing images to obtain image S, image T and image U, and then obtaining source domain sub-image pairs and target domain sub-image pairs based on image S and image T respectively; In this application, the hyperspectral remote sensing image dataset is the HyRANK dataset. The website for obtaining the HyRANK dataset is: https: / / www.jianguoyun.com / p / DT4z7ugQncXpChj19dAEIAA; There are two scenes involved in the HyRANK dataset, namely Dioni and Loukia. In this application, source hyperspectral remote sensing images and target hyperspectral remote sensing images are divided according to the scenes in the hyperspectral remote sensing image dataset. Specifically, Dioni is identified as the source scene, Loukia is identified as the target scene, hyperspectral remote sensing images with source scenes are identified as source hyperspectral remote sensing images, and hyperspectral remote sensing images with target scenes are identified as target hyperspectral remote sensing images; In this application, the ground object coverage categories of Dioni and Loukia shown in Table 1 and the number of samples in each category are the same as TABLE Ⅱ disclosed in the paper "Topological Structure and Semantic Information Transfer Network for Cross-Scene Hyperspectral Image Classification"; In this application, the source hyperspectral remote sensing image and the target hyperspectral remote sensing image are preprocessed to obtain image S, image T and image U, and then based on image S and image T, the source domain sub-image pair and the target domain sub-image pair are obtained respectively, which specifically includes the following steps: S1-1: Fill pixels of the same size on each boundary of the source hyperspectral remote sensing image and the target hyperspectral remote sensing image, and the filled pixel size is 1 / 2 patch size plus 1, to obtain the initially processed source hyperspectral remote sensing image and the initially processed target hyperspectral remote sensing image; in this embodiment, the patch size is 13, and the filled pixel size is 13 / 2 + 1=7.5; since some source hyperspectral remote sensing images in the HyRANK data set have corresponding labels, and some source hyperspectral remote sensing images do not have corresponding labels; similarly, some target hyperspectral remote sensing images in the HyRANK data set have corresponding labels, and some target hyperspectral remote sensing images do not have corresponding labels; therefore, after step S1-1, the following situation exists: some initially processed source hyperspectral remote sensing images have corresponding labels, some initially processed source hyperspectral remote sensing images do not have corresponding labels, some initially processed target hyperspectral remote sensing images have corresponding labels, and some initially processed target hyperspectral remote sensing images do not have corresponding labels; S1-2: 80% of the first-processed source hyperspectral remote sensing images with corresponding labels are randomly selected, and then together with all the first-processed target hyperspectral remote sensing images without corresponding labels, they form a training set, and all the first-processed target hyperspectral remote sensing images with corresponding labels are used as a test set; Then, 5% of the primary processing source hyperspectral remote sensing images are randomly selected from the training set (the primary processing source hyperspectral remote sensing images in the training set are all primary processing source hyperspectral remote sensing images with corresponding labels) for flipping and radiation noise processing to perform data enhancement, and obtain secondary processing source hyperspectral remote sensing images, the secondary processing source hyperspectral remote sensing images all have corresponding labels, and the secondary processing source hyperspectral remote sensing images are also added to the training set; in this application, flipping includes left-right flipping and up-down flipping, and the probabilities of left-right flipping and up-down flipping are the same, and the radiation noise includes Gaussian noise, Poisson noise and salt and pepper noise, and the probabilities of radiating Gaussian noise, Poisson noise and salt and pepper noise are the same; S1-3: All images in the training set and the test set are input into the feature space for mapping. The primary processed source hyperspectral remote sensing images with corresponding labels and the secondary processed source hyperspectral remote sensing images with corresponding labels in the training set are mapped to obtain image S. The primary processed target hyperspectral remote sensing images without corresponding labels in the training set are mapped to obtain image T. The primary processed target hyperspectral remote sensing images with corresponding labels in the test set are mapped to obtain image U. S1-4: Each image S is divided into two source domain sub-images of the same dimensional size. The dimensional size of the two source domain sub-images is 1 / 2 of the dimensional size of the image S. The two source domain sub-images divided based on the same image S are a pair of source domain sub-images, called a source domain sub-image pair; each image T is divided into two target domain sub-images of the same dimensional size. The dimensional size of the two target domain sub-images is 1 / 2 of the dimensional size of the image T. The two target domain sub-images divided based on the same image T are a pair of target domain sub-images, called a target domain sub-image pair.

[0022] S2: Construct a hyperspectral remote sensing image classification network. The network structure of the hyperspectral remote sensing image classification network is as follows: Figure 1 As shown in Figure 1, the hyperspectral remote sensing image classification network includes a feature extraction network, a domain alignment network, and a classifier connected in sequence, as follows: In this application, the feature extraction network is used to extract global features and local features from paired source domain sub-images (i.e., source domain sub-image pairs) to obtain a source domain feature map M s ; It is also used to extract global features and local features of paired target domain sub-images (i.e., target domain sub-image pairs) to obtain the target domain feature map M t ; In this application, the network structure of the feature extraction network is as follows: Figure 2As shown, the feature extraction network includes a Mamba branch and a convolution branch; the input end of the Mamba branch is connected to the input end of the convolution branch, the output end of the Mamba branch and the output end of the convolution branch are connected to the Concat layer, the Concat layer is sequentially connected to the first 2D convolution layer and the first Add layer, and the input end of the first Add layer is also connected to the input end of the convolution branch and the input end of the Mamba branch; In the present application, the Mamba branch includes a first normalization layer, a first linear layer, a depth-separable convolution layer, a self-attention module, a first SiLU layer, a 2D selective scanning module, a second normalization layer, and an element-by-element multiplication module connected in sequence, the output end of the first normalization layer is also connected in sequence to the second linear layer and the second SiLU layer, the output end of the second SiLU layer is connected to the input end of the element-by-element multiplication module, and the element-by-element multiplication unit is also connected in sequence to the third linear layer and the second Add layer; the input end of the first normalization layer is connected to the input end of the second Add layer; In the present application, the convolution branch includes the second 2D convolution layer, the first RELU layer, the third 2D convolution layer, the second RELU layer, the fourth 2D convolution layer, the third normalization layer and the third Add layer connected in sequence; wherein the input end of the second 2D convolution layer is connected to the input end of the first normalization layer in the Mamba branch, and the input end of the second 2D convolution layer is also connected to the input end of the third Add layer; in the present application, the output end of the third Add layer and the output end of the second Add layer are both connected to the Concat layer, and the input end of the second 2D convolution layer is connected to the input end of the first Add layer. In the present application, the convolution kernel size of the second convolution layer and the third convolution layer is 3×3, and the convolution kernel size of the fourth 2D convolution layer is 1×1.

[0023] When the input to the feature extraction network is a pair of source domain sub-images, one of the source domain sub-images enters the convolution branch and obtains a convolution branch with dimensions H×W× The local fusion feature map has rich complementary information and local spatial information, and the other source domain sub-image enters the Mamba branch to obtain a fusion feature map with a dimension of H×W×C and complementary information. The local fusion feature map and the fusion feature map are concatenated using the Concat layer, and then the first 2D convolution layer is used for convolution operation to obtain a feature map with rich global spatial information and local spatial information. Then, the feature map with rich global spatial information and local spatial information and the above two source domain sub-images are added using the first Add layer to obtain a source domain feature map M with rich complementary information, global spatial information and local spatial information. s ; Similarly, when the input to the feature extraction network is a pair of target domain sub-images, the target domain feature map M output by the feature extraction network is t It also contains rich supplementary information, global spatial information and local spatial information.

[0024] This application takes a pair of source domain sub-images as input to the feature extraction network as an example to introduce the working principle of the feature extraction network in this application: In this application, the Mamba branch is used to extract the global features of one of the source domain sub-images to obtain a feature vector with dimensions H×W× The global feature fusion feature map of the present application has rich complementary information, spatial information and contextual information, and also has rich enhanced key spatial information and key contextual information. H represents the height of image S, W represents the width of image S, and C represents the dimension of image S; the convolution branch is used to extract the local features of another source domain sub-image, and the dimension is H×W× The Concat layer is used to concatenate the feature maps output by the Mamba branch and the convolution branch to obtain a fused feature map with a dimension of H×W×C and rich global spatial features and rich local detail spatial features; the first 2D convolution layer is used to further fuse the feature maps output by the first Concat layer; the first Add layer is used to add the two source domain sub-images input to the feature extraction network and the feature maps output by the first 2D convolution layer element by element to obtain a fused feature map with a dimension of H×W×C and complementary information; specifically: In the present application, in the Mamba branch, the first normalization layer is used to normalize the feature map in one of the source domain sub-images in the input feature extraction network; the first linear layer is used to perform a linear transformation on the feature map output by the first normalization layer; the depth-separable convolution layer is used to extract features from the feature map output by the first linear layer to obtain a global feature map with spatial information and contextual information; the self-attention module captures the long-distance dependencies in the feature map output by the depth-separable convolution layer to obtain a global feature map with enhanced spatial information and contextual information; the first SiLU layer is used to activate the feature map output by the self-attention module; the 2D selective scanning module is used to selectively scan the feature map output by the first SiLU layer to obtain a global feature map with enhanced key spatial information and contextual information; the second normalization layer is used to normalize the feature map output by the 2D selective scanning module; The second linear layer is used to perform linear processing on the feature map output by the first normalization layer; the second SiLU layer is used to activate the feature map output by the second linear layer; the element-by-element multiplication module is used to perform element-by-element multiplication on the feature map output by the second SiLU layer and the feature map output by the second normalization layer to obtain a global feature map with richer spatial information and contextual information; the third linear layer is used to perform linear transformation on the feature map output by the element-by-element multiplication unit; the second Add layer is used to perform feature addition on the feature map output by the third linear layer and the feature map input to the first normalization layer to obtain a feature map with a dimension of H×W× The global fusion feature map of the present application has rich complementary information, spatial information and contextual information, and also has rich enhanced key spatial information and key contextual information; In this application, in the convolution branch, the second 2D convolution layer is used to perform a convolution operation on another source domain sub-image of the input feature extraction network to obtain a feature map with local features; the first RELU layer is used to activate the feature map output by the second 2D convolution layer; the third 2D convolution layer is used to perform a convolution operation on the feature map output by the first RELU layer to obtain a local feature map with rich detail information; the second RELU layer is used to activate the feature map input therein; the fourth 2D convolution layer is used to perform a convolution operation on the feature map output by the second RELU layer to obtain a local feature map with a dimension of H×W× The third normalization layer normalizes the feature map output by the fourth 2D convolution layer. The third Add layer is used to add the feature map output by the third normalization layer and the feature map input by the second 2D convolution layer to obtain a feature map with a dimension of H×W× The local fusion feature map of the present application has rich complementary information and rich local spatial information.

[0025] In the present application, the domain alignment network includes a feature enhancement network, the input end of the feature enhancement network is connected to the output end of the feature extraction network, the output end of the feature enhancement network is connected to the distribution alignment module and the graph alignment module, the distribution alignment module includes a teacher network and a student network, the graph alignment module includes a branch I and a branch II, the input end of the teacher network and the input end of the branch I are both connected to the output end of the feature enhancement network, the input end of the student network and the input end of the branch II are also connected to the output end of the feature enhancement network; in the present application, the source domain feature map M output by the feature extraction network s After the feature enhancement network is used for enhancement, the source domain enhanced feature map is obtained, and the target domain feature map M output by the feature extraction network is obtained. tAfter the feature enhancement network is used for enhancement processing, the target domain enhanced feature map is obtained. In this application, the feature extraction network is used for the source domain feature map M s and the target domain feature map M t The enhancement processing method is the same; wherein, the source domain enhanced feature map is used as the input of the teacher network and branch I, and the target domain enhanced feature map is used as the input of the student network and branch II; the output end of the teacher network and the output end of branch I are both connected to the input end of the classifier; in the present application, in the process of training the hyperspectral remote sensing image classification network, the classifier is used to classify and predict the feature maps output by the teacher network and branch I, and output the classification category prediction results; in the testing process of the present application, the classifier is used to classify and predict the feature maps output by the feature enhancement network, and output the classification category prediction results; and in the present application, when classifying and predicting the hyperspectral remote sensing image to be classified and predicted, the classifier is used to classify and predict the feature maps output by the feature enhancement network, and output the classification category prediction results.

[0026] In this application, the network structure of the feature enhancement network is as follows: Figure 3 As shown, the feature enhancement network includes the fifth 2D convolution layer, GELU layer, sixth 2D convolution layer, and channel segmentation layer connected in sequence. The output of the channel segmentation layer is connected to four depth-separable convolution layers. The outputs of the four depth-separable convolution layers are all connected to the Concat layer. The Concat layer is connected to the seventh 2D convolution layer and the Add layer in sequence. The input of the Add layer is also connected to the input of the fifth 2D convolution layer. The convolution kernel size of the fifth 2D convolution layer is 3×3, and the convolution kernel sizes of the sixth and seventh 2D convolution layers are 1×1. In this application, the source domain feature map As an example to introduce the function of feature enhancement network, the fifth 2D convolutional layer is used to map the source domain feature map A convolution operation is performed to obtain a feature map with richer spatial information and the number of channels is changed from 48 to 96. The GELU layer activates the feature map output by the fifth 2D convolution layer. The sixth 2D convolution layer performs a convolution operation on the feature map after the activation of the GELU layer to obtain a feature map with richer local spatial information and the number of channels is changed from 96 to 48. The channel segmentation layer performs channel segmentation on the feature map output by the sixth 2D convolution layer to obtain four segmentation sub-blocks, each of which has 12 channels. In this application, four depth-separable convolution layers are used to perform convolution operations on the four segmentation sub-blocks respectively, and the four The convolution kernel sizes of the four depth-wise separable convolution layers are 3×3, 5×5, 7×7, and 11×11, respectively. This application uses four depth-wise separable convolution layers with different convolution kernel sizes to perform convolution operations on the four segmented sub-blocks, which can effectively enhance the spatial expression ability of the features, capture local features of different scales, and obtain four feature maps with local features of different scales. The Concat layer concatenates the feature maps output by the four depth-wise separable convolution layers, and the seventh 2D convolution layer performs channel fusion on the feature maps output by the Concat layer, and outputs and inputs the feature map of the fifth 2D convolution layer (i.e., the source domain feature map). ) Feature maps with the same number of channels. The Add layer is used to add feature maps output by the seventh 2D convolutional layer and the source domain feature map Features are added to obtain a source domain enhanced feature map in which local spatial information and global spatial information are enhanced.

[0027] Similarly, when the input to the feature enhancement network is the target domain feature map M t When , the target domain enhanced feature map output by the feature enhancement network also contains enhanced local spatial information and global spatial information.

[0028] In this application, the network structure of the teacher network is as follows Figure 4 As shown, the teacher network includes four channel conversion units connected in sequence, each of which includes a convolutional layer, a batch normalization layer, a ReLU layer, and a maximum pooling layer connected in sequence, wherein the convolution kernel size of the convolution layer in the first channel conversion unit is 7×7, the convolution kernel size of the convolution layer in the second channel conversion unit, the convolution kernel size of the convolution layer in the third channel conversion unit, and the convolution kernel size of the convolution layer in the fourth channel conversion unit are all 3×3; In the present application, the number of channels of the source domain enhanced feature map input to the teacher network is 32, and the convolution layers in the four channel conversion units are all used for convolution operations. After the convolution operation is performed on the convolution layer in the first channel conversion unit, a feature map with primary features and 64 channels is obtained. The convolution operation is performed on the convolution layer in the second channel conversion unit to obtain a feature map with intermediate features and 128 channels. The convolution operation is performed on the convolution layer in the third channel conversion unit to obtain a feature map with advanced features and 256 channels. The convolution operation is performed on the convolution layer in the fourth channel conversion unit to obtain a feature map with more advanced features and 512 channels. feature map; the batch normalization layers in the four channel conversion units in the present application are all used to perform batch normalization operations on the feature maps, the ReLU layers are all used to activate the feature maps output by the batch normalization layers, and the maximum pooling layers are all used to perform maximum pooling operations on the feature maps output by the ReLU layers. In the present application, the feature map output by the fourth channel conversion unit is a feature map having inter-class features in the source domain enhanced feature map and having 512 channels; that is, after the source domain enhanced feature map is sequentially processed by the four channel conversion units in the teacher network, a feature map having inter-class features in the source domain enhanced feature map and having 512 channels can be obtained; In this application, the knowledge learned by the teacher network in the process of processing the source domain enhanced feature map is passed to the student network by means of knowledge distillation, so that the feature map obtained by the student network in the process of processing the target domain enhanced feature map can be aligned with the feature map obtained by the teacher network processing the source domain enhanced feature map (this mutual alignment is distribution alignment), so that the feature map output by the teacher network and the feature map output by the student network have a higher similarity. In this application, during the training of the hyperspectral remote sensing image classification network, the feature map obtained by the teacher network processing the source domain enhanced feature map is transmitted to the classifier.

[0029] In this application, the network structure of the student network is as follows Figure 5As shown, the student network includes four dimensional conversion units connected in sequence, and the four dimensional conversion units all include fully connected layers and ReLU layers connected in sequence, wherein the fully connected layer in the first dimensional conversion unit is used to extract primary features from the feature map, and convert it from 1024 dimensions to 64 dimensions, so as to obtain a feature map with a dimension of 64 and primary features, the fully connected layer in the second dimensional conversion unit is used to extract intermediate features from the feature map to enhance semantic information, and convert the feature map from 64 dimensions to 128 dimensions, so as to obtain a feature map with a dimension of 128 and rich intermediate features, the fully connected layer in the third dimensional conversion unit is used to extract advanced features from the feature map, and convert the feature map from 128 dimensions to 256 dimensions, so as to obtain a feature map with a dimension of 256 and advanced features, the fully connected layer in the fourth dimensional conversion unit is used to extract higher-level features from the feature map, and convert the feature map from 256 dimensions to 512 dimensions, so as to obtain a feature map with a dimension of 512 and higher-level features; wherein the ReLU layers in the four dimensional conversion units are all used to perform activation operations on the feature maps.

[0030] In this application, the graph alignment module includes branch I and branch II. Branch I and branch II have the same structure and function. The network structure of branch I is as follows: Figure 6 As shown; Branch I includes two graph convolutional layers GraphSAGE connected in sequence, wherein the first graph convolutional layer GraphSAGE is used to capture global spatial topological structure information (the spatial topological structure information is graph structure information), and the global spatial topological structure information includes the global node relationships, and the second graph convolutional layer GraphSAGE is used to capture more detailed local spatial topological structure information, and the local spatial topological structure information includes the local node relationships; the acquisition website of the graph convolutional layer GraphSAGE is: https: / / proceedings.neurips.cc / paper / 2017 / hash / 5dd9db5e033da9c6fb5ba83c7a7ebea9-Abstract.html; in this application, after the source domain enhanced feature map is processed in sequence by the two graph convolutional layers GraphSAGE in branch I, a feature map with global spatial topological structure information and local spatial topological structure information can be obtained; Then, branch II processes the target domain enhanced feature map. Through optimal transportation, the feature map obtained by branch II in the process of processing the target domain enhanced feature map can be aligned with the feature map obtained by branch I in processing the source domain enhanced feature map (this mutual alignment is image alignment), so that the feature map output by branch I and the feature map output by branch II are aligned. Figure 2 In this application, during the training of the hyperspectral remote sensing image classification network, the feature map output by branch I is transmitted to the classifier.

[0031] S3: Based on the total loss of the hyperspectral remote sensing image classification network, the hyperspectral remote sensing image classification network is trained using the source domain sub-image pairs and the target domain sub-image pairs to obtain a hyperspectral remote sensing image classification network model; specifically, the following steps are included: A pair of source domain sub-images and a pair of target domain sub-images are respectively input into the hyperspectral remote sensing image classification network, the total loss of the hyperspectral remote sensing image classification network is calculated, and then the gradient is optimized and back-propagated, the model parameters of the hyperspectral remote sensing image classification network are updated, and an epoch of training is completed; in this application, a pair of source domain sub-images and a pair of target domain sub-images are input for each epoch of training. During the training process of each epoch, the input source domain sub-image pair is randomly selected from all source domain sub-image pairs, and the input target domain sub-image is also randomly selected from all target domain sub-images; in this application, the training is repeated for 500 epochs. poch completes the training of a training segment. After a training segment, the classifier outputs the classification category prediction result and classification accuracy. In the process of training the hyperspectral remote sensing image classification network in this application, the classification accuracy is calculated by comparing the source domain prediction result with the source domain label; repeat the training of the training segment until the classification accuracy of the next training segment output is greater than the classification accuracy of the current training segment (the next training segment and the current training segment are two adjacent training segments), then save the parameters of the hyperspectral remote sensing image classification network during the last epoch training of the next training segment as the final model parameters of the hyperspectral remote sensing image classification network, and obtain the hyperspectral remote sensing image classification network model.

[0032] In this application, the total loss of the hyperspectral remote sensing image classification network includes distribution alignment and distillation loss. , graph alignment loss , classification loss of the teacher network , classification loss of branch I and consistency constraint loss , as shown in formula (1): = + + + + (1) In formula (1), distribution alignment and distillation loss Used to calculate the loss caused by knowledge distillation and the loss caused by distribution alignment of the teacher network and the student network, distribution alignment and distillation loss The calculation method of is as shown in formula (2): = α· +(1-α)· (2) In formula (2), represents the distillation loss, represents the loss caused by the distribution alignment of the teacher network and the student network, α is a balance parameter, and its value is set to 0.5 in this embodiment; In formula (2), the distillation loss Used to calculate the loss generated during the knowledge distillation process from the teacher network to the student network, distillation loss The calculation method of is shown in formula (3). = (3) In formula (3), and Represent the feature maps output by the teacher network and the student network respectively, T is the temperature coefficient, KL express KL Divergence, S Represents the calculation of the softmax function with temperature.

[0033] In formula (2), the loss generated by the distribution alignment of the teacher network and the student network is The calculation method of is as shown in formula (4): = (4) In formula (2), and denote the i-th feature map output by the teacher network and the i-th feature map output by the student network, respectively. N is the number of feature maps output by the teacher network. In this application, the number of feature maps output by the student network is the same as the number of feature maps output by the teacher network. Represents second normal form computation.

[0034] In formula (1), the graph alignment loss Used to calculate the loss caused by aligning the source domain enhanced feature map and the target domain enhanced feature map, the map alignment loss The calculation method is as shown in formula (5): (5) In formula (5), Indicates The output of the graph convolutional layer is =0, represents the alignment result between the source domain enhanced feature map and the target domain enhanced feature map; when =1, Represents the features with source domain features output by the first graph convolution layer of branch I and the first graph convolution layer of branch II Figure Ⅰ and features with target domain characteristics Figure Ⅰ The alignment result between =2, Represents the features with source domain characteristics output by the second graph convolution layer of branch I and the second graph convolution layer of branch II Figure II and features with target domain characteristics Figure II The alignment results between .

[0035] In the present application formula (5), The calculation method of is the same as that of formula (7) disclosed in the paper "Topological Structure and Semantic Information Transfer Network for Cross-Scene Hyperspectral Image Classification"; In formula (1), Represents the classification loss of the teacher network, which is used to calculate the loss between the feature map output by the teacher network and its corresponding true label, and the convolution classification loss The calculation method of is shown in formula (7): (7) In formula (7), and They represent the i-th feature map output by the teacher network and the true label corresponding to the feature map, S represents the softmax function calculation, represents the cross entropy loss calculation, and ns represents the number of feature maps output by the teacher network.

[0036] In formula (1), the classification loss of branch I is Used to calculate the loss between the feature map output by branch I and its corresponding true label, and the classification loss of branch I The calculation method of is shown in formula (8): (8) In formula (8), and They represent the i-th feature map output by branch I and the true label corresponding to the feature map, S represents the softmax function calculation, represents the cross entropy loss calculation, and NS represents the number of feature maps output by branch I.

[0037] In formula (1), the consistency constraint loss The consistency constraint loss is used to calculate the difference between the probability prediction results obtained by the classifier for the feature map output by the teacher network and the feature map output by branch I. The calculation method of is shown in formula (9): ( = (9) In formula (9), C represents the number of categories, and They respectively represent the probability prediction results of the classifier for the feature maps output by the teacher network and the feature maps output by branch I, and ns represents the number of feature maps output by the teacher network. In this application, the number of feature maps output by the teacher network and branch I is the same.

[0038] S4. Fill pixels of the same size on each boundary of the hyperspectral remote sensing image to be classified and predicted, and the size of the filled pixels is 1 / 2 the patch size plus 1, and then map to obtain image X, divide image X into two sub-images to be predicted of the same size, and the two sub-images to be predicted constitute a sub-image pair to be predicted, and then input the sub-image pair to be predicted into the hyperspectral remote sensing image classification network model obtained in step S3, perform feature extraction through the feature extraction network, and then use the feature enhancement network to perform feature enhancement to obtain an enhanced feature map, and then input the enhanced feature map into the classifier for classification prediction, and the image classification category prediction result and classification accuracy can be obtained.

[0039] test: In order to compare the classification effects of the hyperspectral remote sensing image classification method of the present invention with seven existing hyperspectral image classification methods in the prior art, including the SVM method (from the paper "Support Vector Method of Classification"), the DAN method (from the paper "Learning transferable features with deep adaptation networks"), the MRAN method (from the paper "Deep subdomain adaptation network for image classification"), the TSTnet method (from the paper "Topological Structure and Semantic Information Transfer Network for Cross-Scene Hyperspectral Image Classification"), the DSAN method (from the paper "Deep subdomain adaptation network for image classification"), the DAAN method (from the paper "Deep subdomain adaptation network for image classification"), and the CEGCN method (from the paper "Deep subdomain adaptation network for image classification"), this application uses the above seven existing hyperspectral image classification methods and the remote sensing image classification method described in this application to perform classification tests on the image U obtained in step S1 of this application, and obtains test results such as classification accuracy CA, overall accuracy OA, and Kappa coefficient KAPPA, as shown in Table 2. During the classification test of the present application, the image U is divided into two test sub-images of the same size, and the two test sub-images constitute a test sub-image pair. Then, the test sub-image pair is input into the hyperspectral remote sensing image classification network model obtained in step S3, and feature extraction is performed through the feature extraction network, and then feature enhancement is performed using the feature enhancement network to obtain an enhanced feature map, and then the enhanced feature map is input into the classifier for classification prediction.

[0040] Table 1 shows the types of objects in the HYRANK dataset and the number of Dioin and Loukia samples for different types of objects

[0041] The first column of Table 1 expresses the 12 types of land feature categories in the HYRANK dataset, and the second column of Table 1 shows the corresponding land feature categories; the specific number of samples in Dioin and Loukia for each land feature category is shown in the third and fourth columns of Table 1; the last row of Table 1, which is the row with the total number, shows that the total number of samples in Dioin and Loukia is 20024 and 10317 respectively.

[0042] Table 2 shows the test results of classification accuracy CA, overall accuracy OA and Kappa coefficient KAPPA obtained by different hyperspectral remote sensing image classification methods

[0043] In Table 2, the Ours method represents the hyperspectral remote sensing image classification method described in this application. The category numbers 1 to 12 recorded in the first column of Table 2 represent the 12 categories included in the HYRANK data set, which are consistent with the 12 categories shown in the first column of Table 1; the DAN method, SVM method, MRAN method, TSTnet method, CEGCN method, DSAN method, DAAN method and Ours method recorded in the first row of Table 2 represent different hyperspectral remote sensing image classification methods, and the second to ninth columns and the third to fourteenth rows show the classification accuracy CA of the image U corresponding to each type of ground object when the image U corresponding to 12 types of ground object categories is classified and predicted by eight hyperspectral remote sensing image classification methods such as the DAN method, SVM method, MRAN method, TSTnet method, CEGCN method, DSAN method, DAAN method and Ours method, wherein the classification accuracy CA refers to the ratio between the number of correct classifications obtained by a certain hyperspectral remote sensing image classification method for classifying and predicting an image U corresponding to a certain category and the total number. As shown in the value 11.66% at the intersection of the third row and the second column in Table 2, it means that the classification accuracy CA obtained by testing the image U corresponding to the ground object category 1 using the DAN method is 11.66%.

[0044] In Table 2, the second to last row represents the overall accuracy OA, and the first to last row represents the Kappa coefficient KAPPA. The higher the classification accuracy CA, overall accuracy OA, and Kappa coefficient KAPPA indicators, the higher the classification accuracy of the model. Among them, the overall accuracy OA is the ratio of the sum of correct samples of all categories output by the hyperspectral remote sensing image classification network model to the total number of test samples. The remote sensing image classification method described in this application has the highest accuracy compared to the above seven existing hyperspectral remote sensing image classification methods. Compared with the highest OA value of 62.66% that can be achieved by the above seven existing hyperspectral remote sensing image classification methods, the OA value obtained by the hyperspectral remote sensing image classification method described in this application is improved by ((0.6488-0.6266) / 0.6266)×100%=3.54%, which shows that the correct classification ability of the present application is more outstanding; The KAPPA coefficient is a standardized accuracy indicator that takes into account the accuracy of random classification. The closer the KAPPA value is to 1, the better the performance of the classifier. Compared with the above-mentioned existing hyperspectral remote sensing image classification methods, the KAPPA accuracy obtained by the remote sensing image classification method described in this application is also the highest, with a KAPPA value of 0.5599. The KAPPA value obtained by the remote sensing image classification method described in this application is higher than the highest KAPPA value that can be obtained by the above-mentioned seven existing hyperspectral remote sensing image classification methods by ((0.6012-0.5420) / 0.5420)×100%=10.92%; this shows that the classification categories obtained by the remote sensing image classification method described in this application when performing image classification have higher consistency.

[0045] In addition, it can be seen from Table 2 that the hyperspectral remote sensing image classification method described in the present application has a higher classification accuracy CA for images corresponding to the second and fourth types of land objects in the HYRANK data set than that obtained by the above-mentioned seven existing hyperspectral remote sensing image classification methods. Among them, when the hyperspectral remote sensing image classification method described in the present application classifies and predicts the image U corresponding to the second type of land object, the classification accuracy CA is improved by ((0.5378-0.4344) / 0.4344)×100%=23.80 compared with the highest classification accuracy CA obtained by the above-mentioned seven existing hyperspectral remote sensing image classification methods. %, when the hyperspectral remote sensing image classification method described in the present application classifies and predicts the image U corresponding to the fourth type of land object category, the classification accuracy CA is improved by ((0.2407-0.1177) / 0.1177)×100%=104.50% compared with the highest classification accuracy CA obtained by the above-mentioned seven existing hyperspectral remote sensing image classification methods. This shows that the hyperspectral remote sensing image classification method described in the present application has a more significant classification accuracy for the second type of land object and the fourth type of land object in the HYRANK data set, especially for the fourth type of land object in the HYRANK data set.

Claims

1. A hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features, characterized by: The specific steps include: S1: Obtain source domain sub-image pairs and target domain sub-image pairs for training; S2: Construct a hyperspectral remote sensing image classification network. The hyperspectral remote sensing image classification network includes a feature extraction network, a domain alignment network and a classifier connected in sequence. The feature extraction network is used to extract global features and local features, the domain alignment network is used to achieve distribution alignment and graph alignment, and the classifier is used to predict classification results. The domain alignment network includes a feature enhancement network and a distribution alignment module and a graph alignment module respectively connected to the output end of the feature enhancement network. The distribution alignment module includes a teacher network and a student network. The graph alignment module includes branches I and II. The input end of the teacher network and the input end of branch I are both connected to the output end of the feature enhancement network. The source domain output of the feature enhancement network The enhanced feature map is used as the input of the teacher network and branch I; the input of the student network and the input of branch II are connected to the output of the feature extraction network, and the target domain enhanced feature map output by the feature enhancement network is used as the input of the student network and branch II; the output of the teacher network and the output of branch I are connected to the input of the classifier; the knowledge learned by the teacher network is transferred to the student network by knowledge distillation, so that the feature map obtained by the student network can be aligned with the feature map obtained by the teacher network, and the mutual alignment is distribution alignment; through optimal transportation, the feature map obtained by branch II is aligned with the feature map obtained by branch I, and the mutual alignment is graph alignment; S3: Based on the total loss of the hyperspectral remote sensing image classification network, the hyperspectral remote sensing image classification network is trained using the source domain sub-image pairs and the target domain sub-image pairs to obtain the hyperspectral remote sensing image classification network model; S4. Fill pixels of the same size on each boundary of the hyperspectral remote sensing image to be classified and predicted. The pixel size is 1 / 2 of the patch size plus 1, and then map it to obtain image X. Divide image X into sub-image pairs to be predicted. Then, input the sub-image pairs to be predicted into the hyperspectral remote sensing image classification network model, process them through the feature extraction network and the feature enhancement network in turn, and then perform classification prediction through the classifier to obtain the image classification category prediction result and classification accuracy.

2. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 1 is characterized in that: Step S1 specifically includes the following steps: obtaining a hyperspectral remote sensing image dataset, dividing a source hyperspectral remote sensing image and a target hyperspectral remote sensing image according to the scenes in the hyperspectral remote sensing image dataset; and preprocessing the source hyperspectral remote sensing image and the target hyperspectral remote sensing image to obtain image S, image T and image U, and then obtaining a source domain sub-image pair and a target domain sub-image pair based on image S and image T, respectively.

3. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 1 is characterized in that: In step S2, the feature extraction network includes a Mamba branch and a convolution branch; the input end of the Mamba branch is connected to the input end of the convolution branch, the output end of the Mamba branch and the output end of the convolution branch are both connected to the Concat layer, the Concat layer is sequentially connected to the first 2D convolution layer and the first Add layer, and the input end of the first Add layer is also connected to the input end of the convolution branch and the input end of the Mamba branch.

4. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 3 is characterized in that: In step S2, the Mamba branch includes a first normalization layer, a first linear layer, a depth-wise separable convolution layer, a first SiLU layer, a 2D selective scanning module, a second normalization layer, and an element-by-element multiplication module connected in sequence; the output end of the first normalization layer is also connected to the second linear layer and the second SiLU layer in sequence, the output end of the second SiLU layer is connected to the input end of the element-by-element multiplication module, and the element-by-element multiplication unit is also connected to the third linear layer and the second Add layer in sequence; the input end of the first normalization layer is also connected to the input end of the second Add layer.

5. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 3 is characterized in that: In step S2, the convolution branch includes a second 2D convolution layer, a first RELU layer, a third 2D convolution layer, a second RELU layer, and a fourth 2D convolution layer connected in sequence; the input end of the second 2D convolution layer in the convolution branch is connected to the input end of the first normalization layer in the Mamba branch, and the output end of the fourth 2D convolution layer and the output end of the second Add layer are both connected to the Concat layer; the input end of the second 2D convolution layer is also connected to the input end of the first Add layer.

6. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 1 is characterized in that: In step S2, the teacher network includes four channel conversion units connected in sequence, each of which includes a convolutional layer, a batch normalization layer, a ReLU layer and a maximum pooling layer connected in sequence, wherein the convolution kernel size of the convolution layer in the first channel conversion unit is 7×7, the convolution kernel size of the convolution layer in the second channel conversion unit, the convolution kernel size of the convolution layer in the third channel conversion unit, and the convolution kernel size of the convolution layer in the fourth channel conversion unit are all 3×3.

7. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 1 is characterized in that: In step S2, the student network includes four dimensional conversion units connected in sequence, and each of the four dimensional conversion units includes a fully connected layer and a ReLU layer connected in sequence.

8. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 1 is characterized in that: In step S2, branch I and branch II have the same structure and function; branch I includes two graph convolutional layers connected in sequence, where the first graph convolutional layer is used to capture global spatial topological structure information, and the second graph convolutional layer is used to capture more detailed local spatial topological structure information.

9. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 1, characterized in that: In step S2, the feature enhancement network includes a fifth 2D convolutional layer, a GELU layer, a sixth 2D convolutional layer, and a channel segmentation layer connected in sequence, the output end of the channel segmentation layer is connected to four depth-separable convolutional layers, the output ends of the four depth-separable convolutional layers are all connected to the Concat layer, the Concat layer is sequentially connected to the seventh 2D convolutional layer and the Add layer, and the input end of the Add layer is also connected to the input end of the fifth 2D convolutional layer; wherein the convolution kernel size of the fifth 2D convolutional layer is 3×3, and the convolution kernel size of the sixth 2D convolutional layer and the seventh 2D convolutional layer is 1×1; the feature enhancement network is used to enhance the local spatial information and the global spatial information in the feature map output by the feature extraction network.

10. The hyperspectral remote sensing image classification method based on spatial topological structure and inter-class features according to claim 1, characterized in that: Step S3 specifically includes the following steps: inputting a pair of source domain sub-images and a pair of target domain sub-images into the hyperspectral remote sensing image classification network, calculating the total loss of the hyperspectral remote sensing image classification network, and then optimizing the gradient and back-propagating, updating the model parameters of the hyperspectral remote sensing image classification network, and completing the training of one epoch; repeating the training for 500 epochs to complete the training of one training segment, after one training segment ends, the classifier outputs the classification category prediction result and classification accuracy; repeating the training of the training segment until the classification accuracy output by the next training segment is greater than the classification accuracy of the current training segment, then saving the parameters of the hyperspectral remote sensing image classification network during the last epoch training of the next training segment as the final model parameters, and obtaining the hyperspectral remote sensing image classification network model.

Citation Information

Patent Citations

  • Federal domain adaptation method and system based on knowledge distillation

    CN115761408A

  • Cross-domain remote sensing image scene classification method based on rotation robust feature subclass center alignment

    CN116152671A

  • Method and device for performing target detection on optical remote sensing image

    CN118968016A

  • Remote sensing image unsupervised domain adaptation method based on comparative learning and multi-prototype alignment

    CN119251646A

  • Supervised domain adaptation

    US11170581B1

Cited By

  • Pigment classification method and related device

    CN120747644A