Hyperspectral image classification method based on dynamic scaling and graph optimal transmission
The hyperspectral image classification model with dynamic scale adjustment and graph optimal transmission solves the problems of scale adaptability and graph structure information loss, achieves high-precision classification of hyperspectral images, and improves the compatibility and information expression ability of multi-network feature fusion.
Patent Information
- Application Number
- CN202510859954.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing hyperspectral remote sensing image classification methods face problems of scale adaptability, graph structure information loss, and multi-network feature fusion compatibility, making it difficult to achieve optimal observation scale adaptation and effective information expression, resulting in limited classification accuracy.
A hyperspectral image classification model based on dynamic scale adjustment and graph optimal transmission is adopted. The sampling window size is dynamically adjusted through the adaptive scale ViT network branch and the dynamic graph optimal transmission network branch, and the indirect coupling of different network features and multi-channel feature fusion are realized through the graph optimal transmission theory.
It improves the accuracy and category distinction ability of hyperspectral image classification, optimizes the classification performance, avoids information conflict when fusing different network features, and enhances the expression ability of spectral information in the graph structure.
Smart Images

Figure CN120375100B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image classification, and in particular relates to a hyperspectral image classification method based on dynamic scaling and graph optimal transmission. Background Art
[0002] Hyperspectral remote sensing imaging is an important Earth observation technology that can obtain spectral information about objects in visible light and other wavelengths. Hyperspectral image (HSI) classification, a key research area in this field, has broad applications in urban planning, precision agriculture, ecological monitoring, and national defense and military fields.
[0003] Traditional HSI classification methods, such as support vector machines (SVMs) and random forests, primarily rely on spectral features, offering low computational costs. However, these methods struggle to effectively address spectral variations, which compromises classification accuracy. With the recent development of deep learning techniques, researchers have incorporated spatial information to enhance classification capabilities. Convolutional neural network (CNN)-based models enable joint spectral-spatial analysis, and further developments such as feature pyramid networks (FPNs) and U-Nets have been applied to hyperspectral imagery classification. In addition to CNNs, the self-attention-based Visual Transformer (ViT) has demonstrated excellent performance in image classification tasks, particularly for modeling long-range dependencies in HSI classification. Furthermore, graph neural networks (GNNs) are becoming an increasingly important tool for HSI classification due to their ability to model non-Euclidean data structures. By constructing a graph structure between pixels or superpixels, GNNs can effectively capture relationships between data, and combined with the graph attention (GAT) mechanism, they can adaptively update the adjacency matrix. Although current deep learning-based HSI classification methods have achieved some progress, further improvements in HSI classification performance remain limited by the following issues:
[0004] Scale adaptability: Most classification models rely on fixed-size sampling windows for analysis, and multi-scale methods still require a predefined scale configuration. However, the optimal observation scale varies depending on the object type, and fixed scale constraints make it difficult for HSI classification to adaptively select the optimal sampling scale.
[0005] Graph structure information loss: Traditional GNN methods are usually calculated based on a fixed adjacency matrix, which causes the graph structure information to be compressed and cannot flexibly update the relationship between graph nodes, limiting the information representation capability.
[0006] Compatibility issues with multi-network feature fusion: In a multi-network fusion architecture, different models have different ways of expressing features. For example, GCN feature representations include relationships between nodes, while CNN feature mapping relies on the receptive field range. Direct fusion of multiple networks may cause information conflicts, affecting classification performance. Summary of the Invention
[0007] In view of this, the present invention aims to provide a hyperspectral image classification method based on dynamic scaling and graph optimal transmission to solve the problems in hyperspectral image classification technology based on deep learning. The present invention proposes a hyperspectral image classification model based on dynamic scaling and graph optimal transmission. The hyperspectral image classification model includes two network branches, namely the adaptive scaling ViT (AS-ViT) network branch and the dynamic graph optimal transmission (DGOT) network branch. Specifically, it includes innovative architectures such as dynamic scaling adaptive adjustment, dual-channel graph structure optimization and graph convolution fusion, and multi-network feature fusion based on graph optimal transmission. The present invention can automatically adjust the size of the sampling window during training and dynamically reconstruct the graph node structure. In addition, the dual-channel feature fusion mechanism based on graph optimal transmission can effectively alleviate the compatibility issues of different network features and achieve better classification performance.
[0008] To achieve the above object, the technical solution created by the present invention is implemented as follows:
[0009] A hyperspectral image classification method based on dynamic scaling and graph optimal transmission specifically includes the following steps:
[0010] S1: Obtain a labeled hyperspectral image dataset and use it to construct a training set;
[0011] S2: Construct a hyperspectral image classification model, which includes an adaptive scale ViT network branch and a dynamic graph optimal transmission network branch;
[0012] The adaptive scale ViT network branch is used to dynamically adjust the sampling scale of input data according to the characteristics of different land feature categories, and the dynamic graph optimal transmission network branch is used to achieve the interaction and fusion of different information;
[0013] S3: Construct an overall loss function, use the training set and the overall loss function to train the hyperspectral image classification model, and obtain a trained hyperspectral image classification model;
[0014] S4: Input the hyperspectral image to be classified into the trained hyperspectral image classification model for classification to obtain the classification result.
[0015] Furthermore, in step S1, pixels with labels are extracted from the hyperspectral image dataset, all pixel sets with labels are divided into categories, and 100 pixels are randomly selected from the pixel set of each category. For categories with a total number of less than 100 pixels, half of the total number of pixels are selected.
[0016] Principal component analysis is performed on each hyperspectral image contained in the hyperspectral image dataset, and a patch sample is constructed based on each selected pixel point. The patch sample is composed of the pixel point and the 8 neighborhood pixels of the current pixel point. The patch sample size is 9×9×b, where 9×9 represents the size of the neighborhood range and b is the number of principal components retained after dimensionality reduction. The set of corresponding patch samples constructed using each selected pixel point is used as the training set.
[0017] Furthermore, in step S2, the adaptive scale ViT network branch includes an LSS module, an attention module, a first 3D convolution block, a second 3D convolution block, a third 3D convolution block, a VIT module and a first classifier, wherein the patch sample X in the training set is input into the LSS module for processing to obtain a feature map , the feature map Input to the attention module for processing to obtain the feature map , the feature map Input to the first 3D convolution block for processing to obtain feature map A1, feature map A1 is processed by the second and third convolution blocks to obtain feature map A2, and the processing results of the first 3D convolution block, the second 3D convolution block and the third 3D convolution block are cascaded to obtain feature map , the feature map Input to the VIT module for processing to obtain the feature map , the feature map Input to the first classifier for processing to obtain the classification prediction results of the adaptive scale ViT network branch ;
[0018] The dynamic graph optimal transmission network branch includes the first DGAM module, the second DGAM module, the DLFM module, the DGOM module and the second classifier. The patch sample X in the training set is input into the first DGAM module for processing to obtain the feature map , the feature map Input to the second DGAM module for processing to obtain the feature map , the feature map and feature maps Input to the DLFM module for processing to obtain the Laplace fusion matrix ;
[0019] The feature map Perform learnable linear combination processing on the nodes and use the processing results to generate a dual-channel Laplace matrix, and combine the dual-channel Laplace matrix and the Laplace fusion matrix Input to DGOM module for processing and obtain loss and the set of feature vectors , i is 1 or 2, the feature vector set Input to the second classifier for processing to obtain the prediction result of the classification of the optimal transmission network branch of the dynamic graph .
[0020] Furthermore, the LSS module includes a 3D-CNN encoder, which inputs the patch sample X into the 3D-CNN encoder for processing to obtain the scaling factor , the scale factor Perform normalization to obtain the normalized scale factor, generate a binary mask matrix based on the scale factor, perform Hadamard product operation on the binary mask matrix and the patch sample X, and then crop to obtain the image block with the center area retained. , the image block After bicubic upsampling, samples with optimal observation scale are obtained , the sample Cascade with the patch sample X to obtain the feature map .
[0021] Furthermore, the processing flow of the first DGAM module and the second DGAM module is the same, wherein the processing flow of the first DGAM module is: the patch sample X input to the first DGAM module is dimensionally transformed to obtain the matrix , the matrix With learnable weight matrix After matrix multiplication and dimension transformation, the feature map is obtained , the feature map Feed into the GAT module for processing to obtain the feature map , using feature maps Construct a dual-channel adjacency matrix, which includes the modulus adjacency matrix and the spectral angle adjacency matrix. Input the modulus adjacency matrix and the spectral angle adjacency matrix into the single-layer graph convolution layer for convolution processing, and concatenate the processing results to obtain the dual-channel graph features. ;
[0022] The GAT module includes the first linear layer, the second linear layer, the dimension compression module and the LReLU module. The first linear layer performs feature encoding to obtain the feature map , the feature map Perform dimension transformation and replication, concatenate the feature maps obtained after replication in two different dimensions and feed them into the second linear layer for processing, and perform dimension compression and LReLU activation on the processing results to obtain the attention weight factor , the feature map and attention weight factor Multiply and output feature map .
[0023] Furthermore, the feature map After adding by channel, it is fed into the second DGAM module for processing to obtain the dual-channel image features In the DLFM module, the dual-channel graph features and dual-channel graph features Each channel of each channel constructs a dual-channel adjacency matrix to obtain eight adjacency matrices, and constructs corresponding Laplace matrices for each of the eight adjacency matrices. The two Laplace matrices corresponding to the module value adjacency matrix are weighted summed to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the spectral angle adjacency matrix of the spectral angle are weighted summed to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the modulus adjacency matrix of the weighted sum are obtained to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the spectral angle adjacency matrix of the spectral angle are weighted summed to obtain the Laplace fusion matrix , the Laplacian fusion matrix in the channel dimension and Laplace fusion matrix Cascade to obtain the Laplace fusion matrix , the Laplacian fusion matrix in the channel dimension and Laplace fusion matrix Cascade to obtain the Laplace fusion matrix , the Laplace fusion matrix and Laplace fusion matrix Add together to obtain the Laplace fusion matrix .
[0024] Furthermore, the DGOM module includes a first COPT module and a second COPT module, wherein the Laplace fusion matrix output by the DLFM module is And the feature map output by the VIT module The Laplace matrix Input to DGOM module, Laplace fusion array Including feature maps and feature maps , the Laplace matrix Including feature maps and feature maps , using the first COPT module to optimize the feature map and feature maps The distance is obtained to obtain the first COPT distance, and the second COPT module is used to optimize the feature map and feature maps The distance is obtained to obtain the second COPT distance, and the first COPT distance and the second COPT distance are adaptively fused to obtain the loss After stacking the first COPT distance and the second COPT distance, we get the feature vector set , i is 1 or 2.
[0025] Furthermore, in step S3, the overall loss function for:
[0026] ;
[0027] ;
[0028] ;
[0029] in, is the loss of the classification result produced by the first classifier, is the loss of the classification result produced by the second classifier, For the prediction results and prediction results The corresponding real feature labels, is the loss, i is the index of cn, and cn is the total number of categories.
[0030] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0031] (1) The hyperspectral image classification method based on dynamic scale and image optimal transmission created by the present invention automatically adjusts the sampling scale of the hyperspectral image through a learning mechanism to adapt to the optimal classification scale of different land object categories, thereby improving classification accuracy. In combination with a multi-scale feature fusion strategy, the maximum scale information and the optimal scale information are integrated to optimize the classification performance.
[0032] (2) The hyperspectral image classification method based on dynamic scaling and graph optimal transmission created by the present invention and the dual-channel adjacency matrix construction method based on the modulus of spectral vector differences and spectral angles enhance the expression ability of spectral information in the graph structure, and adaptively updates the graph nodes, i.e., the embedding vectors, through the learnable linear combination of nodes and the GAT mechanism, thereby improving the category differentiation ability.
[0033] (3) The hyperspectral image classification method based on dynamic scaling and graph-optimal transport described in this invention uses graph-optimal transport theory to achieve indirect coupling between CNN, ViT, and GCN features, thus avoiding information conflict when fusing features from different networks. In addition, multi-channel feature fusion is performed using a Laplacian fusion matrix, and the fusion method between different feature channels is optimized through an adaptive weighting strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0035] Figure 1 A schematic diagram of a process for a hyperspectral image classification method based on dynamic scaling and graph optimal transmission according to an embodiment of the present invention;
[0036] Figure 2 This is a schematic diagram of the structure of the hyperspectral image classification model described in the embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of the structure of the LSS module described in the embodiment of the present invention;
[0038] Figure 4 This is a schematic structural diagram of the first DGAM module according to an embodiment of the present invention;
[0039] Figure 5 A schematic diagram of the structure of the DLFM module according to an embodiment of the present invention;
[0040] Figure 6 A schematic diagram of the structure of the DGOM module according to an embodiment of the present invention;
[0041] Figure 7 A schematic diagram comparing the classification results of the embodiment of the present invention on the HS dataset with other cutting-edge methods;
[0042] Figure 8 A schematic diagram comparing the classification results of the embodiment of the present invention on the SV dataset with other cutting-edge methods;
[0043] Figure 9A schematic diagram showing a comparison of the classification indicators of the embodiment of the present invention on the HS dataset with other cutting-edge methods;
[0044] Figure 10 A schematic diagram comparing the classification indicators of the embodiment of the present invention with other cutting-edge methods on the SV dataset. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation of the present invention.
[0046] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0047] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second" and the like are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first", "second" and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0048] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art can understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0049] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0050] like Figure 1As shown, the present invention proposes a hyperspectral image classification method based on dynamic scale and graph optimal transmission, which specifically includes the following steps: S1: obtaining a hyperspectral image dataset with labels, and constructing a training set using the hyperspectral image dataset with labels; S2: constructing a hyperspectral image classification model, the hyperspectral image classification model includes an adaptive scale ViT network branch and a dynamic graph optimal transmission network branch; the adaptive scale ViT network branch is used to dynamically adjust the sampling scale of the input data according to the characteristics of different land object categories, and the dynamic graph optimal transmission network branch is used to realize the interaction and fusion of different information; S3: constructing an overall loss function, and using the training set and the overall loss function to train the hyperspectral image classification model to obtain a trained hyperspectral image classification model; S4: inputting the hyperspectral image to be classified into the trained hyperspectral image classification model for classification to obtain a classification result.
[0051] It should be noted that the specific solution of the present invention is: the HSI data set is divided into a training set, a validation set and a test set. The hyperspectral image (HSI) is reduced in dimension by principal component analysis (PCA), and on this basis, patch samples are constructed for each pixel point one by one. An adaptive scale ViT network branch (AS-ViT for short) is constructed to dynamically adjust the sampling scale of the input data according to the characteristics of different land object categories, thereby effectively solving the limitation of the fixed scale and improving the classification accuracy. A dynamic graph optimal transmission network branch (DGOT for short) is constructed and interacts with AS-ViT: a multi-network feature fusion strategy based on graph optimized transmission (COPT) theory is adopted to indirectly couple the features of CNN, visual transformer (ViT) and GCN, and realize efficient interaction and fusion of information through the Laplace fusion matrix. On this basis, a graph optimized transmission module (DGOM module) is constructed. In the DGOM module, the optimal transmission distance calculated by COPT works together with the adaptive weighting strategy to optimize the feature fusion process and enhance the compatibility between different networks. The overall loss function is designed, combining the cross entropy loss and COPT distance loss (i.e., loss ) and feature similarity constraints, supervised learning is performed through a two-branch classifier structure to improve the robustness and classification performance of the model. The training set and validation set are used to train the entire model and adjust the parameters, and the test set is used to evaluate the model performance.
[0052] In some embodiments, in step S1, pixels with labels are extracted from the hyperspectral image dataset, all sets of labeled pixels are divided into categories, and 100 pixels are randomly selected from the pixel sets of each category. For categories with less than 100 pixels, half of the total number of pixels is selected.
[0053] Principal component analysis is performed on each hyperspectral image contained in the hyperspectral image dataset, and a patch sample is constructed based on each selected pixel point. The patch sample is composed of the pixel point and the 8 neighborhood pixels of the current pixel point. The patch sample size is 9×9×b, where 9×9 represents the size of the neighborhood range and b is the number of principal components retained after dimensionality reduction. The set of corresponding patch samples constructed using each selected pixel point is used as the training set.
[0054] It is important to note that comprehensive training, validation, and test sets are constructed to ensure the effectiveness and robustness of the hyperspectral image classification model. The specific process is as follows: First, the labeled pixels in the hyperspectral image dataset are divided. For each category, 100 pixels are randomly selected from the hyperspectral image dataset to construct the training and validation sets. For categories with fewer than 100 pixels, half of the total number of pixels in that category is selected for sample construction. The remaining labeled pixels are used to evaluate the model's classification accuracy, ensuring its effectiveness in practical applications.
[0055] After the dataset is partitioned, PCA (Principal Component Analysis) is used to reduce the dimensionality of the hyperspectral image and construct patch samples. A patch sample is a local HSI image block consisting of the spectrum of the target pixel and the spectra within its neighborhood. The patch size is 9×9×b, where 9×9 represents the size of the neighborhood and b is the number of principal components retained after dimensionality reduction.
[0056] In some embodiments, in step S2, the adaptive scale ViT network branch includes an LSS module, an attention module, a first 3D convolution block, a second 3D convolution block, a third 3D convolution block, a VIT module and a first classifier, wherein the patch sample X in the training set is input to the LSS module for processing to obtain a feature map , the feature map Input to the attention module for processing to obtain the feature map , the feature map Input to the first 3D convolution block for processing to obtain feature map A1, feature map A1 is processed by the second and third convolution blocks to obtain feature map A2, and the processing results of the first 3D convolution block, the second 3D convolution block and the third 3D convolution block are cascaded to obtain feature map , the feature map Input to the VIT module for processing to obtain the feature map , the feature map Input to the first classifier for processing to obtain the classification prediction results of the adaptive scale ViT network branch ;
[0057] The dynamic graph optimal transmission network branch includes the first DGAM module, the second DGAM module, the DLFM module, the DGOM module and the second classifier. The patch sample X in the training set is input into the first DGAM module for processing to obtain the feature map , the feature map Input to the second DGAM module for processing to obtain the feature map , the feature map and feature maps Input to the DLFM module for processing to obtain the Laplace fusion matrix ;
[0058] The feature map Perform learnable linear combination processing on the nodes and use the processing results to generate a dual-channel Laplace matrix, and combine the dual-channel Laplace matrix and the Laplace fusion matrix Input to DGOM module for processing and obtain loss and the set of feature vectors , i is 1 or 2, the feature vector set Input to the second classifier for processing to obtain the prediction result of the classification of the optimal transmission network branch of the dynamic graph .
[0059] It should be noted that the DGOT module is a dynamic graph optimal transmission module, the DGAM module is a dynamic graph generation module, the DLFM module is a dual-channel Laplace fusion module, the DGOM module is a dual-channel graph optimal transmission module, and the LSS is a learnable scale selector.
[0060] Furthermore, the feature map output by the LSS module Feed it into an attention module, the processing process of the attention module is as follows:
[0061] ;
[0062] ;
[0063] in, (·) represents the attention module, (·) represents global pooling, (·) represents the Softmax function, represents the output features of the attention module, Represents a 3D convolution with 8 channels. The output features Feed into three cascaded convolution blocks and fuse them in the form of skip connections to output feature maps , the processing process is as follows:
[0064] ;
[0065] in, Denotes the i-th convolution block, and Cat denotes cascade. Each convolution block has the same structure except for the number of channels, as shown in the following formula:
[0066] ;
[0067] ;
[0068] ;
[0069] in, is a 3D convolution with 3 channels, is a 3D convolution with 5 channels, It is a 3D convolution with 1 channel.
[0070] Will Feed into a standard ViT output feature map , and the feature map , through the classifier (i.e. the first classifier) outputs the final result. The process is as follows:
[0071] ; ;
[0072] in, is the random dropout function, Structure and Same structure, represents a linear layer with m nodes, Indicates the total number of land feature categories, for The classification results of the network branches. In addition, the feature map It is also used as one of the inputs of DGOT to construct the subsequent graph optimal transmission algorithm.
[0073] like Figure 2 As shown in FIG, the present invention includes an AS-ViT network branch based on a learnable scale selector and a DGOT network branch based on dual-channel graph optimal transmission. After the patch sample is fed into AS-ViT, the category prediction result of the branch is output. In addition, the patch sample is fed into DGOT to construct a dual-channel graph sample, and DGOM is used to perform cross-network feature fusion. The output COPT distance (loss ) and the feature vector set are combined with the output of AS-ViT to construct the final overall loss function.
[0074] In some embodiments, the LSS module includes a 3D-CNN encoder, and the patch sample X is input to the 3D-CNN encoder for processing to obtain the scale factor , the scale factor Perform normalization to obtain the normalized scale factor, generate a binary mask matrix based on the scale factor, perform Hadamard product operation on the binary mask matrix and the patch sample X, and then crop to obtain the image block with the center area retained. , the image block After bicubic upsampling, samples with optimal observation scale are obtained , the sample Cascade with the patch sample X to obtain the feature map .
[0075] It should be noted that a learnable scale selector (LSS module) is constructed. The scale selection is to adaptively crop the spatial sampling scale of the input patch sample to the optimal observation scale, and restore it to the original size by bicubic upsampling while retaining the actual spatial observation range. Specifically, the workflow of the LSS module is divided into the following key steps: First, the maximum field of view range (i.e., the initial patch size) is determined based on the spectral spatial characteristics of the input sample. Subsequently, the sample is fed into a lightweight 3D-CNN encoder, which outputs a scale factor The encoder process is as follows:
[0076] ; ;
[0077] in, represents the original input patch sample, denotes a 3D convolutional layer with c channels, BN(·) denotes a batch normalization layer, ReLU(·) denotes a ReLU activation function, and Flatten(·) denotes a tensor flattening operation. Indicates generation 3D-CNN encoder. After that, we perform the following normalization operation:
[0078] ;
[0079] in, Represents the normalized proportional coefficient, S=[ ], B represents the batch size. Through normalization, the network can autonomously learn the desired field of view length. and the length of the initial field of view (i.e. the optimal ratio coefficient between the side lengths of patch samples) ( ∈(0,1]). i represents the i-th scale factor in the current batch. satisfy:
[0080] ;
[0081] Based on the normalized scale factor Value, the network dynamically generates the corresponding binary mask matrix M∈ H×W , where H and W represent the spatial dimensions of the initial patch size. In the binary mask matrix, the center area value is 1, the edge area value is 0, and the side length of the center area is Through the Hadamard product operation, the mask matrix interacts with the original patch, effectively retaining the key feature information at the current optimal scale. The image block after the operation is cropped to remove the zero values at the edge of the image and only retain the image block in the center area. .Will After bicubic upsampling, the sample with the best observation scale is obtained Finally, Cascaded with the initial input sample to obtain the output feature map of the learnable scale selector , the processing process is as follows:
[0082] ;
[0083] ;
[0084] in, (·) means concatenating n tensors into a single n-channel tensor. (·) indicates bicubic upsampling.
[0085] like Figure 3 As shown in the figure, a corresponding mask matrix is constructed based on the scale factor. The mask matrix interacts with the input patch samples to retain the optimal sampling area in the middle. Based on this, cropping is performed to remove the peripheral all-zero areas and bicubic upsampling is used to construct an image patch of the same size as the original patch. Finally, the processed image patch is concatenated with the original patch as the output of the LSS module, which preserves the optimal sampling scale and contrast information with the original patch.
[0086] In some embodiments, the processing flow of the first DGAM module and the second DGAM module is the same, wherein the processing flow of the first DGAM module is: the patch sample X input to the first DGAM module is dimensionally transformed to obtain a matrix , the matrix With learnable weight matrix After matrix multiplication and dimension transformation, the feature map is obtained , the feature map Feed into the GAT module for processing to obtain the feature map , using feature maps Construct a dual-channel adjacency matrix, which includes the modulus adjacency matrix and the spectral angle adjacency matrix. Input the modulus adjacency matrix and the spectral angle adjacency matrix into the single-layer graph convolution layer for convolution processing, and concatenate the processing results to obtain the dual-channel graph features. ;
[0087] The GAT module includes the first linear layer, the second linear layer, the dimension compression module and the LReLU module. The first linear layer performs feature encoding to obtain the feature map , the feature map Perform dimension transformation and replication, concatenate the feature maps obtained after replication in two different dimensions and feed them into the second linear layer for processing, and perform dimension compression and LReLU activation on the processing results to obtain the attention weight factor , the feature map and attention weight factor Multiply and output feature map .
[0088] It should be noted that the feature map During dimension transformation and replication, the replication process is as follows: First, add a dimension before the dimension containing the number of nodes and replicate it multiple times in that dimension, with the number of replications matching the number of nodes. Then, add a dimension after the dimension containing the number of nodes and replicate it in the same manner.
[0089] Furthermore, in the process of constructing the optimal transmission network branch of the dynamic graph, in order to introduce the concept of free scale in the graph, the present invention designs a dual-channel graph attention module (i.e., DGAM module). The DGAM module adopts a dual update strategy when constructing graph data, including updating the graph nodes and updating the graph embedding vector. Specifically, in the DGAM module, a learnable linear combination of nodes is first performed, that is, the input patch sample is first transformed into a matrix after dimension conversion. , and then combine this matrix with the learnable weight matrix Perform matrix multiplication so that Implement adaptive linear combination and perform dimensional transformation to generate feature maps of the spatial scale required for subsequent linear layers The processing process is shown below:
[0090] ;
[0091] Where DT represents the dimension transformation function. So far, the node update is completed through the learnable linear combination. The generated feature map Feed into the GAT module to update the embedding vector. The specific process is as follows: First, Input the first linear layer to encode the features into feature maps , and then perform dimension transformation and replication. After the two feature maps with different dimensions are cascaded, they are fed into the second linear layer and dimensionally compressed and activated, and the attention weight factor is output. Finally, and Multiply output feature map At this point, the update of the embedding vector is completed. The mathematical process of the GAT mechanism is as follows:
[0092] ;
[0093] ; ;
[0094] in, represents the number of nodes encoded in the first linear layer, and Respectively indicate replication in the last two dimensions. represents the compression dimension function, express activation function, A linear layer with 1 node.
[0095] After completing the double update, according to the output characteristics Construct a dual-channel adjacency matrix, which includes the modulus adjacency matrix and the spectral angle adjacency matrix. The algorithm for the two adjacency matrices is as follows:
[0096] ;
[0097] ;
[0098] in, represents the modular adjacency matrix, represents the spectral angle adjacency matrix, represents the feature map The i-th feature sequence in Representation feature map The jth feature sequence in express After completing the construction of the two adjacency matrices, they are respectively input into a single-layer graph convolution layer, and the convolution results are cascaded to obtain the dual-channel graph features finally output by the DGAM module. .
[0099] like Figure 4 As shown in the figure, the input patch samples are fed into the GAT module after undergoing a learnable linear combination, achieving dual updates of the nodes and embedding vectors. Then, the modulus adjacency matrix and spectral angle adjacency matrix of the updated data are calculated and fed into a single-layer graph convolution layer. Finally, the convolution results are cascaded to form a dual-channel graph sample as the output of the DGAM module.
[0100] In some embodiments, the feature map After adding by channel, it is fed into the second DGAM module for processing to obtain the dual-channel image features In the DLFM module, the dual-channel graph features and dual-channel graph features Each channel of each channel constructs a dual-channel adjacency matrix to obtain eight adjacency matrices, and constructs corresponding Laplace matrices for each of the eight adjacency matrices. The two Laplace matrices corresponding to the modulus adjacency matrix of the weighted sum are obtained to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the spectral angle adjacency matrix of the spectral angle are weighted summed to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the modulus adjacency matrix of the weighted sum are obtained to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the spectral angle adjacency matrix of the spectral angle are weighted summed to obtain the Laplace fusion matrix , the Laplacian fusion matrix in the channel dimension and Laplace fusion matrix Cascade to obtain the Laplace fusion matrix , the Laplacian fusion matrix in the channel dimension and Laplace fusion matrix Cascade to obtain the Laplace fusion matrix , the Laplace fusion matrix and Laplace fusion matrix Add together to obtain the Laplace fusion matrix .
[0101] It should be noted that considering the dynamic graph features of different scales can make the input information more complete. In addition to feeding the input patch sample into the first DGAM module, the present invention also feeds the output results of LSS into the second DGAM module after adding them by channel to generate dual-channel graph features. , then, a dual-channel Laplace fusion module (DLFM module) is constructed. The DLFM module is designed to achieve the fusion of patch features and dynamic scale features and construct input suitable for the graph optimal transmission algorithm. The specific process is as follows: First, and A two-channel adjacency matrix is constructed for each channel (eight adjacency matrices in total), and on this basis, the Laplacian matrix corresponding to each adjacency matrix is constructed. The Laplacian matrix construction method is as follows:
[0102] ;
[0103] ;
[0104] in, , and denote the p-th Laplacian matrix, the p-th degree matrix and the p-th adjacency matrix respectively, represents the value of the i-th diagonal element of the p-th degree matrix, is the element value of the i-th row and j-th column of the p-th adjacency matrix. The four Laplacian matrices corresponding to the modular adjacency matrix are summed up according to the weights to obtain and .by For example, the mathematical process is as follows:
[0105] ;
[0106] in, and are all learnable weight parameters, and express The two Laplacian matrices corresponding to the module value adjacency matrix of . Similarly, the four Laplacian matrices corresponding to the spectral angle adjacency matrix are summed up according to the weights to obtain and On this basis, in the channel dimension and Concatenate to obtain the Laplacian fused matrix of the vector module Similarly, in the channel dimension and Cascade to get , and finally, by and Add to obtain the Laplace fusion matrix of the final output of DLFM , the mathematical expression is as follows:
[0107] .
[0108] like Figure 5 As shown in the figure, two dual-channel adjacency matrices from different networks are fed into the dual-channel adjacency matrix generation algorithm, and the corresponding Laplacian matrices are generated based on them. These Laplacian matrices are weighted fused and concatenated according to learnable weights to output the final Laplacian fusion matrix.
[0109] In some embodiments, the DGOM module includes a first COPT module and a second COPT module, wherein the Laplace fusion matrix output by the DLFM module is And the feature map output by the VIT module The Laplace matrix Input to DGOM module, Laplace fusion array Including feature maps and feature maps , the Laplace matrix Including feature maps and feature maps , using the first COPT module to optimize the feature map and feature maps The distance is obtained to obtain the first COPT distance, and the second COPT module is used to optimize the feature map and feature maps The distance is obtained to obtain the second COPT distance, and the first COPT distance and the second COPT distance are adaptively fused to obtain the loss After stacking the first COPT distance and the second COPT distance, we get the feature vector set , i is 1 or 2.
[0110] It should be noted that in order to improve the compatibility between different network feature data, the Laplace fusion matrix output by the DLFM module in the present invention is And the output feature map of the ViT module In order to improve the compatibility between the two, a domain alignment algorithm based on COPT, namely DGOM module, is designed to achieve the progressive fusion of different network features. The DGOM module mainly compares and optimizes the COPT distance of two dual-channel Laplace matrices to make the information between the two similar, thereby completing information interaction. Therefore, the input of DGOM has two parts. The first part is the Laplace fusion matrix output by DLFM. The second part is the output feature map of ViT The Laplace matrix First, the Laplace matrix is fused The construction method is the same as that of Perform a learnable linear combination of nodes. Then, construct its dual-channel adjacency matrix according to the modulus adjacency matrix and the spectral angle adjacency matrix. Finally, calculate the corresponding Laplacian matrix according to the degree matrix of these two adjacency matrices, and concatenate these two Laplacian matrices into a dual-channel Laplacian matrix. .
[0111] In the DGOM module, as input and Each has two channels. Therefore, the present invention inputs the corresponding channels of the two inputs into two different COPT modules. The COPT module is an existing graph optimal transmission algorithm that can optimize the distance between two graphs. The present invention adaptively fuses the distance output by the two COPT modules as the loss of DSGO (DSGO is a hyperspectral image classification model). , thereby achieving indirect interaction between different features. The mathematical expression of this process is as follows:
[0112] ;
[0113] in, and Represent the average value and absolute value respectively. represents the COPT algorithm, represents the i-th learnable weight parameter, and is the Laplace fusion matrix The feature maps corresponding to the two channels of and is a dual-channel Laplacian matrix The feature maps corresponding to the two channels of
[0114] In addition, the present invention outputs the feature vector set generated by each COPT module during transmission, and stacks the feature vector set obtained by stacking them. Feeding a classifier (i.e. the second classifier), the feature vector set Calculated by the following formula:
[0115] ;
[0116] in, represents the feature vector set of the i-th COPT module, represents the learnable transmission plan matrix. Represents a vector set of calculated feature matrices. The second classifier The mathematical expression is as follows:
[0117]
[0118] ;
[0119] in, represents a 2D convolution operation with one layer of i channels, Represents a cascade operation between two network layers. Represents a classifier The classification result of the output is used to construct the final loss.
[0120] like Figure 6 As shown in the figure, first, the dual-channel Laplacian matrices from two different networks are fed into the DGOM module. Then, the two sets of Laplacian matrices for the corresponding channels are fed into different COPT modules, which output two COPT distances and two sets of feature vectors. Finally, the two COPT distances are adaptively fused and used as the network loss. Simultaneously, the two sets of feature vectors are concatenated and fed into the second classifier to predict the classification result.
[0121] In some embodiments, in step S3, the overall loss function for:
[0122] ;
[0123] ;
[0124] ;
[0125] in, is the loss of the classification result produced by the first classifier, is the loss of the classification result produced by the second classifier, For the prediction results and prediction results The corresponding real feature labels, is the loss, i is the index of cn, and cn is the total number of categories.
[0126] It should be noted that after constructing the total loss function, the prediction results of the forward propagation are used to optimize the total loss function, thus forming a complete DSGO.
[0127] Feed the training samples of the training set into DSGO and iteratively optimize the total loss function. The number of training iterations is 150 (open early stopping), and the optimizer selects Adam optimizer ( =0.9, 0.999 ), with a batch size of 256 and an initial learning rate of 0.001 (reduced to 0.1 every 30 epochs). While optimizing the total loss function, the learnable parameters of the hyperspectral image classification model are updated, resulting in a gradual improvement in the classification accuracy of the training set. During each training epoch, validation samples from the validation set are fed into the hyperspectral image classification model to verify its classification performance on untrained samples, and hyperparameters are adjusted accordingly. Training is terminated when the validation set classification accuracy reaches the highest. The test set samples are fed into the trained HPFN, and the output is the final predicted classification result. Finally, the overall accuracy, average accuracy, and Kappa are calculated based on the classification results of all pixels. A classification map is generated to comprehensively evaluate the classification performance of the model.
[0128] like Figure 7 and Figure 8 As shown, (a)-(l) are specifically: (a) SVM; (b) A 2 S 2 K; (c) MVAHN; (d) MRViT; (e) DSGSF; (f) FM; (g) AMGCFN; (h) HKAN; (i) MFFN; (j) MLP; (k) DSGO (the present invention); (l) Ground Truth. The present invention can achieve the best classification results on both the HS dataset and the SV dataset. Figure 9 and Figure 10 As shown in the figure, the present invention can achieve the best average accuracy, overall accuracy and Kappa on both the HS dataset and the SV dataset.
[0129] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. This is not limited herein.
[0130] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A hyperspectral image classification method based on dynamic scaling and graph optimal transmission, characterized by: The specific steps include: S1: Obtain a labeled hyperspectral image dataset and use it to construct a training set; S2: constructing a hyperspectral image classification model, wherein the hyperspectral image classification model includes an adaptive scale ViT network branch and a dynamic graph optimal transmission network branch; The adaptive scale ViT network branch is used to dynamically adjust the sampling scale of input data according to the characteristics of different ground object categories, and the dynamic graph optimal transmission network branch is used to achieve the interaction and fusion of different information; In step S2, the adaptive scale ViT network branch includes an LSS module, an attention module, a first 3D convolution block, a second 3D convolution block, a third 3D convolution block, a VIT module and a first classifier, wherein the patch sample X in the training set is input into the LSS module for processing to obtain a feature map , the feature map Input to the attention module for processing to obtain the feature map , the feature map Input to the first 3D convolution block for processing to obtain feature map A1, feature map A1 is processed by the second and third convolution blocks to obtain feature map A2, and the processing results of the first 3D convolution block, the second 3D convolution block and the third 3D convolution block are cascaded to obtain feature map , the feature map Input to the VIT module for processing to obtain the feature map , the feature map Input to the first classifier for processing to obtain the classification prediction results of the adaptive scale ViT network branch ; The dynamic graph optimal transmission network branch includes the first DGAM module, the second DGAM module, the DLFM module, the DGOM module and the second classifier. The patch sample X in the training set is input into the first DGAM module for processing to obtain the feature map , the feature map Input to the second DGAM module for processing to obtain the feature map , the feature map and feature maps Input to the DLFM module for processing to obtain the Laplace fusion matrix ; The feature map Perform learnable linear combination processing on the nodes and use the processing results to generate a dual-channel Laplace matrix, and combine the dual-channel Laplace matrix and the Laplace fusion matrix Input to DGOM module for processing and obtain loss and the set of feature vectors , i is 1 or 2, the feature vector set Input to the second classifier for processing to obtain the prediction result of the classification of the optimal transmission network branch of the dynamic graph ; S3: constructing an overall loss function, and training the hyperspectral image classification model using the training set and the overall loss function to obtain a trained hyperspectral image classification model; S4: Input the hyperspectral image to be classified into the trained hyperspectral image classification model for classification to obtain the classification result.
2. The hyperspectral image classification method based on dynamic scaling and graph optimal transmission according to claim 1 is characterized in that: In step S1, the labeled pixels in the hyperspectral image dataset are extracted, all the labeled pixel sets are divided into categories, and 100 pixels are randomly selected from the pixel set of each category. For categories with less than 100 pixels, half of the total number of pixels are selected. Principal component analysis is performed on each hyperspectral image contained in the hyperspectral image dataset, and a patch sample is constructed based on each selected pixel point. The patch sample is composed of the pixel point and the 8 neighborhood pixels of the current pixel point. The patch sample size is 9×9×b, where 9×9 represents the size of the neighborhood range and b is the number of principal components retained after dimensionality reduction. The set of corresponding patch samples constructed using each selected pixel point is used as the training set.
3. The hyperspectral image classification method based on dynamic scaling and graph optimal transmission according to claim 1 is characterized in that: The LSS module includes a 3D-CNN encoder, which inputs the patch sample X into the 3D-CNN encoder for processing to obtain the scale factor , the scale factor Perform normalization to obtain the normalized scale factor, generate a binary mask matrix based on the scale factor, perform Hadamard product operation on the binary mask matrix and the patch sample X, and then crop to obtain the image block with the center area retained. , the image block After bicubic upsampling, samples with optimal observation scale are obtained , the sample Cascade with the patch sample X to obtain the feature map .
4. The hyperspectral image classification method based on dynamic scaling and graph optimal transmission according to claim 1, characterized in that: The processing flow of the first DGAM module and the second DGAM module is the same, wherein the processing flow of the first DGAM module is: the patch sample X input to the first DGAM module is dimensionally transformed to obtain the matrix , the matrix With learnable weight matrix After matrix multiplication and dimension transformation, the feature map is obtained , the feature map Feed into the GAT module for processing to obtain the feature map , using feature maps Construct a dual-channel adjacency matrix, which includes the modulus adjacency matrix and the spectral angle adjacency matrix. Input the modulus adjacency matrix and the spectral angle adjacency matrix into the single-layer graph convolution layer for convolution processing, and concatenate the processing results to obtain the dual-channel graph features. ; The GAT module includes the first linear layer, the second linear layer, the dimension compression module and the LReLU module. The first linear layer performs feature encoding to obtain the feature map , the feature map Perform dimension transformation and replication, concatenate the feature maps obtained after replication in two different dimensions and feed them into the second linear layer for processing, and perform dimension compression and LReLU activation on the processing results to obtain the attention weight factor , the feature map and attention weight factor Multiply and output feature map .
5. The hyperspectral image classification method based on dynamic scaling and graph optimal transmission according to claim 1, characterized in that: The feature map After adding by channel, it is fed into the second DGAM module for processing to obtain the dual-channel image features In the DLFM module, the dual-channel graph features and dual-channel graph features Each channel of each channel constructs a dual-channel adjacency matrix to obtain eight adjacency matrices, and constructs corresponding Laplace matrices for each of the eight adjacency matrices. The two Laplace matrices corresponding to the modulus adjacency matrix of the weighted sum are obtained to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the spectral angle adjacency matrix of the spectral angle are weighted summed to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the modulus adjacency matrix of the weighted sum are obtained to obtain the Laplace fusion matrix , the dual-channel graph features The two Laplace matrices corresponding to the spectral angle adjacency matrix of the spectral angle are weighted summed to obtain the Laplace fusion matrix , the Laplacian fusion matrix in the channel dimension and Laplace fusion matrix Cascade to obtain the Laplace fusion matrix , the Laplacian fusion matrix in the channel dimension and Laplace fusion matrix Cascade to obtain the Laplace fusion matrix , the Laplace fusion matrix and Laplace fusion matrix Add together to obtain the Laplace fusion matrix .
6. The hyperspectral image classification method based on dynamic scaling and graph optimal transmission according to claim 1, characterized in that: The DGOM module includes the first COPT module and the second COPT module, wherein the Laplace fusion matrix output by the DLFM module is And the feature map output by the VIT module The Laplace matrix Input to DGOM module, Laplace fusion array Including feature maps and feature maps , the Laplace matrix Including feature maps and feature maps , using the first COPT module to optimize the feature map and feature maps The distance is obtained to obtain the first COPT distance, and the second COPT module is used to optimize the feature map and feature maps The distance is obtained to obtain the second COPT distance, and the first COPT distance and the second COPT distance are adaptively fused to obtain the loss After stacking the first COPT distance and the second COPT distance, we get the feature vector set , i is 1 or 2.
7. The hyperspectral image classification method based on dynamic scaling and graph optimal transmission according to claim 1, characterized in that: In step S3, the overall loss function for: ; ; ; in, is the loss of the classification result produced by the first classifier, is the loss of the classification result produced by the second classifier, For the prediction results and prediction results The corresponding real feature labels, is the loss, i is the index of cn, and cn is the total number of categories.
Citation Information
Patent Citations
Double-branch hyperspectral image classification method based on graph convolutional neural network and attention mechanism
CN118135306A