Construction method of adaptive gated spectrum-space-graph collaborative fusion network
By constructing an adaptive gating spectrum-space-graph collaborative fusion network, the problems of insufficient fusion of hollow spectral features and spectral redundancy in hyperspectral image classification are solved, efficient feature selection and dynamic coordination are achieved, and classification accuracy and computing efficiency are improved.
Patent Information
- Application Number
- CN202510752767.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-12
AI Technical Summary
The existing hyperspectral image classification methods have the challenges of insufficient fusion of null spectral features, spectral redundancy problems and inefficient computing efficiency. Traditional methods are difficult to effectively model the dynamic interaction between multi-scale spatial structures and discriminant spectral patterns, and have high computational costs or insufficient generalization capabilities.
Adaptive gating spectrum-space-graph collaborative fusion network (SGCFN) is constructed, and superpixel-level and pixel-level spatial features are extracted through improved graph attention network and multi-scale proxy self-attention mechanism. Combined with the pooling-induced spectrogram attention module and the adaptive gating fusion module, adaptive feature selection and cross-modal dynamic coordination are realized.
It effectively improves the accuracy of geographic classification of hyperspectral images, solves complex space-spectral interaction modeling and spectral redundancy problems, and reduces computational costs, improves the generalization ability and classification robustness of the model.
Smart Images

Figure CN120472243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image classification, and in particular to a method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network. Background Art
[0002] Hyperspectral imagery (HSI) provides rich information for precise identification of ground objects thanks to its "spectral fingerprint" characteristics of hundreds of continuous spectral bands. However, existing HSI classification methods face the following core challenges:
[0003] Insufficient fusion of spatial-spectral features: Traditional methods find it difficult to effectively model the dynamic interaction between multi-scale spatial structures and discriminative spectral patterns, resulting in limited classification accuracy in complex scenarios.
[0004] Spectral redundancy problem: A large number of redundant bands are not effectively suppressed, and the extraction of discriminative spectral features lacks an adaptive mechanism.
[0005] There is a contradiction between computational efficiency and model complexity: Transformer-based methods have high computational costs, while graph convolutional network (GCN)-based methods rely on manually preset graph structures and have insufficient generalization capabilities.
[0006] While existing hybrid models attempt to integrate the strengths of different architectures, stacking multimodal modules leads to a surge in the number of parameters, and fusion strategies often rely on simple concatenation or weighted summation, lacking dynamic weight calibration mechanisms. Therefore, an efficient spatial-spectral feature fusion framework is urgently needed to achieve adaptive feature selection and dynamic cross-modal collaboration. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network to solve the problems raised in the above background technology.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network, comprising the following steps:
[0009] Step 1: Build the overall network architecture: We propose a hybrid two-stream fusion model, SGCFN, which extracts superpixel-level and pixel-level spatial features through an improved graph attention network and a multi-scale proxy self-attention mechanism. A pooling-induced spectral graph attention module mines discriminative spectral information, and an adaptive gated fusion module (AGFM) integrates the two-stream features to effectively improve the accuracy of hyperspectral imagery object classification.
[0010] Step 2: Spatial Branching - Hierarchical Spatial Feature Modeling: To address the static structural limitations of graph networks and the high computational complexity of the Transformer method, we design a hierarchical multi-adjacency graph attention module (MAGAT) and a lightweight multi-scale agent-attention module (MSAAT). The MAGAT uses dual-hop graph attention to dynamically capture local geometric details and regional semantic dependencies between superpixels, and combines it with a channel splitting strategy to enhance feature diversity. The MSAAT integrates multi-scale depthwise separable convolution with an agent self-attention mechanism to reduce computational costs while maintaining global context modeling capabilities.
[0011] Step 3: Spectral branching - adaptive spectral feature optimization: We propose a pooling-induced spectral graph attention module (PISGAT). This module generates channel-level node embeddings through multi-scale spatial pooling and dynamically constructs an adaptive spectral graph structure, effectively suppressing spectral redundancy and enhancing discriminative features.
[0012] Step 4. Design of Adaptive Gated Fusion Module (AGFM): Design a pixel-level learnable adaptive gated fusion module (AGFM). Dynamically adjust the feature weights of spatial and spectral branches based on the scene context, enhance spatial structural features in spectrally complex regions to distinguish heterogeneous boundaries, and highlight spectral identification features in spatially fragmented regions to improve intra-class consistency.
[0013] Preferably, the step 1 is specifically as follows: the hyperspectral image data is formalized into an input cube , where H, W, and B represent the image height, width, and number of spectral bands, respectively. To exploit spatial regularity, the original image is segmented into N superpixels by simple linear iterative clustering (SLIC), generating three structured inputs: a superpixel feature matrix containing C-dimensional node embeddings, , a multi-hop adjacency matrix defining the hierarchical connectivity pattern between superpixels , and the projection matrix describing the pixel-superpixel mapping relationship ;
[0014] The hybrid two-stream fusion model SGCFN consists of three collaborative components: a spatial branch for superpixel-level and pixel-level spatial information modeling, a spectral branch for channel-level discriminative feature learning, and an adaptive gated fusion module for cross-modal feature integration.
[0015] Preferably, first, the spatial branch modeling processes the superpixel graph through a multi-adjacency graph attention module (MAGAT) , where the parallel graph attention network is based on different adjacency matrices Capture hierarchical spatial correlations; after decoding by the projection matrix Q, these spatial features are further enhanced by MSAAT, fusing multi-scale depth-separable convolution with lightweight proxy self-attention mechanism to obtain pixel-level spatial information; secondly, the spectral branch uses 1×1 convolution to convert the original hyperspectral data ( ) is mapped to the latent space, and then the channel relationship is optimized through PISGAT: this process uses multi-scale mean pooling to generate spatial descriptors to dynamically construct the channel map, and implements graph convolution propagation to suppress redundant bands; finally, the adaptive gated fusion module bridges the two branches through AGFM, uses learnable convolutional gates to generate pixel-level fusion weights, and selectively fuses features according to contextual relevance; the fused features are processed by a feedforward network with a hidden layer dimension of C / 2 and the classification results are output. , where c is the number of categories.
[0016] Preferably, in the step 2, the hierarchical multi-adjacency graph attention module MAGAT is specifically as follows: the standard graph attention network GAT processes graph structure data through a dynamic attention mechanism, aiming to adaptively capture the dependencies between nodes without relying on a predefined adjacency matrix; the standard GAT is improved so that the network can adjust the attention weight according to the adjacency matrix; specifically, given the input superpixel feature , GAT first passes the learnable weight matrix Project features into latent space:
[0017]
[0018] in and ; For each superpixel i, it is compared with the neighboring superpixels The attention coefficient between is calculated as follows:
[0019]
[0020] Among them, concat(∙) represents the vector concatenation operation, and a is the shared attention vector; these attention coefficients are filtered by the positive values of the corresponding positions in the adjacency matrix (A) and normalized by the SoftMax function to generate the attention weight reflecting the importance of superpixel j to i. :
[0021]
[0022] Then, the attention weight is used to aggregate neighborhood features, and its expression is:
[0023]
[0024] Where σ(∙) represents the activation function of the stabilization process. To enhance the diversity of feature representation, the multi-head attention mechanism uses K independent attention heads to operate in parallel, and then splices the output results of each head. Its mathematical expression is:
[0025]
[0026] The outputs of each head are spliced together to form the final feature ;
[0027] This mechanism dynamically infers the association between superpixels through data, thereby effectively capturing contextual dependencies in irregular spatial structures. However, the standard GAT is limited by the homogeneous adjacency mechanism, which imposes the same interaction range on all nodes. This construction method cannot fully characterize the hierarchical spatial associations of hyperspectral superpixels: neighboring regions need to preserve local patterns, while distant regions benefit from extended context aggregation. To solve this problem, the proposed MAGAT introduces a two-hop hierarchical aggregation architecture. The single-hop GAT layer captures direct neighborhood relationships to preserve geometric details. Its output features are spliced with deep group features and then processed by the two-hop GAT layer to fuse regional semantic information.
[0028] Given the input superpixel features , this module first divides the features into two complementary groups by channel splitting, and its operation is expressed as:
[0029]
[0030] Then, By The guided GAT layer is processed; the process is expressed as:
[0031]
[0032] Where LeakyReLU(∙) represents the LeakyReLU activation function, LayerNorm(∙) is the layer normalization operation, and GAT(∙) represents the graph attention network; then, the intermediate features and Stitching to form enhanced features , and through the linear layer and by Guided secondary GAT layer processing;
[0033]
[0034]
[0035] Hierarchical fusion through learnable coefficients (The coefficient is initialized to 0.5) to achieve the combination of multi-hop features, and its operation is expressed as:
[0036] .
[0037] Preferably, in step 2, the lightweight multi-scale proxy attention module MSAAT is specifically: the lightweight multi-scale proxy attention module MSAAT innovatively synergistically integrates lightweight proxy guided attention and multi-scale depth-separable convolution;
[0038] Based on the hierarchical features extracted by MAGAT, MSAAT enhances spatial feature learning through a local-global hybrid interaction mechanism; given the input features mapped by the projection matrix Q in the superpixel space, , ,in( ), first execute the query Q, key K, value V) shadow:
[0039]
[0040] in , , is a learnable linear transformation matrix; in order to reduce the computational complexity while maintaining global interaction, adaptive pooling and shaping operations are performed to transform Reshape the generation of proxy tokens in the form of compact spatial grids in( ):
[0041] The proxy token acts as a lightweight representation of the global context; it then generates a proxy value by interacting with the key and value:
[0042] In the formula represents the Dropout layer, and Softmax(∙) is the SoftMax operation; then, the query and proxy values are attention-corrected through the proxy token:
[0043] To supplement the local detail features in the attention mechanism, parallel multi-scale depth-wise separable convolution is used to process the numerical terms:
[0044] Finally, the fusion of multi-scale features is achieved through residual connection:
[0045] In the formula Represents the multi-scale feature integration result.
[0046] Preferably, the step three is as follows: the proposed PISGAT enhances the hyperspectral feature representation by fusing adaptive spatial pooling and dynamic graph convolution mechanism; given an input feature map , first compress the spatial dimension by adaptive average pooling:
[0047] Where S is the spatial resolution after downsampling; the pooled features are projected into the latent space through 1×1 convolution:
[0048] Where r is the channel reduction rate; then Flatten to node embedding features ( ) to process the graph structure; design the graph channel attention GCAT module to generate channel weights for the input features, by introducing the static unit matrix To maintain the intrinsic self-connection relationship between nodes and ensure the stability of the basic topology; to inject input-specific adaptability, the pooled node embedding After 1×1 convolution and SoftMax normalization, a data-dependent adjustment matrix is generated:
[0049]
[0050] This matrix captures the relationship between nodes under the input feature conditions; introduces the learnable matrix (The initial value is close to zero matrix) to supplement the above components, so that the graph topology is gradually optimized during the training process to achieve task adaptability adjustment; the final adjacency matrix is composed of the combination of each component:
[0051] in Represents element-by-element multiplication; spectral graph convolution updates node features through the following formula:
[0052] In the formula is the learnable weight matrix, σ(∙) represents the ReLU activation function; the updated embedding features After 1×1 convolution expansion, the channel attention weight is generated by global average pooling:
[0053] These weights recalibrate the input features through channel-wise multiplication:
[0054] Where ⊗ represents element-wise multiplication under the broadcast mechanism.
[0055] Preferably, the step 4 is specifically as follows: dynamically calibrating cross-branch feature contributions through learnable spatial-spectral attention; given spatial features and spectral characteristics , first concatenate the two along the channel dimension and generate gating weights through a hierarchical convolutional layer:
[0056] Where the gated tensor Contains two spatial attention maps To achieve weighted fusion:
[0057] where ⊙ represents element-wise multiplication.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] This paper proposes a Spectral-Spatial-Graph Cooperative Fusion Network (SGCFN), a dual-stream architecture that addresses the challenges of complex spatial-spectral interaction modeling, spectral redundancy, and computational inefficiency in hyperspectral image classification. By integrating hierarchical spatial feature learning, adaptive spectral optimization, and a dynamic cross-modal fusion mechanism, the network effectively combines the complementary strengths of graph attention, agent self-attention, and data-driven spectral modeling. Experimental results on three benchmark datasets validate the advanced performance of the Spectral-Spatial-Graph Cooperative Fusion Network (SGCFN), and ablation studies confirm the irreplaceable role of each module in enhancing classification robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 This is the SGCFN architecture diagram proposed by the present invention;
[0061] Figure 2 is a diagram of the GAT adjacency architecture of the present invention;
[0062] Figure 3 This is the MAGAT architecture diagram proposed by the present invention;
[0063] Figure 4 This is the MSAAT architecture diagram proposed by the present invention;
[0064] Figure 5 This is the PISGAT architecture diagram proposed by the present invention;
[0065] Figure 6 This is the AGFM architecture diagram proposed by the present invention;
[0066] Figure 7 It is a full-factor classification diagram of the IP dataset of the present invention;
[0067] In the figure, (a) DBDA; (b) PCIA; (c) SSGCA; (d) MSCA; (e) DBCT; (f) TECC; (g) DGFNet; (h) AMGCF; (i) MRCAG; (j) PCCGC; (k) the method of the present invention; (l) false color image (60, 30, 10);
[0068] Figure 8 It is the full-factor classification map of the UP dataset of the present invention;
[0069] In the figure, (a) DBDA; (b) PCIA; (c) SSGCA; (d) MSCA; (e) DBCT; (f) TECC; (g) DGFNet; (h) AMGCF; (i) MRCAG; (j) PCCGC; (k) this method; (l) false color image (50, 30, 10);
[0070] Figure 9 is the full-factor classification diagram of the MUUFL dataset of the present invention; (a) DBDA; (b) PCIA; (c) SSGCA; (d) MSCA; (e) DBCT; (f) TECC; (g) DGFNet; (h) AMGCF; (i) MRCAG; (j) PCCGC; (k) the method of the present invention; (l) false color image (31, 14, 2);
[0071] Figure 10 is the influence diagram of the MAGAT, MSAAT, PISGAT and AGFS modules of the present invention;
[0072] Figure 11 is a classification accuracy graph under different sample ratios of the present invention;
[0073] In the figure, (a) Indianapolis dataset; (b) University of Pavia dataset; (c) University of Mississippi multispectral dataset;
[0074] Figure 12 This is the label map in Table 1 of the embodiment, which is derived from Ground truth;
[0075] Figure 13 This is the label map in Table 2 of the embodiment, which is derived from Ground truth;
[0076] Figure 14 This is the label map in Table 3 in the embodiment, which is derived from Ground truth. DETAILED DESCRIPTION
[0077] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0078] See also Figure 1-11 The present invention provides a method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network, comprising the following steps:
[0079] Step 1: Build the overall network architecture: A hybrid two-stream fusion model, SGCFN, is proposed. It uses an improved graph attention network and a multi-scale proxy self-attention mechanism to extract superpixel-level and pixel-level spatial features, respectively. It then uses a pooling-induced spectral graph attention module to mine discriminative spectral information. Finally, an adaptive gated fusion module is used to integrate two-stream features, effectively improving the accuracy of hyperspectral imagery object classification.
[0080] Step 2: Spatial Branching - Hierarchical Spatial Feature Modeling: To address the static structural limitations of graph networks and the high computational complexity of the Transformer method, we design a hierarchical multi-adjacency graph attention module (MAGAT) and a lightweight multi-scale proxy attention module (MSAAT). The hierarchical multi-adjacency graph attention module (MAGAT) uses dual-hop graph attention to dynamically capture local geometric details and regional semantic dependencies between superpixels, and combines it with a channel splitting strategy to enhance feature diversity. The lightweight multi-scale proxy attention module (MSAAT) integrates multi-scale depthwise separable convolution with a proxy self-attention mechanism to reduce computational costs while maintaining global context modeling capabilities.
[0081] Step 3: Spectral branching - adaptive spectral feature optimization: We propose a pooling-induced spectral graph attention module (PISGAT), which generates channel-level node embeddings through multi-scale spatial pooling and dynamically constructs an adaptive spectral graph structure, effectively suppressing spectral redundancy and enhancing discriminative features.
[0082] Step 4. Design of Adaptive Gated Fusion Module (AGFM): Design a pixel-level learnable adaptive gated fusion module (AGFM). Dynamically adjust the feature weights of spatial and spectral branches based on the scene context, enhance spatial structural features in spectrally complex regions to distinguish heterogeneous boundaries, and highlight spectral identification features in spatially fragmented regions to improve intra-class consistency.
[0083] The above step 1 is specifically as follows: the hyperspectral image data is formalized into an input cube , where H, W, and B represent the image height, width, and number of spectral bands, respectively. To exploit spatial regularity, the original image is segmented into N superpixels by simple linear iterative clustering (SLIC), generating three structured inputs: a superpixel feature matrix containing C-dimensional node embeddings, , a multi-hop adjacency matrix defining the hierarchical connectivity pattern between superpixels , and the projection matrix describing the pixel-superpixel mapping relationship ;
[0084] The hybrid two-stream fusion model SGCFN consists of three collaborative components: a spatial branch for superpixel-level and pixel-level spatial information modeling, a spectral branch for channel-level discriminative feature learning, and an adaptive gated fusion module for cross-modal feature integration.
[0085] First, spatial branch modeling processes superpixel maps through MAGAT , where the parallel graph attention network is based on different adjacency matrices Capture hierarchical spatial correlations; after decoding by the projection matrix Q, these spatial features are further enhanced by MSAAT, fusing multi-scale depth-separable convolution with lightweight proxy self-attention mechanism to obtain pixel-level spatial information; secondly, the spectral branch uses 1×1 convolution to convert the original hyperspectral data ( ) is mapped to the latent space, and then the channel relationship is optimized through PISGAT: this process uses multi-scale mean pooling to generate spatial descriptors to dynamically construct the channel map, and implements graph convolution propagation to suppress redundant bands; finally, the adaptive gated fusion module bridges the two branches through AGFM, uses learnable convolutional gates to generate pixel-level fusion weights, and selectively fuses features according to contextual relevance; the fused features are processed by a feedforward network with a hidden layer dimension of C / 2 and the classification results are output. , where c is the number of categories.
[0086] In the above step 2, the hierarchical multi-adjacency graph attention module MAGAT is specifically as follows: the standard graph attention network GAT processes graph structure data through a dynamic attention mechanism, aiming to adaptively capture the dependencies between nodes without relying on a predefined adjacency matrix; the standard GAT is improved so that the network can adjust the attention weight according to the adjacency matrix (the improved adjacency GAT architecture is shown in Figure 2). Figure 2 Specifically, given the input superpixel feature , GAT first passes the learnable weight matrix Project features into latent space:
[0087]
[0088] in and ; For each superpixel i, it is compared with the neighboring superpixels The attention coefficient between is calculated as follows:
[0089]
[0090] Among them, concat(∙) represents the vector concatenation operation, and a is the shared attention vector; these attention coefficients are filtered by the positive values of the corresponding positions in the adjacency matrix (A) and normalized by the SoftMax function to generate the attention weight reflecting the importance of superpixel j to i. :
[0091]
[0092] Then, the attention weight is used to aggregate neighborhood features, and its expression is:
[0093]
[0094] Where σ(∙) represents the activation function of the stabilization process. To enhance the diversity of feature representation, the multi-head attention mechanism uses K independent attention heads to operate in parallel, and then splices the output results of each head. Its mathematical expression is:
[0095]
[0096] The outputs of each head are spliced together to form the final feature ;
[0097] This mechanism dynamically infers the association between superpixels through data, thereby effectively capturing contextual dependencies in irregular spatial structures; however, the standard GAT is limited by the homogeneous adjacency mechanism and imposes the same interaction range on all nodes; this construction method is difficult to fully characterize the hierarchical spatial associations of hyperspectral superpixels: neighboring regions need to retain local patterns, while distant regions benefit from extended context aggregation; to solve this problem, the proposed MAGAT introduces a double-hop hierarchical aggregation architecture; the single-hop GAT layer captures the direct neighborhood relationship to retain geometric details, and its output features are spliced with the deep group features and then processed by the double-hop GAT layer to fuse regional semantic information (the specific architecture is as follows: Figure 3 shown);
[0098] Given the input superpixel features , this module first divides the features into two complementary groups by channel splitting, and its operation is expressed as:
[0099]
[0100] Then, By The guided GAT layer is processed; the process is expressed as:
[0101]
[0102] Where LeakyReLU(∙) represents the LeakyReLU activation function, LayerNorm(∙) is the layer normalization operation, and GAT(∙) represents the graph attention network; then, the intermediate features and Stitching to form enhanced features , and through the linear layer and by Guided secondary GAT layer processing;
[0103]
[0104]
[0105] Hierarchical fusion through learnable coefficients (The coefficient is initialized to 0.5) to achieve the combination of multi-hop features, and its operation is expressed as:
[0106] .
[0107] In the above step 2, the lightweight multi-scale agent attention module MSAAT is specifically as follows: The lightweight multi-scale agent attention module MSAAT innovatively integrates lightweight agent guided attention with multi-scale depth-separable convolution (the specific architecture is as follows Figure 4 shown);
[0108] Based on the hierarchical features extracted by MAGAT, MSAAT enhances spatial feature learning through a local-global hybrid interaction mechanism; given the input features mapped by the projection matrix Q in the superpixel space, , ,in( ), first execute the query Q, key K, value V) shadow:
[0109]
[0110] in , , is a learnable linear transformation matrix; in order to reduce the computational complexity while maintaining global interaction, adaptive pooling and shaping operations are performed to transform Reshape the generation of proxy tokens in the form of compact spatial grids in( ):
[0111] The proxy token acts as a lightweight representation of the global context; it then generates a proxy value by interacting with the key and value:
[0112] In the formula represents the Dropout layer, and Softmax(∙) is the SoftMax operation; then, the query and proxy values are attention-corrected through the proxy token:
[0113] To supplement the local detail features in the attention mechanism, parallel multi-scale depth-wise separable convolution is used to process the numerical terms:
[0114] Finally, the fusion of multi-scale features is achieved through residual connection:
[0115] In the formula Represents the multi-scale feature integration result.
[0116] The above step 3 is as follows: the proposed PISGAT enhances the hyperspectral feature representation by fusing adaptive spatial pooling and dynamic graph convolution mechanism (the specific architecture is as follows Figure 5 As shown); given the input feature map , first compress the spatial dimension by adaptive average pooling:
[0117] Where S is the spatial resolution after downsampling; the pooled features are projected into the latent space through 1×1 convolution:
[0118] Where r is the channel reduction rate; then Flatten to node embedding features ( ) to process the graph structure; design the graph channel attention GCAT module to generate channel weights for the input features, by introducing the static unit matrix To maintain the intrinsic self-connection relationship between nodes and ensure the stability of the basic topology; to inject input-specific adaptability, the pooled node embedding After 1×1 convolution and SoftMax normalization, a data-dependent adjustment matrix is generated:
[0119]
[0120] This matrix captures the relationship between nodes under the input feature conditions; introduces the learnable matrix (The initial value is close to zero matrix) to supplement the above components, so that the graph topology is gradually optimized during the training process to achieve task adaptability adjustment; the final adjacency matrix is composed of the combination of each component:
[0121] in Represents element-by-element multiplication; spectral graph convolution updates node features through the following formula:
[0122] In the formula is the learnable weight matrix, σ(∙) represents the ReLU activation function; the updated embedding features After 1×1 convolution expansion, the channel attention weight is generated by global average pooling:
[0123] These weights recalibrate the input features through channel-wise multiplication:
[0124] Where ⊗ represents element-wise multiplication under the broadcast mechanism.
[0125] The above step 4 is specifically: dynamically calibrate cross-branch feature contributions through learnable spatial-spectral attention (specific architecture as follows Figure 6 Given spatial features and spectral characteristics , first concatenate the two along the channel dimension and generate gating weights through a hierarchical convolutional layer:
[0126] Where the gated tensor Contains two spatial attention maps To achieve weighted fusion:
[0127] where ⊙ represents element-wise multiplication.
[0128] Example:
[0129] Data Description
[0130] To validate the effectiveness of the proposed network, we tested it using three publicly available hyperspectral datasets: the Indiana Pine Forest Dataset (IP), the University of Pavia Dataset (UP), and the University of Mississippi Urban Land Cover Dataset (MUUFL). The basic characteristics of these three datasets are described below (detailed ground truth maps, class labels, legends, and sample sizes are shown in Tables 1-3):
[0131] The IP dataset was collected by the Airborne Visible / Infrared Imaging Spectrometer (AVIRIS) over agricultural areas in northwestern Indiana, USA. The hyperspectral image is 145 × 145 pixels in size. After removing 20 noise-interfering bands, 200 valid spectral bands (covering the 400-2500 nm spectral range) are retained. The spatial resolution is approximately 20 meters per pixel and includes 16 typical ground object categories.
[0132] The UP dataset was acquired by the Reflecting Optical System Imaging Spectrometer (ROSIS) in Pavia, Italy, in 2003. The image size is 610 × 340 pixels, with a spatial resolution of 1.3 meters per pixel. After removing 12 noise bands, 103 spectral channels (430–860 nm wavelength range) are retained, covering nine urban land cover types.
[0133] The MUUFL dataset was collected in 2010 using the ITRES CASI-1500 sensor over the University of Southern Mississippi campus. The imagery consists of 325 × 220 pixels with a high spatial resolution of 0.54–1.0 meters. It contains 64 spectral bands ranging from 375–1050 nm and is annotated with 11 fine-grained urban land cover classes.
[0134] Table 1: Real annotations, categories, legends, and sample numbers of the IP dataset
[0135]
[0136] Table 2: Real annotations, categories, legends, and sample numbers of the UP dataset
[0137]
[0138] Table 3: Ground truth data, category division, legend and sample number statistics of the MUUFL dataset
[0139]
[0140] Experimental setup and evaluation metrics
[0141] In order to comprehensively evaluate the performance of the proposed method, we conducted comparative experiments with ten cutting-edge methods. These methods are divided into four categories according to their technical paradigms:
[0142] Dual-branch hybrid network: DBDA and PCIA use parallel spatial-spectral convolution and self-attention branches to fuse local context and global dependency information through feature splicing.
[0143] Spatial-spectral attention model: SSGCA combines spatial and spectral attention to extract discriminative features; MSCA
[53] enhances the multi-scale context modeling capability through the combined structure of pyramid convolution and cross-dimensional attention mechanism, and achieves robust capture of hierarchical spatial-spectral interaction features.
[0144] Convolution-Transformer hybrid model: DBCT jointly models local spatial patterns and global long-range dependencies by cascading 2D CNN and Transformer layers; TECC proposes a two-stream convolutional framework based on Transformer enhancement to efficiently learn complementary spatial-spectral dependencies.
[0145] Graph neural network architecture: AMGCF adopts a superpixel-guided graph convolutional network combined with a multi-hop message passing mechanism to capture the structured correlation of hyperspectral images; MRCAG integrates multi-scale random shape convolution and adaptive graph convolution to process heterogeneous spatial structures; PCCGC constructs a hierarchical graph topology driven by multi-scale spatial-semantic cues to achieve contextual feature aggregation; DGFNet collaboratively models the spectral correlation and spatial structure information of hyperspectral images through a graph attention network and a two-stream strategy.
[0146] To ensure the fairness of the comparative experiments, all baseline methods use the hyperparameter configurations reported in their original literature. In the SGCFN framework proposed in this paper, superpixel generation uses the SLIC algorithm, and the segmentation scale is set according to the spatial resolution and scene complexity characteristics of different datasets: the IP and MUUFL datasets are set to , the UP dataset is (in is the input data space size). All datasets uniformly set the hidden layer dimension to 64.
[0147] The model is optimized using the Adam optimizer, and the basic learning rate is set to 5×10 -4 , momentum parameter (0.9, 0.999), numerical stability parameter 1×10 -8 The learning rate is dynamically adjusted through the cosine annealing strategy to escape the local minimum. At the same time, an early stopping mechanism is set to terminate the training when the validation set accuracy does not improve for 20 consecutive epochs, effectively preventing overfitting.
[0148] Performance evaluation used three classic metrics: overall classification accuracy (OA), average classification accuracy (AA), and Kappa coefficient (Kappa). To ensure statistical reliability, all experiments were repeated 10 times and the results were averaged. The training and validation sets were randomly split between each trial. The experimental platform was equipped with an Intel Xeon E5-2680v4 CPU and an NVIDIA RTX2080Ti GPU, running CUDA 11.2, PyTorch 1.10, and Python 3.8.
[0149] Experimental results
[0150] Experimental results on the Indiana Pine Forest (IP) dataset (shown in Table 4) reveal significant performance differences between our proposed method and ten baseline methods. Class-by-class classification accuracy analysis shows that our proposed method (our method) demonstrates significant superiority across most object classes, particularly for complex objects: achieving 99.44% accuracy for alfalfa (C1), 99.07% for tree grassland (C6), 99.63% for formed hay (C8), and 100% perfect classification for wheat (C13). Our method also performs well in heterogeneous scenes such as woodland (C14, 99.91%) and mixed areas of buildings, grassland, and trees (C15, 99.77%), demonstrating its ability to model hierarchical spatial-spectral features. However, when distinguishing agricultural subclasses with similar spectral characteristics (e.g., 94.07% accuracy for conventionally tilled corn (C3) and 87.64% for cornfield (C4), our proposed method's accuracy is slightly lower than that of the comparison methods, including SSGCA, MSCA, and PCCGC. Notably, these values remain within a reasonable range and lack extreme anomalies, reflecting the inherent challenges of classifying similar crop types rather than structural flaws in the model. Overall, these results demonstrate the robustness of the method across diverse landforms and suggest optimization strategies for enhancing the ability to discern subtle spectral differences within agricultural subcategories.
[0151] From the perspective of comprehensive performance indicators, the present invention achieved a leading level with 97.73% OA, 97.29% AA and a Kappa coefficient of 0.9740, surpassing all baseline methods in all aspects. Among them, compared with the similar graph convolution method AMGCF, the OA, AA and Kappa coefficients were improved by 1.45%, 1.82% and 1.64% respectively, verifying its excellent discrimination ability in complex scenarios. Among the other comparison methods, PCCGC and TECC were close to the optimal level with 97.63% and 97.18% OA respectively; MRCAG performed the lowest with only 76.39% OA, and the OA of other methods was concentrated in the range of 95.06%-97.07%. These results show that the proposed algorithm has significant advantages in classification accuracy and result consistency.
[0152] In terms of model efficiency, the method of the present invention has 120.98K parameters, which is at a medium complexity level. The high training time of 244.79 seconds is due to the computational complexity of the multi-scale convolution and dual attention modules, as well as the additional training cycles required to achieve the optimal solution. It is worth noting that the method takes only 0.02 seconds to infer a single image, which is on par with AMGCF (0.02 seconds) and significantly better than methods such as TECC (3.53 seconds) and PCCGC (3.23 seconds). This high efficiency is due to the fact that the model generates classification results through a single forward propagation during the inference phase, while most comparison methods (except AMGCF) rely on a pixel-by-pixel iterative processing mechanism, resulting in a sharp increase in computational overhead.
[0153] Visual analysis of classification maps of different methods for the Indiana Pine Forest (IP) dataset (e.g. Figure 7 (as shown) highlights the significant advantages of the proposed method (k). Compared to other methods (aj), the classification map generated by (k) is closer to the actual distribution of objects in terms of regional segmentation accuracy and preservation of spatial details: for example, the areas corresponding to alfalfa (C1) and arbor grassland (C6) in (k) appear compact, clearly defined patches with minimal misclassified pixels, effectively alleviating the spectral confusion common in similar methods. In contrast, baseline methods such as (g) and (i) exhibit blurred boundaries and misclassification in heterogeneous areas such as conventionally cultivated corn (C3) and cornfields (C4), where subtle spectral differences make discrimination difficult. This visual consistency confirms the advantages of this method in quantitative indicators such as 99.44% classification accuracy for alfalfa (C1) and 97.73% overall classification accuracy, jointly verifying its enhanced ability to capture discriminative spatial-spectral features, making the classification results as close to the actual distribution of objects (l) as possible.
[0154] On the University of Pavia (UP) dataset (experimental results shown in Table 5), our proposed method (our) consistently demonstrates strong performance compared to the 10 baseline methods used in the IP dataset analysis, demonstrating its universal effectiveness in diverse hyperspectral scenarios. In terms of per-class classification accuracy, our method achieves excellent results for most object categories, maintaining the robust discrimination capabilities demonstrated by the IP dataset: asphalt pavement (C1) reaches 99.10%, gravel areas (C3) reaches 99.96%, and woodlands (C4) reaches 99.96%. Metal roofs (C5) achieve 100% perfect classification, and bare soil (C6) reaches 99.94%. Even for brick structures (C8), which are susceptible to spatial-spectral aliasing, our method achieves 98.28% accuracy, highlighting its ability to handle the fine-grained structural differences typical of urban areas.
[0155] In terms of comprehensive performance indicators, the method of the present invention continues its excellent performance on the IP dataset: it ranks first among the 11 compared methods with 99.18% OA, 99.07% AA and 0.9891 Kappa coefficient, which are 0.06% OA, 0.12% AA and 0.08% Kappa higher than the second-best method AMGCF, respectively, showing the consistency advantage of multiple evaluation dimensions.
[0156] In terms of model efficiency, this method maintains a parameter count of 120.64K, comparable to the configuration complexity of the IP dataset. Due to the computational demands of multi-scale feature fusion and the attention mechanism, training takes 162.43 seconds, the longest of all methods. However, inference takes only 0.03 seconds, remaining highly competitive.
[0157] Comparative analysis of the classification map (ak) and false color image (l) of the University of Pavia (UP) dataset (e.g. Figure 8 The superior classification performance of the proposed method (k) is demonstrated in the classification map generated by the proposed method. In the classification map generated by the proposed method, asphalt pavement (C1) appears as coherent spatial clusters with sharp boundaries, closely aligned with the visual representation in the false color image (l); the delineation of grassland (C2) matches its continuous range shown in (l), effectively reducing misclassification with adjacent urban surface features. The spatial distribution pattern of gravel areas (C3) in the classification map (k) closely reproduces the irregular shapes visible in the false color image, while the segmentation of arbor woodland (C4) is exceptionally accurate, capturing the granular texture characteristics unique to the tree canopy structure in (l). For categories such as sheet metal roof (C5), bare soil (C6), and brick structure (C8), the color-coded regions in (k) closely correspond to the spectral features evident in the false color image, demonstrating the method's ability to convert spectral-spatial features into accurate spatial classifications. Even for challenging categories that may have slight spectral overlap in false color displays (such as asphalt roof C7 and shadow C9), the classification map of the proposed method still maintains spatial coherence consistent with (l), effectively reducing the classification ambiguity in heterogeneous urban areas.
[0158] On the MUUFL dataset, our proposed method (our) demonstrates competitive performance across a diverse range of fine-grained urban land cover classes (see Table 6 for details). In terms of per-class classification accuracy, our method achieves significant results in challenging classes: 82.64% for yellow curb (C10), 95.23% for canvas awning (C11), 92.04% for road (C5), and 94.09% for building (C8). While our method achieves an accuracy of 77.32% for predominantly grass (C2)—slightly lower than comparable methods such as PCIA (79.88%) and PCCGC (80.67%)—this difference remains within a reasonable range, likely due to the spectral similarity between this class and neighboring vegetation and ground surfaces. Importantly, no classes exhibit extreme misclassification, demonstrating the robustness of our method even in complex scenes with class imbalance.
[0159] In terms of comprehensive performance, the proposed algorithm achieved optimal performance, achieving an overall classification accuracy (OA) of 90.70%, an average classification accuracy (AA) of 84.90%, and a Kappa coefficient of 0.8773, ranking first among all methods. Compared to the next-best method, AMGCF, the proposed method achieved improvements of 0.46% OA, 6.99% AA, and 0.59% Kappa, respectively. In terms of model efficiency, the proposed method had a medium parameter count of 120.73 K, a training time of 129.93 seconds, and a testing time of 0.03 seconds.
[0160] Visual analysis of the classification maps shows that the proposed method exhibits excellent matching with the false color image (l) in multiple categories (e.g. Figure 9 As shown in Figure 2 ). Trees (C1) appear as spatially compact regions with sharp boundaries, eliminating the fragmented misclassification common in baseline methods such as (i) and highly consistent with the continuous canopy structure visible in the false color visualization. The demarcation accuracy of the main grassland (C2) area is significantly improved, and its distribution range more accurately matches the uniform green area in (l), effectively reducing confusion with adjacent categories with similar spectral characteristics such as mixed surface (C3) or soil and gravel (C4). Roads (C5) in the classification map generated by the method of the present invention show stronger continuity and structural integrity, can capture the details of complex road networks and minimize fractures / dislocations. Similarly, the building (C8) area in (k) has a regularized geometry and consistent spatial distribution characteristics, accurately reproducing the structural characteristics of the man-made environment shown in (l).
[0161] Table 4: Overall classification accuracy (%), average accuracy (%), Kappa coefficient (×100), number of parameters (K), training time (s), and testing time (s) of the IP dataset
[0162]
[0163] Table 5: Overall classification accuracy (%), average accuracy (%), Kappa coefficient (×100), number of parameters (K), training time (seconds), and testing time (seconds) for the UP dataset
[0164]
[0165] Table 6: Overall classification accuracy (%), average accuracy (%), Kappa coefficient (×100) and training time (s) of the MUUFL dataset
[0166]
[0167] Ablation experiments
[0168] To verify the performance of each module in the network, we conducted ablation experiments on three datasets. For ease of illustration, the modules are numbered as follows: 1. MAGAT; 2. MSAAT; 3. PISGAT; 4. AGFS.
[0169] Verification of Module Independence: We first evaluated the independent performance of each module, retaining only a single module at a time through ablation experiments. The results show that MAGAT (ab-1) achieves the best single-module performance on all datasets (IP: 95.59% OA, UP: 98.7% OA, MUUFL: 89.71% OA), significantly outperforming other modules. Taking the IP dataset as an example, ab-1 outperforms ab-2 (MSAAT), ab-3 (PISGAT), and ab-4 (AGFS) by approximately 24.77%, 24.26%, and 24.39%, respectively. This demonstrates the effectiveness of MAGAT in capturing superpixel-level spatial dependencies through its dual-hop graph attention mechanism, enabling robust structural feature extraction even without the introduction of auxiliary modules. In comparison, MSAAT (ab-2)'s performance is limited in the absence of superpixel semantic guidance (IP: 70.82% OA, UP: 93.79% OA, MUUFL: 83.6% OA). Its pixel-level multi-scale texture extraction capability struggles to distinguish spectrally similar classes, confirming its reliance on hierarchical spatial context. PISGAT (ab-3), which focuses on spectral optimization, achieves IP: 71.33% OA, UP: 93.35% OA, and MUUFL: 83.8% OA. While slightly better than MSAAT, it still lags significantly behind MAGAT, indicating that its spectral optimization effectiveness relies on the synergy of spatial context features. AGFS (ab-4), as a fusion module, achieves the lowest performance (IP: 71.2% OA, UP: 94.12% OA, MUUFL: 83.44% OA), confirming its role as an auxiliary module for feature fusion rather than a standalone feature extractor.
[0170] A second set of ablation experiments validated the synergistic effect of the module combination: ab-12 retained the spatial branch modules MAGAT and MSAAT, while ab-34 retained the spectral branch modules PISGAT and AGFS. On the IP dataset, ab-12 achieved 97.37% OA, a 1.78% improvement over the single-module MAGAT (ab-1) and a 26.0% improvement over ab-34 (71.38% OA). This demonstrates that the superpixel-level semantic modeling of MAGAT and the pixel-level multi-scale texture extraction of MSAAT work synergistically to effectively enhance hierarchical spatial feature representation and alleviate the misclassification of spectrally similar agricultural subcategories. On the UP dataset, ab-12 achieved 98.95% OA, only 0.23% lower than the full model, demonstrating its ability to accurately capture spatial details. However, ab-34's 93.48% OA was 5.47% lower than ab-12, highlighting the limitations of pure spectral optimization in handling complex objects without spatial feature support. On the MUUFL dataset, ab-12 has an OA advantage of 6.65% over ab-34, further verifying the superiority of the spatial module combination in distinguishing fine-grained urban features such as roads and building shadows by integrating superpixel context dependencies and pixel-level texture details.
[0171] The third set of ablation experiments examined the synergistic effects of multiple modules (ab-123, ab-124, ab-134, and ab-234) to explore the interactions of the core components. ab-123, which retains MAGAT, MSAAT, and PISGAT, achieved OA of 97.58%, 99.02%, and 90.58% on the IP, UP, and MUUFL datasets, respectively. The differences from the full model (our) were only 0.15%, 0.16%, and 0.12%, respectively. This demonstrates that the spatial dual modules (MAGAT and MSAAT) form an efficient framework through superpixel-level semantic modeling and pixel-level texture extraction, combined with the spectral feature optimization of the spectral module (PISGAT). AB-124, which retains MAGAT, MSAAT, and AGFS, achieves 99.07% OA on the UP dataset (only 0.11% lower than the full model). However, its OA on IP and MUUFL decreases by 0.27% and 0.42%, respectively. This suggests that while the dynamic fusion of AGFS can compensate for the lack of PISGAT in the spectral dimension, spectral filtering mechanisms are still required to ensure accuracy. AB-134, which retains MAGAT, PISGAT, and AGFS, experiences a sharp drop in OA to 95.48% on the IP dataset (2.1% lower than AB-123), with a significant increase in misclassification of spectrally similar classes. OA on the UP and MUUFL datasets reaches 98.83% and 90.11%, respectively (0.19% and 0.47% lower than AB-123), demonstrating that superpixel semantics can partially compensate for the performance loss caused by the lack of texture detail in urban scenes. In contrast, ab-234, which retains MSAAT, PISGAT, and AGFS, shows a sharp performance drop on all datasets (IP: 73.8% OA, UP: 93.99% OA, MUUFL: 84.83% OA), confirming that MAGAT's superpixel-level structural modeling is irreplaceable. Its absence causes the model to lose the ability to capture hierarchical spatial dependencies, making the remaining modules ineffective in complex scenes.
[0172] In summary, ablation experiments show that MAGAT is the core feature extraction component of this network, ensuring the accuracy of the model; MSAAT and PISGAT play a complementary role by optimizing pixel-level multi-scale spatial features and spectral features respectively; AGFS realizes the dynamic fusion of spectral and spatial features, and together builds the robust classification capability of SGCFN.
[0173] The impact of the percentage of training samples
[0174] To evaluate the impact of varying labeled data ratios on algorithm performance, experiments were conducted on three datasets. Experimental results on the IP dataset showed that as the training sample ratio increased from 1% to 10%, the OA of all methods increased, indicating that more labeled data generally enhances model performance by providing richer spectral and contextual information. Among the baseline methods, MRCAG performed worst at low sample ratios (59.02% at 1%), gradually increasing to 85.23% at 10% as the ratio increased. SSGCA also struggled with minimal data, achieving an OA of 38.69% at 1% (the lowest of all methods), but recovering to 97.75% at 10%. In contrast, our method (our) demonstrated excellent stability and accuracy at all ratios: reaching 87.15% at 1%, significantly higher than most baseline methods, and steadily improving to 98.98% at 10%, comprehensively surpassing its competitors. While competing methods such as MSCA, DBCT, and PCCGC perform well, our method remains in the lead, especially at medium and high sample ratios. This is due to its adaptive feature fusion and multi-scale attention mechanism, which can efficiently extract discriminative representations from limited data and alleviate overfitting. DGFNet exhibits volatility: its OA reaches 97.01% at 9%, but drops to 96.08% at 10%, indicating instability when optimizing model capacity with sufficient data. Overall, our method balances the performance of small samples with the scalability of multiple samples, demonstrating strong adaptability to different training data sizes and validating its practical advantages in real-world scenarios with limited labeled data.
[0175] On the UP dataset, experimental results show a similarly steady upward trend in overall classification accuracy (OA) as the training sample ratio increases from 0.5% to 10%, consistent with observations on the IP dataset. This consistency suggests that increasing labeled data generally enhances the model's ability to effectively learn spatial-spectral features. Similar to the IP dataset, the baseline method MRCAG performs poorly at low sample ratios: its OA is 84.23% at 0.5% of samples, significantly lower than other methods. This performance gradually improves to 97.67% at 10% of samples, but still lags behind most compared methods. In contrast, the excellent performance of our method on the IP dataset persists on the UP dataset: achieving 98.13% OA at 0.5% of samples, significantly exceeding baseline methods such as MRCAG (84.23%) and DGFNet (92.04%), demonstrating robust small-shot learning capabilities. Its OA steadily improves with increasing sample ratios, reaching 99.93% at 10%, comparable to high-performing baseline methods such as DBCT (99.98%) and AMGCF (99.90%). Notably, at 1% sample size, the proposed method achieves 99.18% OA, surpassing PCIA (98.33%) and SSGCA (98.80%). At 5% sample size, the proposed method achieves 99.88% OA, surpassing the accuracy of most baseline methods at 10% sample size. This advantage stems from its adaptive gated fusion and multi-scale attention mechanisms, which can efficiently extract discriminative features from limited data, as well as the regularization technique's mitigation of overfitting. While methods such as DBCT perform well at high sample ratios, the proposed method's advantages are most pronounced at medium and low sample ratios. This finding is consistent with observations from the IP dataset, validating its robustness and generalization capabilities at varying data scales. This reinforces its practical value in real-world scenarios where labeled data is scarce or requires incremental acquisition, demonstrating a unique balance between small-sample efficiency and scalability.
[0176] Experimental results on the MUUFL dataset show that the overall classification accuracy (OA) of all methods shows a consistent upward trend as the training sample ratio increases from 0.5% to 10%, a pattern similar to that observed on the IP and UP datasets. Among the baseline methods, MRCAG again significantly underperforms at low sample ratios: its OA is only 76.33% at 0.5% and gradually improves to 91.04% at 10%. Our method (our) consistently demonstrates robust performance across all sample ratios on the MUUFL dataset, achieving 88.6% OA at 0.5% training sample, significantly exceeding baseline methods such as MRCAG (76.33%) and DGFNet (83.47%), demonstrating strong small-shot learning capabilities. Its OA steadily improves with increasing sample ratios, reaching 97.13% at 10%, outperforming high-performing baseline methods such as PCIA (96.77%), AMGCF (96.24%), and MSCA (96.07%). Notably, our method achieves 90.70% OA at 1% sample size, surpassing SSGCA (87.42%) and TECC (88.68%). At 5% sample size, it reaches 96.17%, surpassing the performance of most baseline methods at 10% sample size. While methods such as DBCT and PCCGC show steady improvement, our method maintains a leading advantage across all sample sizes, and its ability to extract discriminative information from sparse data is particularly impressive at low and medium sample sizes. These results, combined with findings from the IP and UP datasets, demonstrate the generalizability and robustness of our method across diverse hyperspectral datasets.
[0177] in conclusion
[0178] This paper proposes the Spectral-Spatial-Graph Cooperative Fusion Network (SGCFN), a two-stream architecture designed to address the challenges of modeling complex spatial-spectral interactions, spectral redundancy, and computational inefficiency in hyperspectral image classification. By integrating hierarchical spatial feature learning, adaptive spectral optimization, and a dynamic cross-modal fusion mechanism, the network effectively combines the complementary strengths of graph attention, agent self-attention, and data-driven spectral modeling. Experimental results on three benchmark datasets validate the advanced performance of SGCFN, and ablation studies confirm the irreplaceable role of each module in enhancing classification robustness.
[0179] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network, characterized by: The following steps are involved: Step 1: Build the overall network architecture: A hybrid two-stream fusion model, SGCFN, is proposed. It uses an improved graph attention network and a multi-scale proxy self-attention mechanism to extract superpixel-level and pixel-level spatial features, respectively. It then uses a pooling-induced spectral graph attention module to mine discriminative spectral information. Finally, an adaptive gated fusion module is used to integrate two-stream features, effectively improving the accuracy of hyperspectral imagery object classification. Step 2: Spatial Branching - Hierarchical Spatial Feature Modeling: To address the static structural limitations of graph networks and the high computational complexity of the Transformer method, we design a hierarchical multi-adjacency graph attention module (MAGAT) and a lightweight multi-scale proxy attention module (MSAAT). The hierarchical multi-adjacency graph attention module (MAGAT) uses dual-hop graph attention to dynamically capture local geometric details and regional semantic dependencies between superpixels, and combines it with a channel splitting strategy to enhance feature diversity. The lightweight multi-scale proxy attention module (MSAAT) integrates multi-scale depthwise separable convolution with a proxy self-attention mechanism to reduce computational costs while maintaining global context modeling capabilities. Step 3: Spectral branching - adaptive spectral feature optimization: We propose a pooling-induced spectral graph attention module (PISGAT), which generates channel-level node embeddings through multi-scale spatial pooling and dynamically constructs an adaptive spectral graph structure, effectively suppressing spectral redundancy and enhancing discriminative features. Step 4. Design of Adaptive Gated Fusion Module (AGFM): Design a pixel-level learnable adaptive gated fusion module (AGFM). Dynamically adjust the feature weights of spatial and spectral branches based on the scene context, enhance spatial structural features in spectrally complex regions to distinguish heterogeneous boundaries, and highlight spectral identification features in spatially fragmented regions to improve intra-class consistency.
2. The method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network according to claim 1, characterized in that: The first step is specifically as follows: the hyperspectral image data is formalized into an input cube , where H, W, and B represent the image height, width, and number of spectral bands, respectively. To exploit spatial regularity, the original image is segmented into N superpixels by simple linear iterative clustering (SLIC), generating three structured inputs: a superpixel feature matrix containing C-dimensional node embeddings, , a multi-hop adjacency matrix defining the hierarchical connectivity pattern between superpixels , and the projection matrix describing the pixel-superpixel mapping relationship ; The hybrid two-stream fusion model SGCFN consists of three collaborative components: a spatial branch for superpixel-level and pixel-level spatial information modeling, a spectral branch for channel-level discriminative feature learning, and an adaptive gated fusion module for cross-modal feature integration.
3. The method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network according to claim 2, characterized in that: First, spatial branch modeling processes superpixel maps through MAGAT , where the parallel graph attention network is based on different adjacency matrices Capturing hierarchical spatial correlations; after being decoded by the projection matrix Q, these spatial features are further enhanced by MSAAT, fusing multi-scale depthwise separable convolution with a lightweight proxy self-attention mechanism to obtain pixel-level spatial information; Secondly, the spectral branch uses 1×1 convolution to transform the original hyperspectral data ( ) is mapped to the latent space, and then the channel relationship is optimized through PISGAT: this process uses multi-scale mean pooling to generate spatial descriptors to dynamically construct the channel map, and implements graph convolution propagation to suppress redundant bands; finally, the adaptive gated fusion module bridges the two branches through AGFM, uses learnable convolutional gates to generate pixel-level fusion weights, and selectively fuses features according to contextual relevance; the fused features are processed by a feedforward network with a hidden layer dimension of C / 2 and the classification results are output. , where c is the number of categories.
4. The method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network according to claim 1, characterized in that: In the step 2, the hierarchical multi-adjacency graph attention module MAGAT is specifically as follows: the standard graph attention network GAT processes graph structure data through a dynamic attention mechanism, aiming to adaptively capture the dependencies between nodes without relying on a predefined adjacency matrix; the standard GAT is improved so that the network can adjust the attention weight according to the adjacency matrix; specifically, given the input superpixel feature , GAT first passes the learnable weight matrix Project features into latent space: in and ; For each superpixel i, it is compared with the neighboring superpixels The attention coefficient between is calculated as follows: Among them, concat(∙) represents the vector concatenation operation, and a is the shared attention vector; these attention coefficients are filtered by the positive values of the corresponding positions in the adjacency matrix (A) and normalized by the SoftMax function to generate the attention weight reflecting the importance of superpixel j to i. : Then, the attention weight is used to aggregate neighborhood features, and its expression is: Where σ(∙) represents the activation function of the stabilization process. To enhance the diversity of feature representation, the multi-head attention mechanism uses K independent attention heads to operate in parallel, and then splices the output results of each head. Its mathematical expression is: The outputs of each head are spliced together to form the final feature ; This mechanism dynamically infers the association between superpixels through data, thereby effectively capturing contextual dependencies in irregular spatial structures. However, the standard GAT is limited by the homogeneous adjacency mechanism, which imposes the same interaction range on all nodes. This construction method cannot fully characterize the hierarchical spatial associations of hyperspectral superpixels: neighboring regions need to preserve local patterns, while distant regions benefit from extended context aggregation. To solve this problem, the proposed MAGAT introduces a two-hop hierarchical aggregation architecture. The single-hop GAT layer captures direct neighborhood relationships to preserve geometric details. Its output features are spliced with deep group features and then processed by the two-hop GAT layer to fuse regional semantic information. Given the input superpixel features , this module first divides the features into two complementary groups by channel splitting, and its operation is expressed as: Then, By The guided GAT layer is processed; the process is expressed as: Where LeakyReLU(∙) represents the LeakyReLU activation function, LayerNorm(∙) is the layer normalization operation, and GAT(∙) represents the graph attention network; then, the intermediate features and Stitching to form enhanced features , and through the linear layer and by Guided secondary GAT layer processing; Hierarchical fusion through learnable coefficients To realize the combination of multi-hop features, the operation is expressed as: 。 5. The method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network according to claim 1, characterized in that: In step 2, the lightweight multi-scale proxy attention module MSAAT is specifically as follows: the lightweight multi-scale proxy attention module MSAAT innovatively integrates lightweight proxy guided attention with multi-scale depth-separable convolution; Based on the hierarchical features extracted by MAGAT, MSAAT enhances spatial feature learning through a local-global hybrid interaction mechanism; Given the input features mapped by the superpixel space through the projection matrix Q , ,in( ), first execute the query Q, key K, value V) shadow: in , , is a learnable linear transformation matrix; in order to reduce the computational complexity while maintaining global interaction, adaptive pooling and shaping operations are performed to transform Reshape the generation of proxy tokens in the form of compact spatial grids in( ): The proxy token acts as a lightweight representation of the global context; it then generates a proxy value by interacting with the key and value: In the formula represents the Dropout layer, and Softmax(∙) is the SoftMax operation; Then, the query and proxy values are attention-corrected via the proxy token: To supplement the local detail features in the attention mechanism, parallel multi-scale depth-wise separable convolution is used to process the numerical terms: Finally, the fusion of multi-scale features is achieved through residual connection: In the formula Represents the multi-scale feature integration result.
6. The method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network according to claim 1, characterized in that: The specific steps of step 3 are as follows: the proposed PISGAT enhances the hyperspectral feature representation by fusing adaptive spatial pooling and dynamic graph convolution mechanism; given an input feature map , first compress the spatial dimension by adaptive average pooling: Where S is the spatial resolution after downsampling; the pooled features are projected into the latent space through 1×1 convolution: Where r is the channel reduction rate; then Flatten to node embedding features ( ) to process the graph structure; design the graph channel attention GCAT module to generate channel weights for the input features, by introducing the static unit matrix To maintain the intrinsic self-connection relationship between nodes and ensure the stability of the basic topology; to inject input-specific adaptability, the pooled node embedding After 1×1 convolution and SoftMax normalization, a data-dependent adjustment matrix is generated: This matrix captures the relationship between nodes under the input feature conditions; introduces the learnable matrix To supplement the above components, the graph topology is gradually optimized during training to achieve task adaptability; the final adjacency matrix is composed of the following components: in Represents element-by-element multiplication; spectral graph convolution updates node features through the following formula: In the formula is the learnable weight matrix, σ(∙) represents the ReLU activation function; the updated embedding features After 1×1 convolution expansion, the channel attention weight is generated by global average pooling: These weights recalibrate the input features through channel-wise multiplication: Where ⊗ represents element-wise multiplication under the broadcast mechanism.
7. The method for constructing an adaptive gated spectral-spatial-graph collaborative fusion network according to claim 1, characterized in that: The step 4 is specifically as follows: dynamically calibrating cross-branch feature contributions through learnable spatial-spectral attention; Given spatial features and spectral characteristics , first concatenate the two along the channel dimension and generate the gating weights through a hierarchical convolutional layer: Where the gated tensor Contains two spatial attention maps To achieve weighted fusion: where ⊙ represents element-wise multiplication.
Citation Information
Cited By
Intelligent urine component detection method based on spectrum identification and deep learning
CN121190854A
An intelligent urine component detection method based on spectral recognition and deep learning
CN121190854B
Hyperspectral image classification method and system based on multi-scale spatial-spectral joint representation and dynamic context modeling
CN121438106A
Hyperspectral image unmixing method and system based on double-path attention gating fusion
CN121921623A
Hyperspectral image unmixing method and system based on dual-path attention gate fusion
CN121921623B