Hyperspectral image self-supervised deep clustering based on superpixel guided semantic transformer and system
Patent Information
- Application Number
- CN202610348882.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-20
- Publication Date
- 2026-09-25
AI Technical Summary
二者在双分支网络的结构分工与全局语义建模路径上存在明显差异
[0053]本发明的有益效果是:本发明利用MSCB分支和超像素语义Transformer分支分别提取高光谱图像局部细节特征和全局语义特征,通过双重自监督机制利用KL散度作为聚类对齐损失并结合重构约束来指导整个网络模型进行训练,从而实现对高光谱图像的无监督分类。具体来说:
Smart Images

Figure CN122821183A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unsupervised hyperspectral image technology, specifically relating to a self-supervised deep clustering method and system for hyperspectral images based on superpixel-guided semantic Transformer. Background Technology
[0002] Hyperspectral images are three-dimensional cubic data obtained by imaging ground targets using a hyperspectral imager. Compared to three-channel visible light images, they contain tens to hundreds of continuous spectral bands, providing rich spectral and spatial information simultaneously, facilitating the differentiation of ground cover categories, and are widely used in environmental monitoring, precision agriculture, urban planning, and other fields. Hyperspectral image clustering aims to divide pixels into clusters based on their spectral characteristics and spatial context when annotations are lacking. However, sample annotation usually relies on field surveys, which is time-consuming, labor-intensive, and costly, and difficult to carry out in disaster-prone or inaccessible areas. Large-scale hyperspectral images are characterized by high spectral dimensionality, subtle inter-class differences, and irregular spatial distribution, which places higher demands on feature representation learning for unsupervised clustering. In addition, hyperspectral images suffer from spectral redundancy, noise interference, and class mixing, and pixels in boundary regions are easily confused by their neighbors. As the image coverage area increases and the spatial resolution improves, the data size grows, and existing methods struggle to balance computational complexity and clustering accuracy.
[0003] In recent years, deep learning methods have been widely applied to unsupervised clustering tasks in hyperspectral images due to their strong feature extraction capabilities. Self-supervised deep clustering typically optimizes feature representation and cluster partitioning by constructing pseudo-labels or target distributions. For example, DeepCluster alternates between feature extraction and K-means assignment, using the clustering results as pseudo-labels to supervise network training. Xie et al.'s Deep Embedded Clustering (DEC) uses a KL divergence minimization strategy to learn the target distribution, refining the cluster structure in the embedding space. Building on this, the improved IDEC introduces autoencoder reconstruction constraints in addition to the clustering loss to maintain local structure and prevent feature degradation during cluster optimization. Furthermore, contrastive learning has been introduced into clustering to enhance feature discriminativity, constructing positive and negative sample pairs through data augmentation and jointly performing instance-level and cluster-level consistency learning. SwAV et al. further employ prototype assignment and exchange prediction mechanisms to achieve scalable unsupervised representation learning. The aforementioned methods mostly use convolutional neural networks as the backbone, which can uncover local spectral-spatial details. However, fixed convolutional kernels and preset receptive fields limit their ability to model long-range dependencies, resulting in insufficient understanding of cross-regional semantic relationships. With the introduction of the visual Transformer, the self-attention-based Transformer can process long-range dependencies in parallel and model global context, showing outstanding performance in tasks such as image classification and dense prediction. However, the computational complexity of traditional fully connected self-attention increases quadratically with spatial resolution, making it difficult to directly apply to large-scale hyperspectral images (where high spatial resolution and high spectral dimensionality coexist). To reduce complexity, researchers have proposed window-based visual Transformers, such as the Swing Transformer, which enables cross-window information interaction by moving the window; PVT, which introduces a pyramid structure to generate multi-scale features and shorten sequence length; and Shuffle Transformer and HaloNet, which compute attention within local windows to improve efficiency. Meanwhile, CSWin Transformer computes attention along horizontal and vertical intersecting windows to enhance connectivity; RegionViT utilizes region partitioning and intra-region attention to optimize computation; and Twins Transformer combines local window attention with a global subsampling mechanism to strengthen global modeling. However, in hyperspectral images, the same semantic feature region may exhibit a fragmented, spatially discontinuous, and irregularly shaped distribution. Local window attention often assumes spatial proximity equals semantic consistency, easily leading to feature mixing at boundaries and difficulty in associating spatially separated but spectrally consistent regions, thus weakening inter-cluster separability. On the other hand, single-branch models, while pursuing global semantic consistency, often sacrifice the preservation of fine-grained local structures, causing small targets, texture details, or weakly contrasting categories to be ignored or misclassified in clustering.Furthermore, deep clustering objectives often rely on the quality of initial cluster centers or pseudo-labels; poor initial segmentation can easily lead to local optima. In hyperspectral images, subtle differences between categories and large intra-class variations amplify the interference of pseudo-label noise on training. While window segmentation reduces the complexity of the Transformer, each query only interacts with tokens within the window, making it difficult to establish connections between semantically homogeneous regions distributed across windows. Simultaneously, fixed grid windows are unsuitable for terrain boundary shapes, easily causing the mixing of dissimilar features and fragmented clustering results near boundaries. Existing research utilizes superpixel segmentation to obtain regional priors to improve spatial consistency, but these studies largely remain at the post-processing level and have failed to efficiently apply them to attention computation and cross-regional information aggregation in end-to-end self-supervised clustering. Moreover, they struggle to simultaneously address the requirements of preserving local details and multi-scale context fusion. Therefore, existing technologies urgently need a self-supervised clustering framework that can combine multi-scale local detail extraction and cross-regional semantic aggregation with lower computational overhead to improve the accuracy and stability of unsupervised clustering of large-scale hyperspectral images.
[0004] The clustering and Transformer methods mentioned above often employ a single structure or fixed window attention, making it difficult to preserve local details and model cross-regional semantic relationships. In large-scale images, problems such as boundary blending and category confusion arise.
[0005] I. Upon searching, Chinese invention patent CN118072059A provides a method for clustering ground features in hyperspectral remote sensing images using a self-supervised dual-branch Transformer structure, which differs from the application in the following ways:
[0006] 1. This patent focuses on constructing a local-global collaborative self-supervised deep clustering framework for large-scale hyperspectral images. Its core lies in extracting multi-scale local detail features using the MSCB module, based on shallow spatial-spectral joint features extracted by ResNet, and simultaneously modeling semantic consistency and long-range semantic associations across spatially continuous regions using the superpixel semantic Transformer module. In contrast, patent CN118072059A's core technology involves first performing superpixel segmentation on the hyperspectral remote sensing image, then learning the global attribute information of the HSI data through a shared Autoformer module, and extracting more accurate graph structure features through a Dual-Former Graph module. Essentially, it is a dual-branch clustering scheme based on "global attribute representation + graph structure enhancement." The two differ significantly in their structural division of labor in the dual-branch network and their global semantic modeling paths.
[0007] 2. This patent, in its use of superpixels, does not merely treat them as preprocessed region partitioning results, but directly introduces them as semantic priors into the attention computation process: on the one hand, it uses Semantic Continuous Attention (SCA) to constrain and aggregate within the superpixel region and adjacent regions to alleviate boundary aliasing; on the other hand, it uses Semantic Global Attention (SGA) to globally associate spatially non-adjacent but semantically consistent regions, thereby simultaneously addressing the issues of insufficient consistency in continuous regions and insufficient association of similar targets across regions. In contrast, patent CN118072059A primarily uses superpixels as the foundation for graph node construction and then performs graph structure feature learning on this basis. The former is an attention aggregation mechanism under superpixel prior constraints, while the latter is a graph structure modeling mechanism after superpixel nodeization.
[0008] 3. This patent employs a strategy of soft assignment of Student t distribution, construction of target distribution, three-way KL divergence consistency alignment (local branch / semantic branch / fusion branch), and joint constraints on reconstruction loss, thereby forming an end-to-end self-supervised optimization framework. In comparison, although patent CN118072059A also adopts the joint optimization idea, its optimization focus lies in the collaborative learning between the shared Autoformer module and the Siamese Dual-Former Graph module. The former is a clustering framework that performs multi-branch consistency constraints around "local details + global semantics + fusion representation," while the latter is a clustering framework that performs dual-structure collaborative optimization around "Autoformer features + graph structure features."
[0009] II. According to the search, Chinese invention patent with publication number CN117746079A provides a clustering prediction method, system, storage medium and device for hyperspectral images, which differs from this application in the following ways:
[0010] 1. This patent focuses on constructing a self-supervised deep clustering scheme for hyperspectral images with "local detail preservation + global semantic enhancement" as its core. The core lies in extracting multi-scale local spatial spectral details through the MSCB module and explicitly modeling the semantic relationships between continuous and non-contiguous homogeneous regions through the superpixel semantic Transformer module. In contrast, the core technology of patent CN117746079A involves performing superpixel segmentation on the hyperspectral image, extracting pixel features and superpixel features separately, combining contrastive learning to obtain deep representations, and then using K-means clustering to generate pseudo-labels and output clustering prediction results. The two patents differ fundamentally in their feature learning mechanisms and clustering generation logic.
[0011] 2. This patent employs an integrated design of "multi-scale convolutional local branch + superpixel-guided semantic attention branch + target distribution alignment + reconstruction constraints," enabling the clustering target to directly participate in the network representation learning process. In contrast, the method in patent CN117746079A focuses more on the implementation path of "pixel-level / superpixel-level dual feature extraction + contrastive learning + K-means pseudo-label generation," with its overall framework relying on the clustering results generated after deep feature learning. The former is an end-to-end deep clustering framework driven by representation learning, while the latter is a clustering prediction framework driven by dual-granularity feature learning.
[0012] 3. This patent emphasizes the consistent alignment of cluster distributions between branches and end-to-end updates in the training process. Its output is built upon an embedding space where local branches, semantic branches, and fusion branches converge. It focuses on solving problems such as boundary aliasing, semantic fragmentation, and difficulty in associating similar targets across regions in large-scale hyperspectral images. Patent CN117746079A, on the other hand, focuses on obtaining clustering prediction results through a pseudo-label strategy under unlabeled conditions. Its output is more biased towards land cover labels based on contrastive learning features and K-means clustering results. The former's data utilization emphasizes continuous optimization under semantic constraints, while the latter's data utilization emphasizes the combination of dual-granularity features and pseudo-label results.
[0013] III. A search revealed that Chinese invention patent CN117853769A provides a rapid clustering method for hyperspectral remote sensing images based on multi-scale image fusion, which differs from this application in the following ways:
[0014] 1. This patent focuses on constructing a high-precision self-supervised deep clustering framework for large-scale hyperspectral images. Its core lies in extracting local detail information through multi-scale convolution and enhancing global semantic consistency through a superpixel-guided semantic Transformer, thereby improving clustering accuracy and semantic integrity. In contrast, patent CN117853769A's core technology involves performing superpixel segmentation on the hyperspectral image, then executing superpixel-level graph convolutional subspace clustering to obtain a self-representation coefficient matrix. This coefficient matrix is then fused with multi-scale local spatial information, and finally, spectral clustering yields the final result. Essentially, it is a fast clustering scheme based on "graph convolutional subspace clustering + multi-scale fusion of coefficient matrices." The two patents differ fundamentally in their technical foundation and optimization objectives.
[0015] 2. This patent employs a strategy combining multi-scale convolutional feature extraction, semantic continuous attention, semantic global attention, and a three-branch self-supervised alignment loss. Its focus is on directly learning pixel-level embedding representations that balance local discriminability and global semantic consistency. In contrast, the method in patent CN117853769A emphasizes multi-scale local spatial information fusion of the self-representation coefficient matrix and relies on spectral clustering to output the final result. The former is a deep representation learning-driven self-supervised clustering framework, while the latter is a self-representation matrix-driven graph clustering and spectral clustering framework.
[0016] 3. This patent emphasizes improving clustering accuracy, boundary quality, and cross-regional semantic consistency in large-scale hyperspectral image scenarios. Its training process is an end-to-end update, and the output is the clustering result formed on a unified embedding space. Patent CN117853769A, on the other hand, emphasizes shortening clustering time and improving fast clustering performance under conditions of limited computing resources. Its technical effects are mainly reflected in efficiency improvement and computational cost control. The former focuses more on the semantic expressiveness of the clustering results, while the latter focuses more on the speed advantage and practical efficiency of the clustering process. Summary of the Invention
[0017] To address the problems existing in the prior art, this invention utilizes a multi-scale convolution module to extract local detail features of hyperspectral images, and simultaneously uses a semantic Transformer module based on superpixel priors to extract global semantic features with long-range dependencies. This invention proposes a novel self-supervised deep clustering method and system for hyperspectral images based on superpixel-guided semantic Transformer, which effectively improves the recognition accuracy of hyperspectral ground features.
[0018] To achieve the above objectives, the present invention is implemented through the following technical solution:
[0019] This invention is a self-supervised deep clustering method for hyperspectral images based on superpixel-guided semantic Transformer, specifically including the following steps:
[0020] Step 1: Preprocess the hyperspectral image to construct a 3D pixel block centered on each pixel.
[0021] Step 2: Input the 3D pixel block obtained in Step 1 into the ResNet module to obtain the shallow depth features of the hyperspectral image;
[0022] Step 3: Inject the shallow depth features obtained in Step 2 into the dual-path network module. The dual-path network module includes a multi-scale convolution (MSCB) module and a superpixel semantic Transformer module. The MSCB module extracts local detail features and the superpixel semantic Transformer module extracts global semantic features.
[0023] Step 4: Based on the local detail features obtained by the MSCB module in Step 3, use the kernel function to calculate the similarity between the feature vector of the local detail features and the cluster center vector of each category using the kmeans algorithm, to obtain the predicted distribution of each sample belonging to the category cluster, i.e., the local semantic probability distribution. Then, optimize the local detail feature representation by learning high confidence assignment to improve the cohesion of the category cluster and calculate the target distribution.
[0024] Step 5: Feed the global semantic features obtained by the superpixel semantic Transformer module in Step 3 into the feedforward neural network and use the Softmax function to obtain the global semantic probability distribution;
[0025] Step 6: Based on the predicted distribution and target distribution obtained in Step 4 and the global semantic probability distribution obtained in Step 5, construct a network loss function through a dual self-supervised module to guide the update of the entire local-global dual-branch network.
[0026] A further improvement of the present invention is that the specific process by which the superpixel semantic Transformer module extracts global semantic features is as follows:
[0027] Step 3.1: Convert the shallow features obtained in Step 2 into embedded representations through a 1×1 convolutional layer. Here, dim represents the dimensions of Q, K, and V in the attention mechanism, and is flattened into a two-dimensional vector sequence. And add learnable positional codes;
[0028] Step 3.2, First, the matrix passes through a LayerNorm layer, multiplied by three learnable weight matrices to obtain Q, K, and V, which are then divided into h parts along the channel dimension, as shown in the following formula:
[0029]
[0030]
[0031] Where h represents the number of heads in the multi-head self-attention;
[0032] Step 3.3: Construct an attention constraint set based on superpixel segmentation: Perform feature aggregation on each superpixel region to obtain a superpixel semantic representation, and construct a continuous semantic set of "intra-region representation + adjacent region representation" for each query position for Semantic Continuous Attention (SCA). Simultaneously, combine all superpixel semantic representations into a global set for Semantic Global Attention (SGA). Perform a dot product operation between the query of each head and the corresponding Key to obtain an attention score. To prevent the gradient vanishing problem, divide by a scale factor, and then normalize using Softmax to obtain the attention weight. The attention weight is multiplied by the Value, specifically as follows:
[0033]
[0034]
[0035] in, This represents a set of key / value pairs consisting of "within a superpixel region + adjacent regions". This represents a set of key / value pairs composed of global superpixel semantic representations. The dimension representing the key;
[0036] Step 3.4: After obtaining the attention results for each head in Step 3.3, the results of h heads are concatenated and then fused through a linear layer, as follows:
[0037]
[0038]
[0039] The outputs of the semantic continuous branch and the semantic global branch are then fused to obtain... :
[0040]
[0041] in, It is the projection matrix of the linear layer;
[0042] The results obtained in steps 3.5 and 3.4 After residual join, the specific representation is as follows:
[0043]
[0044] The results obtained in steps 3.6 and 3.5 After passing through a LayerNorm layer, MLP, and residual connections, the specific representation is as follows:
[0045]
[0046]
[0047] in, It is the final output of a semantic Transformer module, used to generate global semantic feature representations and participate in the construction of subsequent clustering prediction distributions.
[0048] The present invention also provides a hyperspectral image self-supervised deep clustering system based on superpixel guided semantic Transformer to implement the above method. The system includes a dual-path network module, a ResNet module, a decoding and reconstruction module, and a self-supervised optimization module.
[0049] The ResNet module includes a multi-layer residual structure, each of which includes a convolutional layer. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function. The 3D pixel block obtains shallow depth features through the ResNet module.
[0050] The MSCB module in the dual-path network module includes a multi-scale convolutional parallel branch and a channel adaptive recalibration structure. The multi-scale convolution is used to capture local spatial structure information under different receptive fields, and the channel recalibration is used to suppress spectral redundancy and enhance discriminability, where z is the number of output channels in the MSCB module.
[0051] The superpixel semantic Transformer module in the dual-path network module is composed of stacked multi-layer semantic Transformer models, and each layer of semantic Transformer introduces semantic continuous attention and semantic global attention to simultaneously enhance the semantic consistency of spatially continuous regions and the ability of cross-regional long-range semantic association.
[0052] The self-supervised module uses KL divergence as the cluster alignment loss and introduces reconstruction loss as a stability constraint, thereby jointly guiding the update of the entire local-global dual-branch network.
[0053] The beneficial effects of this invention are as follows: This invention utilizes the MSCB branch and the superpixel semantic Transformer branch to extract local detail features and global semantic features of hyperspectral images, respectively. Through a dual self-supervised mechanism, it uses KL divergence as a clustering alignment loss and combines it with reconstruction constraints to guide the training of the entire network model, thereby achieving unsupervised classification of hyperspectral images. Specifically:
[0054] (1) The present invention extracts shallow depth features of hyperspectral images through the ResNet module. The residual connection structure in ResNet can correlate features at different levels, thereby improving the performance of the model.
[0055] (2) This invention improves the recognition accuracy of hyperspectral ground features by introducing a local-global dual-path network architecture to mine local detail features of hyperspectral images and global features with consistent semantics across regions.
[0056] (3) This invention uses a dual self-supervised module and uses KL divergence clustering alignment loss and reconstruction loss for joint supervision to effectively supervise and guide the training and updating of the entire network model, forming an end-to-end learnable joint optimization network framework, which effectively improves the accuracy of hyperspectral land cover classification. Attached Figure Description
[0057] Figure 1 This is a flowchart of the hyperspectral image self-supervised deep clustering method using the superpixel-guided semantic Transformer of this invention.
[0058] Figure 2 This is a schematic diagram of the hyperspectral image self-supervised deep clustering method and the overall framework of the local-global dual-branch network of the present invention. In this diagram, 2(a) is the overall framework of STSDC, 2(b) is a schematic diagram of the multi-scale convolutional block MSCB structure, and 2(c) is a schematic diagram of the semantic Transformer structure (including semantic continuous attention SCA, semantic global attention SGA and feature enhancement structure). The connection relationship between the decoding and reconstruction module and the self-supervised alignment loss is also shown.
[0059] Figure 3 These are visualizations of clustering methods DEC, DEKM, IDEC, KDAE, DCN, BootSC, DEC-DA, DWSC, and the STSDC method of this invention on the MUUFL dataset. Among them, 3(a) is the actual ground feature distribution map of the MUUFL dataset, and 3(b) to 3(j) are clustering effect maps of different methods.
[0060] Figure 4 These are visualizations of clustering methods DEC, DEKM, IDEC, KDAE, DCN, DEC-DA, and the STSDC method of this invention on the Pavia University dataset. 4(a) is the actual ground feature distribution map of the Pavia University dataset, and 4(b) to 4(h) are clustering results of different methods.
[0061] Figure 5 These are visualizations of clustering methods DEC, DEKM, IDEC, KDAE, DCN, BootSC, DEC-DA, and the STSDC method of this invention on the Yangzhou dataset. Among them, 5(a) is the actual distribution map of ground features in the Yangzhou dataset, and 5(b) to 5(i) are clustering results of different methods. Detailed Implementation
[0062] The embodiments of the present invention will be disclosed below with reference to the drawings. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential.
[0063] This invention is a semantic Transformer self-supervised deep clustering method for large-scale hyperspectral images, denoted as STSDC.
[0064] like Figure 2 As shown, the STSDC local-global dual-branch network of the present invention specifically includes a ResNet18 backbone module, a multi-scale convolutional block (MSCB) module, a superpixel semantic Transformer module, a self-supervised optimization module, and a decoding and reconstruction module. The ResNet18 backbone module is used to extract shallow spatial-spectral joint features; the MSCB module is used to extract multi-scale local details and suppress spectral redundancy; the superpixel semantic Transformer module is used to simultaneously model spatially continuous semantics and cross-regional long-range semantics under superpixel prior guidance; the self-supervised optimization module is used to perform consistent alignment of the clustering distributions of the MSCB branch, the semantic Transformer branch, and the fusion branch; and the decoding and reconstruction module is used to reconstruct the features of the MSCB branch and introduce reconstruction constraints to stabilize training and avoid feature degradation caused solely by clustering objectives.
[0065] The MSCB module, as follows Figure 2 As shown in (b), multi-scale parallel convolution is used to capture local structural information under different receptive fields, and the information is fused in the channel dimension to obtain a multi-scale representation; preferably, the set of convolution kernel sizes is as follows: The multi-scale convolutional fusion process can be represented as:
[0066]
[0067] in, Output shallow features for ResNet18 and These represent the weights and biases of the convolution at the corresponding scale. This represents the convolution operation. It is a non-linear activation function; furthermore, MSCB introduces adaptive channel recalibration (Squeeze-and-Excitation) for... Channel weighting is performed to enhance discriminative power and suppress spectral redundancy, thereby obtaining the MSCB output characteristics. .
[0068] The superpixel semantic Transformer module, as follows: Figure 2As shown in (c), it includes superpixel segmentation prior, semantic continuous attention (SCA), semantic global attention (SGA), and feature enhancement structure; among which, the Entropy Rate Superpixel (ERS) superpixel segmentation algorithm is preferably used. Segmentation is performed to obtain A set of non-overlapping superpixel regions Furthermore, based on superpixel adjacency relationships, a semantic set of regions within and adjacent regions is constructed to reduce boundary aliasing caused by fixed window attention.
[0069] The SCA is used to enhance semantic consistency in spatially contiguous regions: for each superpixel Computational region center representation And construct a nearest neighbor set within the superpixel region. Simultaneously, a neighbor set is constructed in the adjacent superpixel region. The two are then combined to form a Key / Value set:
[0070]
[0071] The semantically continuous attention output is obtained by performing attention calculations on the query vector and the above Key / Value set:
[0072]
[0073] The SGA is used to associate spatially non-adjacent but semantically consistent regions: for each superpixel Average pooling is used to obtain superpixel-level semantic representations. and construct a global collection. As a global key / value pair, it yields semantic global attention output:
[0074]
[0075] The outputs of SCA and SGA are fused additively to obtain semantically enhanced features:
[0076]
[0077] The self-supervised optimization module uses a clustering target distribution to perform consistency alignment on the multi-branch prediction distribution: embedding the MSCB branches into the representation. With semantic Transformer branch embedding representation Element-by-element addition and fusion are performed to obtain And construct soft-assignment probabilities using the Student t distribution kernel function. :
[0078]
[0079] in, The cluster center vectors (initialized by Kmeans). The number of clusters. For the degree of freedom parameter, the preferred value is... Furthermore, the predicted distributions of the MSCB branches were analyzed separately. Semantic Transformer branch prediction distribution and fusion branch prediction distribution With target distribution Calculate the KL divergence and correlate it with the decoding and reconstruction loss. Together they form the total loss function:
[0080]
[0081] in, This is a weighting factor.
[0082] like Figure 1 As shown, the hyperspectral image self-supervised deep clustering method of the present invention specifically includes the following steps:
[0083] Step 1: Preprocess the hyperspectral image to construct the network input. Specifically: Step 1.1: Obtain the original hyperspectral image and represent it as an input tensor. ,in, and These are the height and width of the image, respectively. The number of spectral channels; Step 1.2, for the input tensor Normalization is performed, and a set of pixel samples is established according to pixel position, which serves as the input for subsequent feature extraction and clustering.
[0084] Step 2: Convert the input tensor obtained in Step 1 into a single input tensor. The shallow spatial-spectral joint features of the hyperspectral image are obtained by feeding them into the ResNet18 backbone module. .
[0085] Step 3: Obtain the shallow features from Step 2 Injected into the MSCB module to extract multi-scale local detail features, resulting in the MSCB branch output features. And obtain the local branch embedding representation through the branch prediction head. .
[0086] Step 4: The results obtained in Step 3 Perform superpixel segmentation to obtain a set of superpixel regions. It also considers the adjacency relationships of the superpixel region and constructs a Key / Value set for the SCA for each query pixel (including the semantic representation set within the superpixel region and the semantic representation set of its adjacent superpixel regions); at the same time, it performs feature aggregation on each superpixel region to obtain a superpixel-level semantic representation and constructs a global Key / Value set for the SGA.
[0087] Step 5: Apply the shallow features obtained in Step 2 The Key / Value set constructed in step 4 is fed into the Superpixel Semantic Transformer module to extract global semantic features and obtain semantic branch embedding representations. Specifically:
[0088] Step 5.1: In Semantic Continuous Attention (SCA), attention aggregation is performed only within the Key / Value set constructed in Step 4. Step 5.2: In the Semantic Global Attention (SGA), long-range semantic aggregation is performed using the global superpixel semantic representation set constructed in Step 4 as the Key / Value pair. Step 5.3, and The fusion yields semantically enhanced features, which are then output through a feature enhancement structure. .
[0089] Step 6: Embed the local branch representation obtained in Step 3. With the semantic branch embedding representation obtained in step 5 Element-wise addition and fusion are performed to obtain the fused embedding representation. And initialize the cluster center vectors based on the Kmeans algorithm. The similarity between the fused embedding representation and the cluster centers is calculated using the Student t-distribution kernel function to obtain the target distribution. .
[0090] Step 7: Embed the local branches from Step 3 into the representations respectively. Semantic branch embedding representation in step 5 and the fusion embedding representation in step 6 The prediction head is fed in and the prediction distribution is obtained using the Softmax function. , and At the same time, the result obtained in step 3 The data is fed into the decoding and reconstruction module for feature reconstruction to obtain the reconstruction loss. .
[0091] Step 8: Based on the target distribution obtained in Step 6 Compared with the predicted distribution obtained in step 7 , , A self-supervised alignment loss is constructed and combined with the reconstruction loss to form the total network loss function. This guides the end-to-end updates of the entire network; after training, the clustering label for each pixel is output, and the final unsupervised clustering result is obtained.
[0092] In this embodiment, the STSDC includes a ResNet18 backbone module, a three-layer MSCB module, a three-layer superpixel semantic Transformer module, and a three-layer... The decoding and reconstruction module consists of stacked convolutional layers; the Adam optimizer is preferably used during training, with a learning rate of 0.005 and 200 training iterations; the output feature dimension of ResNet18 is preferably set to [32, 48, 128] on different datasets; the number of superpixels... The preferred settings are [300, 1000, 1000]; the preferred settings for the loss term coefficient are... The specific experimental data are shown in Table 1-5:
[0093] Table 1
[0094]
[0095] Table 2
[0096]
[0097] Table 3
[0098]
[0099] Table 4
[0100]
[0101] Table 5
[0102]
[0103] Table 1 shows the quantitative evaluation results of clustering using different methods on the MUUFL dataset (ACC, ARI, AMI, NMI, FMI, and computational costs FLOPs and Params); Table 2 shows the quantitative evaluation results of clustering using different methods on the Pavia University dataset; Table 3 shows the quantitative evaluation results of clustering using different methods on the Yangzhou dataset; Table 4 shows the ablation experiment results of the effectiveness of the MSCB and Semantic Transformer modules; and Table 5 shows the comparative experiment results of different self-attention mechanisms.
[0104] As can be seen from Tables 1 to 3, the STSDC proposed in this invention achieves high overall clustering performance on the MUUFL, Pavia University, and Yangzhou datasets. Specifically, the ACC reaches 55.73% on the MUUFL dataset, 64.60% on the Pavia University dataset, and 57.05% on the Yangzhou dataset. At the same time, it also maintains advantages in consistency metrics such as ARI, AMI, NMI, and FMI, indicating that the joint mechanism of "multi-scale local structure + superpixel semantic attention + multi-branch consistency alignment" can effectively improve intra-cluster compactness and enhance inter-cluster separability. Furthermore, Table 4 shows that using only MSCB (STSDC-M) or only introducing SCA / SGA for one-sided semantic modeling (STSDC-C, STSDC-G) is difficult to achieve the performance of the complete STSDC, verifying the complementarity of SCA and SGA in modeling spatial continuous semantics and cross-regional semantic consistency. Table 5 shows that compared with attention / convolution mechanisms such as ScConv, RefConv, STViT and EfficientViT, the superpixel guided semantic attention of this invention can achieve higher ACC on all three types of datasets, demonstrating its effectiveness in large-scale hyperspectral clustering tasks.
[0105] Figures 3-5 Clustering visualization results are presented on the MUUFL, Pavia University, and Yangzhou datasets, respectively. Figure 3 As can be seen, the STSDC of this invention can effectively reduce boundary overlap and maintain internal consistency at complex terrain boundaries; by Figure 4 As can be seen, this invention provides more coherent clustering of areas such as roads, buildings, and shadows in urban and campus scenarios; Figure 5 As can be seen, the present invention can still maintain good regional consistency and cross-regional semantic aggregation capability in scenarios with fragmented spatial distribution and mixed categories, thereby obtaining clearer clustering partitions.
[0106] This invention introduces a local-global dual-path architecture of the MSCB branch and the superpixel semantic Transformer branch, and combines self-supervised alignment loss and decoding reconstruction constraints for end-to-end training, which enables effective clustering of large-scale hyperspectral images without the need for manual annotation.
[0107] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A self-supervised deep clustering method for hyperspectral images based on superpixel-guided semantic Transformer, characterized in that: Includes the following steps: Step 1: Preprocess the hyperspectral image, construct input samples centered on each pixel, and represent the original hyperspectral image as an input tensor; Step 2: Input the input samples obtained in Step 1 into the ResNet module to obtain the shallow spatial-spectral joint features of the hyperspectral image; Step 3: Inject the shallow spatial-spectral joint features obtained in Step 2 into the dual-path network module. The dual-path network module includes a multi-scale convolution module and a superpixel semantic Transformer module. The multi-scale convolution module extracts multi-scale local detail features and obtains the local branch embedding representation Z'. At the same time, the superpixel semantic Transformer module extracts global semantic features and obtains the semantic branch embedding representation Z″. Step 4: Embed the local branch representation With semantic branch embedding representation Element-wise addition and fusion are performed to obtain the fused embedding representation. And initialize the cluster center vectors based on the Kmeans algorithm. The fused embedding vector satisfies: ; Step 5: Embed the local branches into the representation respectively. Semantic branch embedding representation and fusion embedding representation The prediction head is fed in and the corresponding branch prediction distribution is obtained through the Softmax function. , and Simultaneously, the local features output by the MSCB module in step 3 are fed into the decoding and reconstruction module for feature reconstruction to obtain the reconstruction loss. ; Step 6: Based on the target distribution obtained in Step 4 Compared with the branch prediction distribution obtained in step 5 , , A network loss function is constructed through a self-supervised optimization module to guide the end-to-end update of the entire local-global dual-branch network.
2. The hyperspectral image self-supervised deep clustering method based on superpixel-guided semantic Transformer according to claim 1, characterized in that: The specific process of extracting global semantic features by the superpixel semantic Transformer module in step 3 is as follows: Step 3.1: Convert the shallow spatial-spectral joint features obtained in Step 2 into an embedded representation through linear mapping and flatten it into a two-dimensional vector sequence, while adding learnable positional encoding; Step 3.2: Perform superpixel segmentation on the embedding representation from Step 3.1 to obtain a set of superpixel regions. And its adjacency relationships; construct a Key / Value set for Semantic Continuous Attention (SCA) for each query position, including the representation set within the superpixel region. Set of adjacent superpixel regions Its construction formula is: ; Step 3.3, based on the construction in step 3.2 Calculate the Semantic Continuous Attention (SCA) output, whereby the formula for calculating SCA is: ; in, For query vector, The dimension of the Key; Step 3.4: Perform feature aggregation on each superpixel region to obtain a superpixel-level semantic representation set. and with the stated As the key / value set of the Semantic Global Attention (SGA), the output of the SGA is calculated. The formula for calculating the SGA is as follows: ; Step 3.5: Fuse the semantic continuous attention output obtained in Step 3.3 with the semantic global attention output obtained in Step 3.4, and obtain the semantic branch embedding representation through residual connections and a feedforward network. .
3. The hyperspectral image self-supervised deep clustering method based on superpixel-guided semantic Transformer according to claim 1, characterized in that: Step 4 also includes: The similarity between the fused embedding representation and the cluster center vector is calculated using the Student t-distribution kernel function to obtain the soft assignment prediction distribution of each sample to its category cluster. The calculation formula is as follows: ; And based on the soft allocation prediction distribution Construct the target distribution To enhance the clustering constraints of high-confidence samples, the calculation formula is as follows: ; in, For fusion embedding representation The Middle The embedding vector of each sample, The number of clusters. These are the degrees of freedom parameters.
4. A hyperspectral image self-supervised deep clustering system based on superpixel-guided semantic Transformer, used to implement the method as described in any one of claims 1-3, characterized in that: The system includes a dual-path network module, a ResNet module, a decoding and reconstruction module, and a self-supervised optimization module; The dual-path network module includes a multi-scale convolution module and a superpixel semantic Transformer module; The ResNet module includes a multi-layer residual structure, each residual structure includes a convolutional layer, and each convolutional layer is followed by a batch normalization layer and a ReLU activation function to obtain shallow spatial-spectral joint features for subsequent dual-path feature extraction. The self-supervised optimization module utilizes KL divergence on the target distribution. With branch prediction distribution , , Perform consistency alignment and compare it with the reconstruction loss generated by the decoding and reconstruction module. Together, we construct the total network loss function, which is: ; in, This is a weighting factor.
5. A hyperspectral image self-supervised deep clustering system based on superpixel-guided semantic Transformer according to claim 4, characterized in that: The multi-scale convolution module includes parallel multi-scale convolution branches and a channel adaptive recalibration structure. The parallel multi-scale convolution branches are used to extract local spatial structure information under different receptive fields, and the channel adaptive recalibration structure is used to suppress spectral redundancy and enhance discriminability. The multi-scale convolution outputs of the parallel multi-scale convolution branches are concatenated along the channel dimension to form a multi-scale representation, the calculation formula of which is: ; in, The shallow spatial-spectral joint features obtained in step 2, For the set of convolution kernel sizes, This represents the convolution operation. It is a non-linear activation function. This represents the multi-scale features after splicing.
6. The hyperspectral image self-supervised deep clustering system based on superpixel-guided semantic Transformer according to claim 4, characterized in that: The superpixel semantic Transformer module consists of multiple layers of semantic Transformer encoding layers stacked together. Each layer includes semantic continuous attention (SCA) and semantic global attention (SGA) to simultaneously model the semantic consistency of spatially continuous regions and long-range semantic associations across regions.
Citation Information
Patent Citations
Clustering prediction method and system of hyperspectral image, storage medium and equipment
CN117746079A
Hyperspectral remote sensing image rapid clustering method based on multi-scale image fusion
CN117853769A
Hyperspectral remote sensing image ground object clustering method of self-supervised double-branch Transform structure
CN118072059A