Self-supervised meta transfer learning hyperspectral target detection method and system
Through the self-supervised metatransfer learning framework, global-local spectral comparison learning and twin network fine-tuning, the problems of insufficient training samples and poor adaptability of complex scenes in hyperspectral target detection are solved, and efficient target detection effect is achieved.
Patent Information
- Application Number
- CN202510265457.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing hyperspectral target detection methods are inadequate training samples and poor adaptability in complex scenarios, especially in transfer learning, which makes it difficult to effectively identify and separate targets.
The self-supervised metatransfer learning framework is adopted, and the global-local spectral contrast learning (GLSL) module and the maximum distance triple (MDTriplet) loss function is pre-trained, combined with twin networks and contrast loss is fine-tuned, and the adaptive spatial spectral enhancement model is used for feature fusion and constraints to realize the model's transfer and target detection.
It improves the robustness and generalization ability of the model, can effectively identify and separate targets on different hyperspectral image data sets, and significantly improves the target detection effect.
Smart Images

Figure CN120147618A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and relates to a self-supervised meta-transfer learning hyperspectral target detection method and system. Background Art
[0002] In the process of continuous innovation of remote sensing technology, the spatial resolution and spectral resolution of remote sensing images collected by sensors have achieved qualitative leaps. Hyperspectral images are images generated by capturing the radiation data of a target scene within a large number of continuous wavelength ranges. These images are usually composed of hundreds or even thousands of continuous narrow bands, and each band corresponds to a specific region of the electromagnetic spectrum. The spectral resolution of hyperspectral images is very high, and usually the bandwidth of each band is narrow. This high resolution enables the precise differentiation of the characteristics of objects or scenes between different bands. Among them, hyperspectral target detection focuses on locating and identifying specific target pixels in hyperspectral images, aiming to separate the target of interest from various backgrounds, usually only requiring very little prior target spectral information. However, due to limited prior target knowledge and the existence of the two phenomena of "same spectrum, different objects" and "same object, different spectra", hyperspectral target detection faces long-term challenges.
[0003] In the early exploration stage, numerous methods have emerged in the field of hyperspectral target detection. Typical methods include minimizing the constrained energy with finite impulse response filters and target detection by projecting pixel signals onto the orthogonal subspace of each background endmember. However, these methods are difficult to utilize the non-linear characteristics of the spectrum and have poor adaptability to complex detection scenarios. Therefore, many researchers have proposed target detectors based on kernel methods, such as kernel orthogonal subspace projection, kernel-based constrained energy minimization, and kernel matched subspace detector. However, traditional machine learning models usually extract shallow features, which are limited in cases where the target and background are complex and non-linearly differentiable, and ideal detection results cannot be obtained. In recent years, deep learning, with its excellent feature extraction ability and powerful parallel computing ability, has shown broad application prospects and significant advantages in multiple fields of hyperspectral image processing. Its highly automated feature learning mechanism enables deep learning models to automatically extract hierarchical and discriminative feature representations from complex hyperspectral data. It is worth mentioning that deep learning is even more superior and efficient in hyperspectral target detection. Researchers have proposed a series of frameworks for hyperspectral target detection algorithms based on deep learning. For example, in order to reduce the loss of spectral information, some researchers have proposed a two-stream convolutional network framework. In addition, hyperspectral target detection algorithms based on interpretable representation networks, hyperspectral target detection algorithms based on deep metric learning, and hyperspectral target detection algorithms based on lightweight convolutional neural networks have also shown good performance. The above algorithms are usually trained and tested on the same image scene and rely on a small number of target priors to synthesize training samples, which requires a large amount of computing time. In addition, the training and testing of a single image result in poor generalization of these algorithms, making it difficult for the learned features to be transferred to a new dataset for target detection, which greatly limits the application scenarios of these algorithms. To alleviate the above problems, researchers have proposed some algorithms based on few-shot learning. For example, some researchers have proposed a semi-supervised domain adaptive few-shot learning detector. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a self-supervised meta-transfer learning hyperspectral target detection method and system, the core idea of which is to use a self-supervised meta-learning - transfer learning framework to achieve effective transfer from the source domain to the target domain. The method of the present invention consists of two modules: a self-supervised meta-transfer learning and pre-training framework and an adaptive spatial-spectral enhancement detector. First, in the meta-training stage, we use the labeled hyperspectral image classification data to randomly form multiple positive and negative sample pairs from different land covers, train the model to effectively distinguish the similarities and differences between spectra, and improve the model's sensitivity to spectral changes in hyperspectral images. Then, the global-local spectral contrast learning (GLSL) module and the maximum distance triplet (MDTriplet) loss are used to train the model to effectively distinguish spectral differences, and the pre-trained model is transferred to different target detection tasks and fine-tuned using a single target and background sample. Finally, an adaptive spatial-spectral enhancement model is adopted to jointly learn the spatial information constraint and the spectral information constraint to obtain the final detection result. To achieve the above object, the present invention provides the following technical solutions:
[0005] A self-supervised meta-transfer learning hyperspectral target detection method and system, the method comprising the following steps: S1: Preparation of hyperspectral image source domain data and target domain data; S2: Using the hyperspectral classification source data for self-supervised pre-training to obtain a global-local spectral contrast learning module (GLSL) that can effectively distinguish the similarities and differences between spectra; S3: Transferring the GLSL to the target detection dataset for spectral similarity detection to obtain a preliminary target detection result map; S4: Further using the spatial constraint learning module (SCLM) to complete the joint learning constraint of spatial information and spectral information to obtain the final target detection result.
[0006] Further, in step S1, preparing the hyperspectral image source domain data and the target domain data specifically includes: using a hyperspectral classification dataset containing rich ground object information for the source domain data. To enhance the model's ability to perceive spectral changes, contrast learning at the spectral level is carried out by constructing positive and negative sample pairs. Using the idea of meta-representation, two types of spectra are randomly selected in each task to form a set of positive and negative sample pairs, increasing the diversity of tasks. For the target domain dataset, 4 different hyperspectral target domain data are prepared for detection.
[0007] Further, in step S2, during the pre-training phase, in order to alleviate the problem of insufficient training samples for hyperspectral target detection, the proposed GLSL network is first pre-trained using an open-source labeled hyperspectral image classification dataset. To more effectively extract spectral feature information, the proposed GLSL network adopts an information extraction method that gradually delves from local to global. Specifically, the GLSL module consists of three branches: local vision (LV), global vision (GV), and spatial frequency fusion (SFFM). In spectral analysis, the subtle differences between adjacent bands often contain important information. Convolution can effectively extract the features on these adjacent bands, helping the model understand the interactions between different bands. Specifically, we designed the LV module to split the input spectral data into two parts, F1 and F2, for separate processing. F1 uses 1×1 convolution to adjust the dimension of the features and serves as a non-linear transformation to introduce additional non-linearity. F2 can extract multi-level local features from the spectral data by applying convolution kernels of different sizes layer by layer. Finally, the features obtained from parallel processing are fused, and a fully connected layer is used to map the extracted features to a lower-dimensional space to obtain the feature map L. Convolution operations can, to a certain extent, preserve the local details of the spectrum, but the spectral bands usually have dozens or even hundreds of bands, ignoring the global information. Therefore, in the GV module, we use the Transformer structure to establish global relationships and long-range dependencies through the self-attention mechanism. Specifically, the input spectral data is sequentially passed through normalization, multi-head self-attention (MSA), an MLP layer, and a fully connected layer to extract spectral features from multiple adjacent bands, and the residuals are concatenated and fused to reduce information loss from the shallow layer to the deep layer. Generally speaking, the GV module extracts and integrates global features by using the MSA mechanism and MLP components. Through the combination of these components, the module can capture the complex dependencies between different frequency bands when processing hyperspectral data, thus realizing the learning of global features. To further improve the quality of feature representation, the SFFM module was continuously designed, and the main goal is to integrate the local details and global context information of the spectrum. Therefore, in the SFFM, the one-dimensional (1D) input is reshaped into two-dimensional (2D) and feature extraction is performed through 2D convolution, which is beneficial to obtaining the local correlations in the spectral data and learning the potential relationships between originally non-adjacent spectral bands. Specifically, the one-dimensional spectrum 1×1×d is reshaped into 1×1×b×b, where b is the nearest number that can be square-rooted, and then passed into three 3×3 2D convolutional layers for spatial feature learning, and the ReLU activation function is used after each step. Then the 2D feature map is flattened into a 1D vector, and a fully connected layer is used to convert the flattened feature vector back to 1D data. Finally, a 1D convolutional layer is used to perform the final feature extraction on the data converted back to 1D. Through the above series of operations, the SFFM can effectively fuse the spatial and frequency domain features of the input spectrum, thereby improving the ability to capture high-dimensional features. Thus, enhancing the strong feature expression ability to obtain the feature map S.
[0008] Furthermore, in step S2, a method of maximum distance triplet (MDTriplet) loss is proposed. During the pre-training phase, contrastive learning between spectral beams is carried out using a siamese network. Two homogeneous spectra and one heterogeneous spectrum are sequentially input. For this purpose, a maximum distance triplet (MDTriplet) loss function is designed. First, the general Triplet loss is a loss function for learning contrastive feature representations and is commonly used to train contrastive learning models. The goal of Triplet loss is to make the distance between the anchor and the positive sample as small as possible, and the distance between the anchor and the negative sample as large as possible. The specific loss function of the general Triplet loss can be defined as (taking the Euclidean distance as an example):
[0009]
[0010] where F(A i )、F(P i )、F(G i ) represent the feature extraction functions applied to the anchor sample A i , the positive sample P i , and the negative sample G i . α is the margin parameter.
[0011] Although the general Triplet loss plays a certain role in classification tasks, if there are cases where the positive and negative samples are very close, the model may over-optimize these samples. Therefore, based on this, the distance between the positive and negative samples is considered simultaneously, and the MDTriplet Loss is designed. It aims to more comprehensively evaluate the performance of the model by comparing the distance between the anchor and the negative sample and the distance between the positive and negative samples, thereby improving the robustness and generalization ability of the model. The specific description of this MDTriplet Loss is as follows:
[0012]
[0013] where N is the number of samples in the batch.
[0014] Furthermore, in step S3, transfer learning is used to transfer the trained GLSL network to different hyperspectral target detection tasks. In the proposed method, considering the certain gap between the source domain and the target domain, the fine-tuning method in transfer learning is selected for transfer. At the same time, considering the scarcity of target spectra, only one target sample and one background sample are used for fine-tuning the model. Then, the sample to be measured and the target prior are input into the fine-tuned GLSL network for spectral feature enhancement, and the similarity between the enhanced spectrum and the prior knowledge is calculated to obtain the initial target detection result map.
[0015] Furthermore, in step S3, a Siamese network and contrastive loss are used for fine-tuning. Siamese neural networks perform well with a small amount of labeled data and are particularly suitable for contrastive learning tasks. Therefore, in the fine-tuning stage, we adopt the Siamese neural network structure, embed the pre-trained GLSL module into the network structure, and use a small sample to fine-tune the entire GLSL module. Specifically, randomly select one target and one background sample from the hyperspectral target detection dataset. We input the target or background sample and the target prior into the network for spectral feature enhancement, and calculate the similarity between the enhanced spectrum and the prior knowledge. Contrastive loss is used to measure the similarity between sample pairs and is commonly used to train deep learning models for similarity learning. It optimizes the model by minimizing the distance between similar sample pairs and maximizing the distance between dissimilar sample pairs. More importantly, the contrastive loss function can effectively learn useful features with a limited number of samples, reducing the need for a large amount of labeled data. Therefore, in the fine-tuning stage, we use contrastive loss to minimize the distance between positive samples and the target prior, and maximize the distance between negative samples and the target prior. Specifically, let the embedding vectors of the sample pair (x 1 , x 2 ) be z 1 and z 2 , where z i is the embedding representation output from the network, and calculate the similarity between positive and negative samples and the target prior:
[0016]
[0017] Next, the mathematical formula of the contrastive loss function is:
[0018]
[0019] where y is the label of the sample pair. If the sample pair belongs to the same class, then y = 1; if the sample pair belongs to different classes, then y = 0. Margin is a hyperparameter representing the minimum distance between dissimilar sample pairs. Here, N is the number of sample pairs, and N = 1. If the input is a positive sample, we hope this distance is as small as possible, otherwise we hope this distance is greater than margin.
[0020] Furthermore, an initial target detection result map is obtained in step S3. Through the combination of pre-training and fine-tuning, GLSL not only maintains excellent spectral resolution ability but also shows higher flexibility and adaptability in specific tasks. Specifically, after inputting the target prior t and the spectrum to be measured x into the GLSL module for feature enhancement, the extracted features f t and are used to calculate the cosine similarity between the spectrum to be measured and the target prior. Finally, a similarity map S is generated through these similarity values, effectively capturing and identifying subtle spectral differences:
[0021]
[0022] Where S(i,j) represents the similarity value at the position (i,j).
[0023] Furthermore, in step S4, the spatial features of the fusion target are combined to further optimize the target detection result map, achieving spatial-spectral joint constraints. In the learning stage of spatial-spectral joint constraints, in order to refine the target region and remove noise, a variety of techniques including neighborhood operations and morphological operations are adopted. In the neighborhood operation, by analyzing each target pixel and its neighborhood, the confidence of these pixels is dynamically adjusted to better identify and retain the true target region. Specifically, in combination with the spectral similarity map S, pixels with a confidence greater than r are set as target pixels, and the label is set to 1. The r value is set to 0.6. Taking the target pixel as the center point, the neighborhood pixels are analyzed. The specific operations of the analysis include traversing all target pixels and calculating the sum of labels in the small and large neighborhoods of the target pixel position (i,j). For example, in order to update the confidence of each target pixel point in the small range of a 3×3 neighborhood, we first calculate the sum N 3 (i,j). If N 3 (i,j) is less than a set threshold k 1 , it means that this point is more likely to be an isolated noise point, so the confidence level of this point needs to be reduced. Then, we use a decreasing exponential function q based on N 3 (i,j) to calculate the weight w 3 :
[0024]
[0025] For N 3 (i,j) < k 1 , the weights are applied to update the confidence map S:
[0026] S'(i,j) = max(S(i,j) - w 3 , 0)
[0027] Assume that the q function is a decreasing function. When the sum of labels in the neighborhood N 3 (i,j) is small, the calculated weight w 3 has a large value. This indicates that the target pixel is more likely to be an isolated noise point and requires a greater reduction in its confidence level. On this basis, the confidence level of each target pixel point is updated in a 7×7 large neighborhood. If the sum of target pixel points in the neighborhood of point N 7 (i,j) is greater than the set threshold k 2, it indicates that the neighborhood of the target pixel is likely to be a relatively large false target. Therefore, it is necessary to reduce the confidence level of this point. Then, use the incremental exponential function 1 - q to calculate the weight based on N 7 (i, j):
[0028]
[0029] For N 7 (i, j) > k 2 , apply weights to update the confidence map S':
[0030] S”(i, j) = max(S'(i, j) - w 7 , 0)
[0031] The incremental function is used because the larger N 7 (i, j) is, the more likely it is to become a false target, and more weights need to be reduced.
[0032] To optimize the boundary of the target region, remove the noise at the edges, and enhance the target segmentation effect, after each neighborhood operation, two morphological operations, erosion and dilation, are adopted. Specifically, first use a 3×3 kernel to perform an erosion operation on the target region. This operation helps to remove smaller noise points without affecting the larger target structure. Next, use a 7×7 kernel to perform a dilation operation on the eroded target region. The purpose of dilation is to connect adjacent target regions, making the segmented target more complete and coherent. Through the above operations, we can not only reduce false detections but also ensure that the boundary of the target region is more accurate. Finally, the updated similarity map is used as the target detection map.
[0033] The beneficial effects of the present invention are as follows:
[0034] The self-supervised meta-transfer learning hyperspectral target detection algorithm proposed by the present invention is divided into two stages: pre-training and spatial-spectral joint constraint, realizing the effective transfer of the network model, enhancing the model's ability to capture different features in hyperspectral data, enhancing the robustness and generalization ability of the network, and significantly improving the target detection effect. The experimental results on four hyperspectral image datasets compared with the effects of multiple SOTA methods show that the performance of the proposed method is superior to the current advanced hyperspectral image target detection methods.
[0035] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably below in conjunction with the accompanying drawings, where:
[0037] Figure 1 is the flowchart of the method of the present invention;
[0038] Figure 2 is the schematic diagram of self-supervised meta-transfer learning (SelfMTL) for hyperspectral target detection based on contrast representation;
[0039] Figure 3 is the structural diagram of the global-local spectral contrast learning network of the present invention;
[0040] Figure 4 are the visualization results of different target detection methods on four hyperspectral image datasets, where (a) ground truth map, (b) CEM, (c) HCEM, (d) CSCR, (e) DSC, (f) MLSN, (g) LCNN-CD, (h) HTD-IRN, (i) SelfMTL;
[0041] Figure 5 is the box plot of target-background separation of the present invention. Specific Embodiments
[0042] The technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings.
[0043] Figure 1 is the flowchart of the method of the present invention. The present invention provides a self-supervised meta-transfer learning hyperspectral target detection method and system. As shown in the figure, the source domain data uses a hyperspectral classification dataset containing rich ground object information. To enhance the model's ability to perceive spectral changes, contrast learning at the spectral level is performed by constructing positive and negative sample pairs. Using the idea of meta-representation, two types of spectra are randomly selected in each task to form a set of positive and negative sample pairs, increasing the diversity of tasks. For the target domain dataset, 4 different hyperspectral target domain data are prepared for detection. The global-local spectral contrast learning network is as Figure 2As shown, it can enhance the model's ability to capture different features in hyperspectral data. To alleviate the problem of insufficient sample training, we first pre-train the proposed GLSL network using an open-source labeled hyperspectral classification dataset. Subsequently, we adopt the MDTriplet loss and input the constructed positive and negative sample pairs into the GLSL module for spectral-level contrast learning. In addition, to effectively transfer the GLSL module that can distinguish significant spectral differences to various target detection tasks, we use two positive and negative sample spectra from the target and the background (one from the target and one from the background) to fine-tune the model parameters. In the test phase, we input the prior target and the spectra to be tested into the GLSL module to obtain a feature-enhanced spectral beam representation. Then, we calculate the similarity between these spectral beam features and use it to form a spectral similarity map. Finally, to make full use of the spatial information of the target, the GSLS algorithm adopts an adaptive spectral-spatial enhancement module to optimize the final target detection result map by using spatial-spectral constraints learning. The self-supervised meta-transfer learning framework designed in the present invention effectively solves the problem of limited training samples in the mode of randomly constructing positive and negative sample pairs. The adaptive spectral-spatial enhancement module designed in the present invention fully integrates the spatial and spectral information in the hyperspectral image. Specifically, the technical solution of the present invention includes the following content:
[0044] 1. Source domain and target domain data preparation: The source domain data uses a hyperspectral classification dataset containing rich ground object information. To enhance the model's ability to perceive spectral changes, contrast learning at the spectral level is performed by constructing positive and negative sample pairs. Using the idea of meta-representation, two types of spectra are randomly selected in each task to form a set of positive and negative sample pairs, increasing the diversity of the tasks. For the target domain dataset, 4 different hyperspectral target domain data are prepared for detection.
[0045] 2. Self-supervised pre-training: As Figure 3As shown in the figure, in the pre-training stage, in order to alleviate the problem of insufficient training samples for hyperspectral target detection, the proposed GLSL network is first pre-trained using an open-source labeled hyperspectral image classification dataset. In order to extract spectral feature information more effectively, the proposed GLSL network adopts a method of gradually deepening information extraction from local to global. Specifically, the GLSL module consists of three branches: Local Vision (LV), Global Vision (GV), and Spatial Frequency Fusion Module (SFFM). In spectral analysis, the subtle differences between adjacent bands often contain important information. Convolution can effectively extract the features on these adjacent bands, helping the model understand the interactions between different bands. Specifically, we designed the LV module to split the input spectral data into two parts, F1 and F2, for separate processing. F1 uses 1×1 convolution to adjust the dimension of the features and serves as a non-linear transformation to introduce additional non-linearity. F2 can extract multi-level local features from the spectral data by applying convolution kernels of different sizes layer by layer. Finally, the features obtained from parallel processing are fused, and a fully connected layer is used to map the extracted features to a lower-dimensional space to obtain the feature map L. The convolution operation can preserve the local details of the spectrum to a certain extent, but the spectral bands usually have dozens or even hundreds of bands, ignoring the global information. Therefore, in the GV module, we use the Transformer structure to establish global relationships and long-range dependencies through the self-attention mechanism. Specifically, the input spectral data is passed through normalization, multi-head self-attention (MSA), MLP layer, and fully connected layer in sequence to extract spectral features from multiple adjacent bands, and the residuals are concatenated and fused to reduce the information loss from the shallow layer to the deep layer. Generally speaking, the GV module extracts and integrates global features by using the MSA mechanism and MLP components. Through the combination of these components, the module can capture the complex dependencies between different frequency bands when processing hyperspectral data, thus realizing the learning of global features. In order to further improve the quality of feature representation, the SFFM module is designed. The main goal is to integrate the local details and global context information of the spectrum. Therefore, in the SFFM, the one-dimensional (1D) input is reshaped into two-dimensional (2D) and feature extraction is performed through 2D convolution, which is beneficial to obtaining the local correlations in the spectral data and learning the potential relationships between originally non-adjacent spectral bands. Specifically, the one-dimensional spectrum 1×1×d is reshaped into 1×1×b×b, where b is the nearest number that can be square-rooted, and then passed into three 3×3 2D convolution layers for spatial feature learning, and the ReLU activation function is used after each step. Then the 2D feature map is flattened into a 1D vector, and a fully connected layer is used to convert the flattened feature vector back to 1D data. Finally, a 1D convolution layer is used to perform the final feature extraction on the data converted back to 1D. Through the above series of operations, the SFFM can effectively fuse the spatial and frequency domain features of the input spectrum, thereby improving the ability to capture high-dimensional features. Thus, the strong feature expression ability is obtained, and the feature map S is obtained.
[0046] 3. Maximum Distance Triplet Loss Function: As Figure 2 shown, a method of Maximum Distance Triplet (MDTriplet) loss is proposed. During the pre-training stage, contrastive learning between spectral beams is carried out using a Siamese network. Two homogeneous spectra and one heterogeneous spectrum are sequentially input. For this purpose, a Maximum Distance Triplet (MDTriplet) loss function is designed. First of all, the general Triplet loss is a loss function for learning contrastive feature representations and is often used to train contrastive learning models. The goal of Triplet loss is to make the distance between the anchor and the positive sample as small as possible, and the distance between the anchor and the negative sample as large as possible. The specific loss function of the general Triplet loss can be defined as (taking the Euclidean distance as an example):
[0047]
[0048] where F(A i )、F(P i )、F(G i ) represent the feature extraction functions applied to the anchor sample A i , the positive sample P i and the negative sample G i . α is the margin parameter.
[0049] Although the general Triplet loss plays a certain role in classification tasks, if there is a situation where the positive and negative samples are very close, the model may over-optimize these samples. Therefore, based on this, the distance between the positive and negative samples is considered simultaneously, and the MDTriplet Loss is designed. The aim is to more comprehensively evaluate the performance of the model by comparing the distance between the anchor and the negative sample and the distance between the positive and negative samples, so as to improve the robustness and generalization ability of the model. The specific description of this MDTriplet Loss is as follows:
[0050]
[0051] where N is the number of samples in the batch.
[0052] 4. Use transfer learning to transfer the trained GLSL network to different hyperspectral target detection tasks. In the proposed method, considering the certain gap between the source domain and the target domain, we choose the fine-tuning method in transfer learning for transfer. At the same time, considering the scarcity of target spectra, only one target sample and one background sample are used to fine-tune the model. Then, the sample to be measured and the target prior are input into the fine-tuned GLSL network for spectral feature enhancement, and the similarity between the enhanced spectrum and the prior knowledge is calculated to obtain the initial target detection result map.
[0053] 5. Use a Siamese network and contrastive loss for fine-tuning. Siamese neural networks perform well in the case of a small amount of labeled data and are especially suitable for contrastive learning tasks. Therefore, in the fine-tuning stage, we adopt the Siamese neural network structure, embed the trained GLSL module into the network structure, and use a small sample to fine-tune the entire GLSL module. Specifically, randomly select one target and one background sample from the hyperspectral target detection dataset. We input the target or background sample and the target prior into the network for spectral feature enhancement, and calculate the similarity between the enhanced spectrum and the prior knowledge. Contrastive loss is used to measure the similarity between sample pairs and is often used to train deep learning models for similarity learning. It optimizes the model by minimizing the distance between similar sample pairs and maximizing the distance between dissimilar sample pairs. More importantly, the contrastive loss function can effectively learn useful features in the case of a limited number of samples, reducing the need for a large amount of labeled data. Therefore, in the fine-tuning stage, we use contrastive loss to minimize the distance between the positive sample and the target prior and maximize the distance between the negative sample and the target prior. Specifically, let the embedding vectors of the sample pair (x 1 , x 2 ) be z 1 and z 2 , where z i is the embedding representation output from the network, and calculate the similarity between the positive and negative samples and the target prior:
[0054]
[0055] Then, the mathematical formula of the contrastive loss function is:
[0056]
[0057] where y is the label of the sample pair. If the sample pair belongs to the same class, then y = 1; if the sample pair belongs to different classes, then y = 0. margin is a hyperparameter representing the minimum distance between dissimilar sample pairs. Here N is the number of sample pairs, and N = 1. If the input is a positive sample, we hope this distance is as small as possible, otherwise we hope this distance is greater than margin.
[0058] 6. Obtain the initial target detection result map. Through the combination of pre-training and fine-tuning, GLSL not only maintains excellent spectral resolution ability but also shows higher flexibility and adaptability in specific tasks. Specifically, after inputting the target prior t and the spectrum x to be measured into the GLSL module for feature enhancement, the extracted feature f t and are used to calculate the cosine similarity between the spectrum to be measured and the target prior. Finally, a similarity map S is generated through these similarity values, effectively capturing and identifying subtle spectral differences:
[0059]
[0060] where S(i,j) represents the similarity value at position (i,j).
[0061] 7. Fuse the spatial features of the target, further optimize the target detection result map, and achieve spatial-spectral joint constraints. In the learning stage of spatial-spectral joint constraints, in order to refine the target area and remove noise, a variety of techniques including neighborhood operations and morphological operations are adopted. In neighborhood operations, by analyzing each target pixel and its neighborhood, the confidence of these pixels is dynamically adjusted to better identify and retain the real target area. Specifically, combined with the spectral similarity map S, pixels with a confidence greater than r are set as target pixels and the label is set to 1. The r value is set to 0.6, and the neighborhood pixels are analyzed with the target pixel as the center point. The specific operations of the analysis include traversing all target pixels and calculating the sum of labels in the small and large ranges of the neighborhood of the target pixel position (i,j). For example, in order to update the confidence of each target pixel point in the small range of a 3×3 neighborhood, we first calculate the sum N 3 (i,j). If N 3 (i,j) is less than a set threshold k 1 , it means that this point is more likely to be an isolated noise point, so the confidence level of this point needs to be reduced. Then, we use a decreasing exponential function q based on N 3 (i,j) to calculate the weight w 3 :
[0062]
[0063] For N 3 (i,j) < k 1 , apply the weight to update the confidence map S:
[0064] S'(i,j) = max(S(i,j) - w 3 , 0)
[0065] Assume that the q function is a decreasing function. When the neighborhood N3 When the sum of the labels in (i,j) is small, the calculated weight w 3 has a larger value. This indicates that the target pixel is more likely to be an isolated noise point and requires a greater reduction in its confidence level. On this basis, the confidence level of each target pixel is updated in a large 7×7 neighborhood. If the point N 7 the sum of the target pixel points within the neighborhood of (i,j) is greater than the set threshold k 2 , it means that the neighborhood of the target pixel is likely to be a larger false target, so the confidence level of this point needs to be reduced. Then, the incremental exponential function 1-q is used to calculate the weight based on N 7 (i,j):
[0066]
[0067] For N 7 (i,j) > k 2 , the weight is applied to update the confidence map S':
[0068] S”(i,j) = max(S'(i,j) - w 7 , 0)
[0069] The incremental function is used because the larger N 7 (i,j) is, the more likely it is to become a false target and the more weight needs to be reduced.
[0070] To optimize the boundary of the target region, remove the noise at the edges, and enhance the target segmentation effect, after each neighborhood operation, two morphological operations, erosion and dilation, are adopted. Specifically, first, a 3×3 kernel is used to perform an erosion operation on the target region. This operation helps to remove smaller noise points without affecting larger target structures. Next, a 7×7 kernel is used to perform a dilation operation on the eroded target region. The purpose of dilation is to connect adjacent target regions, making the segmented target more complete and coherent. Through the above operations, we can not only reduce false detections but also ensure that the boundary of the target region is more accurate. Finally, the updated similarity map is used as the target detection map.
[0071] 8. The input sample is discriminated by the trained target detection framework, and the target detection result map is output. As Figure 4This is the comparison result graph of the SelfMTL hyperspectral target detection network described in the present invention with the present invention method and other existing methods CEM, HCEM, CSCR, DSC, MLSN, LCNN-CD, and HTD-IRN methods on the hyperspectral natural scene dataset. It can be seen that the target area is well detected. At the same time, the area under the curve (AUC) value is calculated. Among them, the larger the area under the curve (AUC) value, the better the comprehensive evaluation of the target detection result. Table 1 shows the index results of different methods in four different test sets:
[0072] Table 1 Comparison of SelfMTL and various methods on four hyperspectral image datasets (average value)
[0073]
[0074] The method of the present invention achieves the best accuracy on most datasets. Figure 5 The box plot of target-background separation of the SelfMTL method and other comparison methods on four real hyperspectral image datasets is given. It can be seen that the target samples generated by the method described in the present invention achieve better separation between the target and the background to a great extent. The method proposed by the present invention can effectively highlight the target and suppress the background, and can be efficiently migrated to different target detections.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A self-supervised meta-transfer learning hyperspectral target detection method, characterized in that The following steps are involved: S1: Initial source domain and target domain data preparation for hyperspectral images; S2: Self-supervised pre-training of GLSL network using source data; S3: Migrate the GLSL network to the target detection dataset for spectral similarity detection; S4: Use the spatial constraint module to learn constraints and obtain target detection results.
2. The self-supervised meta-transfer learning hyperspectral target detection method according to claim 1, characterized in that: In step S1, the hyperspectral image source domain data and target domain data are prepared, specifically including: the source domain data uses a hyperspectral classification dataset containing rich ground object information. In order to enhance the model's ability to perceive spectral changes, spectral level comparative learning is performed by constructing positive and negative sample pairs. Using the idea of meta-representation, two types of spectra are randomly selected in each task to form a set of positive sample pairs and negative sample pairs to increase the diversity of the task. For the target domain dataset, 4 different hyperspectral target domain data are prepared for detection.
3. The self-supervised meta-transfer learning hyperspectral target detection method according to claim 1, characterized in that: In step S2, in the pre-training stage, in order to alleviate the problem of insufficient training samples for hyperspectral target detection, the proposed GLSL network is first pre-trained using an open source labeled hyperspectral image classification dataset. In order to more effectively extract spectral feature information, the proposed GLSL network adopts a gradually deepening information extraction method from local to global. Specifically, the GLSL network consists of three branches: local vision (LV), global vision (GV), and spatial frequency fusion (SFFM). The local vision (LV) divides the input spectral data into two parts, F1 and F2, and processes them separately. F1 uses 1×1 convolution to adjust the dimension of the feature and introduces additional nonlinearity as a nonlinear transformation. F2 can extract multi-level local features from the spectral data by applying convolution kernels of different sizes layer by layer. Finally, the features obtained by parallel processing are fused, and the extracted features are mapped to a lower dimensional space using a fully connected layer to obtain the feature map L; in the global vision (GV), we use the Transformer structure to automatically The attention mechanism establishes global relationships and long-range dependencies. Specifically, the input spectral data is sequentially normalized, multi-head self-attention (MSA), MLP layer and fully connected layer to extract spectral features from multiple adjacent bands, and the residuals are spliced and fused to reduce the information loss from shallow to deep layers. Therefore, in space-frequency fusion (SFFM), the one-dimensional (1D) input is reorganized into two-dimensional (2D) and features are extracted through 2D convolution. Specifically, the one-dimensional spectrum is reshaped into 1×1×b×b, where b is the nearest square root, and then passed into three 2D convolution layers for spatial feature learning. After each step, the ReLU activation function is used, and then the 2D feature map is flattened into a 1D vector. The flattened feature vector is converted back to 1D data using a fully connected layer, and finally the 1D convolution layer is used to perform the final feature extraction on the converted 1D data. Through the above series of operations, SFFM can effectively fuse the spatial and frequency domain features of the input spectrum, thereby improving the ability to capture high-dimensional features, thereby strengthening the feature expression ability and obtaining the feature map S.
4. The self-supervised meta-transfer learning hyperspectral target detection method according to claim 3 is characterized in that: In step S2, a maximum distance triplet (MDTriplet) loss method is proposed. In the pre-training stage, the twin network is used to perform comparative learning between spectral beams. Two similar spectra and one heterogeneous spectrum are sequentially input. For this purpose, a maximum distance triplet (MDTriplet) loss function is designed. The specific description of the MDTriplet Loss is as follows: Where N is the number of samples in the batch, F(A i )、F(P i )、F(G i ) indicates that it is applied to anchor sample A i , positive sample P i And negative samples G i The feature extraction function of , α is the marginal parameter.
5. The self-supervised meta-transfer learning hyperspectral target detection method according to claim 1, characterized in that: Firstly, transfer learning is used to migrate the trained GLSL network to different hyperspectral target detection tasks. In the proposed method, we take into account the certain gap between the source domain and the target domain, so we choose the fine-tuning method in transfer learning for migration. At the same time, considering the scarcity of target spectrum, we only use one target sample and one background sample to fine-tune the model. Then, the sample to be tested and the target prior are input into the fine-tuned GLSL network for spectral feature enhancement. The similarity between the enhanced spectrum and the prior knowledge is calculated to obtain the initial target detection result map.
6. The self-supervised meta-transfer learning hyperspectral target detection method according to claim 5, characterized in that: Fine-tuning is performed using a twin network and contrast loss. In the fine-tuning stage, we use a twin neural network structure to embed the trained GLSL module into the network structure, and use a small sample to fine-tune the entire GLSL module. Specifically, we randomly select a target and a background sample from the hyperspectral target detection dataset. We input the target or background sample and the target prior into the network for spectral feature enhancement, and calculate the similarity between the enhanced spectrum and the prior knowledge. Contrastive loss is used to measure the similarity between sample pairs and is often used to train deep learning models for similarity learning. It optimizes the model by minimizing the distance between similar sample pairs and maximizing the distance between heterogeneous sample pairs. More importantly, the contrast loss function can effectively learn useful features with limited samples, reducing the need for a large amount of labeled data. Therefore, in the fine-tuning stage, we use contrastive loss to minimize the distance between positive samples and target priors and maximize the distance between negative samples and target priors. Specifically, let the embedding vectors of the sample pair (x1, x2) be z1 and z2, where z i It is the embedding representation output from the network, calculating the similarity between positive and negative samples and the target prior: Next, the mathematical formula of the contrast loss function is: Where y is the label of the sample pair. If the sample pair belongs to the same class, y = 1; if the sample pair belongs to different classes, y = 0. Margin is a hyperparameter, which indicates the minimum distance between heterogeneous sample pairs. N is the number of sample pairs, where N = 1. If the input is a positive sample, we hope that this distance is as small as possible, otherwise we hope that this distance is greater than margin.
7. The self-supervised meta-transfer learning hyperspectral target detection method according to claim 6, characterized in that: The initial target detection result map is obtained. Through the combination of pre-training and fine-tuning, GLSL not only maintains excellent spectral resolution ability, but also shows higher flexibility and adaptability in specific tasks. Specifically, after the target prior and the spectrum to be measured are input into the GLSL module for feature enhancement, the extracted feature f t and It is used to calculate the cosine similarity between the measured spectrum and the target prior. Finally, these similarity values are used to generate a similarity graph S, which can effectively capture and identify subtle spectral differences: Where S(i,j) represents the similarity value at position (i,j).
8. The self-supervised meta-transfer learning hyperspectral target detection method according to claim 1, characterized in that: The spatial features of the target are integrated to further optimize the target detection result map and realize the spatial-spectral joint constraint. In the learning stage of the spatial-spectral joint constraint, in order to refine the target area and remove noise, a variety of techniques including neighborhood operations and morphological operations are used. In the neighborhood operation, each target pixel and its neighborhood are analyzed, and the confidence of these pixels is dynamically adjusted to better identify and retain the real target area. Specifically, combined with the spectral similarity map S, the pixels with confidence greater than r are set as target pixels, the label is set to 1, the r value is set to 0.6, and the target pixel is analyzed as the center point. Neighborhood pixel points, the specific operations of the analysis include traversing all target pixels, calculating the sum of labels in the small range and large range neighborhood of the target pixel position (i, j). For example, in order to update the confidence of each target pixel in the small range of the 3×3 neighborhood, we first calculate the sum of the target pixels in the 3×3 neighborhood of the point N3(i, j). If N3(i, j) is less than a set threshold k1, it means that the point is more likely to be an isolated noise point, so the confidence level of the point needs to be reduced. Then, we use a decreasing exponential function q based on N3(i, j) to calculate the weight w3: For N3(i,j)<k1, apply the weights to update the confidence map S: S'(i,j)=max(S(i,j)-w3,0) Assuming that the q function is a decreasing function, when the sum of the labels in the neighborhood N3(i,j) is small, the calculated weight w3 is large, which indicates that the target pixel is more likely to be an isolated noise point and needs to reduce its confidence level. On this basis, the confidence level of each target pixel is updated in a large 7×7 neighborhood. If the sum of the target pixels in the neighborhood of point N7(i,j) is greater than the set threshold k2, it means that the neighborhood of the target pixel is likely to be a large pseudo target, so the confidence level of the point needs to be reduced. Then, the incremental exponential function 1-q is used to calculate the weight based on N7(i,j): For N7(i,j)>k2, apply the weights to update the confidence map S': S”(i,j)=max(S'(i,j)-w7,0) The incremental function is used because the larger N7(i,j) is, the more likely it is to become a pseudo target and the more weight needs to be reduced. In order to optimize the boundary of the target area and remove the noise at the edge while enhancing the target segmentation effect, after each neighborhood operation, two morphological operations, erosion and dilation, are used. Specifically, first, a 3×3 kernel is used to erode the target area. This operation helps to remove smaller noise points without affecting the larger target structure. Next, a 7×7 kernel is used to dilate the eroded target area. The purpose of dilation is to connect adjacent target areas, so that the segmented target is more complete and coherent. Through the above operations, we can not only reduce false detections, but also ensure that the boundaries of the target area are more accurate. Finally, the updated similarity map is used as the target detection map.
9. A self-supervised meta-transfer learning hyperspectral target detection system, characterized by: The system is configured with a control program for implementing the self-supervised meta-transfer learning hyperspectral target detection method described in any one of claims 1-8.
Citation Information
Patent Citations
Real-time infrared dynamic scene simulation method for multiple objects in sea and sky background
CN103186906A
Hyperspectral image classification method based on deep transfer learning
CN113705580A
Generative self-supervision hyperspectral image target detection method based on spatial spectrum mask
CN117746235A
Contrast type self-supervision hyperspectral image target detection method based on dual-path network
CN117911673A
Unsupervised Latent Low-Rank Projection Learning Method for Feature Extraction of Hyperspectral Images
US20230114877A1
Cited By
Steel wire material rapid classification method based on comparative learning
CN122067022A