Robust HSI classification composite deep network construction method

By adopting a hybrid architecture method of dual-stream network and CVT in hyperspectral remote sensing image classification, combining spectral and spatial empowerment strategies, the fuzzy supervision function is introduced, which solves the problems of inaccurate feature extraction and noise sensitivity in hyperspectral remote sensing image classification, and achieves a more efficient and robust image classification effect.

CN120068977APending Publication Date: 2025-05-30GUILIN UNIV OF ELECTRONIC TECH +1

Patent Information

Application Number
CN202510190219.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

There are high-dimensionality, redundancy and spectral uncertainty problems in hyperspectral remote sensing image classification, resulting in inaccurate feature extraction. Traditional machine learning and deep learning methods such as convolutional neural networks are limited in performance in this field, making it difficult to effectively capture complex patterns and subtle differences, and are also sensitive to noise.

Method used

A robust HSI classification composite deep network construction method is proposed, using a dual-stream network structure, combining 1D CNN and 2D CNN, spectral empowerment strategy and spatial empowerment strategy are introduced, feature mapping and convolutional encoding are performed through the CVT hybrid architecture, and fuzzy supervision function is introduced to suppress the influence of noise samples.

Benefits of technology

The induction of local depth features and extraction of long-distance feature dependencies is realized, which enhances the robustness and noise resistance of the network, and improves the classification accuracy of hyperspectral images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing, and particularly relates to a robust HSI classification composite deep network construction method, which comprises the following steps: S1, based on an inaccurate label introduction mechanism, defining noise types and proportions, and creating a noise label training set; s2, local spectrum mode extraction and discriminative spectrum feature enhancement are carried out in the spectrum flow branches; s3, receiving the data subjected to spectral processing by a spatial stream branch, and extracting and stacking spatial features of the data along a spectral dimension; s4, performing feature mapping and convolutional coding by adopting a CVT hybrid architecture, and modeling a depth relationship of local-global semantic mark features; and S5, introducing a fuzzy supervision function so as to dynamically adjust the loss value according to the prediction probability and pay attention to the samples which are difficult to classify correctly. The method can solve the problems of semantic feature extraction and noise training sample information suppression in deep learning, further improves the precision of hyperspectral image surface feature interpretation, and has a good market application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing, and particularly relates to a method for constructing a robust HSI classification composite deep network. Background Art

[0002] Hyperspectral remote sensing can detect targets at a long distance in a specific narrow channel, obtain continuous spectral characteristics at a fine spectral scale, contain rich spectral information, and contribute to the accurate identification of ground objects and improve the accuracy of fine classification of ground objects. Since the advent of hyperspectral remote sensing technology, it has received extensive attention in the scientific community and become a research hotspot. Various hyperspectral sensors have emerged like mushrooms after a spring rain, and corresponding image processing and information extraction methods have also developed by leaps and bounds.

[0003] Hyperspectral remote sensing image classification, as an important part of the application of remote sensing technology, is an important means to obtain ground object information. With the development of modern society, many fields represented by resource exploration, urban planning, environmental monitoring, disaster prevention and mitigation, etc. have put forward higher requirements for the classification accuracy of remote sensing images. However, affected by the characteristics of hyperspectral data itself, there are still many difficulties in its image classification: high dimensionality, redundancy, and spectral uncertainty lead to inaccurate feature extraction; the quantity of prior knowledge affects classification accuracy and network generalization ability; the quality of prior knowledge affects modeling performance and network inference results.

[0004] Specifically, traditional machine learning has limited performance in hyperspectral image classification. They mainly rely on manually designed shallow features and it is difficult to achieve good generalization. Deep learning methods, due to their powerful hierarchical feature representation ability and automatic learning characteristics, can effectively capture the complex patterns and subtle differences in hyperspectral images (HSIs). For example, convolutional neural networks are widely used in HSI classification, but they also face the following dilemmas. One-dimensional or two-dimensional convolutional neural networks (CNNs) can only use spatial or spectral features separately for classification, resulting in feature loss and making it difficult to handle complex ground object types and phenomena such as "same object with different spectra" and "same spectrum with different objects". For large-scale datasets, three-dimensional CNNs require very high time and computing resources for training and inference; they have a large number of parameters, which not only increases the complexity of the network but may also lead to overfitting problems. As the convolutional layer deepens, CNNs can extract deep abstract features, but the accompanying problems are vanishing or exploding gradients. Although CNNs can capture local patterns in images through local receptive fields and weight sharing mechanisms, they still cannot capture both fine-grained local details and large-scale global context information at the same time. In addition, CNNs have high requirements for data quality and usually require high-quality and accurately labeled data. Any noise or inaccurate labeling may seriously affect the performance of the network, and enhancing the noise resistance of the network often requires sacrificing some feature extraction capabilities.

[0005] The information disclosed in this background section is only intended to enhance the overall understanding of the present invention and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention

[0006] The object of the present invention is to provide a method for constructing a robust HSI classification composite deep network, aiming to provide a lightweight and efficient deep learning network for hyperspectral image classification, realizing local deep feature induction and long-distance feature dependency extraction, and at the same time embedding a sample learning constraint strategy to suppress the expression of noise sample knowledge, so that the network is robust in both clean and noisy environments.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A method for constructing a robust HSI classification composite deep network, comprising the following steps:

[0009] S1. Based on an inaccurate label introduction mechanism, define the proportion of noise labels to be introduced in each dataset, and modify the original labels according to the selected noise type and proportion to create a noise label training set;

[0010] S2. In the two-stream network, a spectral weighting strategy is introduced in the spectral stream branch, allowing the network to dynamically assign weights to different bands for local spectral pattern extraction and discriminative spectral feature enhancement;

[0011] S3. The spatial stream branch receives the spectrally processed data, extracts and stacks spatial features along the spectral dimension, and obtains output features with enhanced spatial features through a spatial weighting strategy;

[0012] S4. The feature extraction results of the two-stream network are input into the CVT hybrid architecture for feature mapping and convolutional encoding, and the depth relationship of local-global semantic labeled features is modeled. A module with the same architecture as the mapping layer is used to perform output projection on the features to obtain the class information of the image patches;

[0013] S5. A fuzzy supervision function is introduced to convert hard labels into soft labels, reducing the network's overconfidence in specific labels. In each training batch, the loss value is dynamically adjusted according to the prediction probability to suppress the expression of mislabeled sample information.

[0014] Preferably, the inaccurate label introduction mechanism described in S1 is specifically as follows:

[0015] First, for each class, randomly select n samples with true labels to form an initial training set; then, uniformly select samples from other classes and label their labels as the current class; then, mix all the samples in each class to ensure that the noisy samples can be fully integrated into the initial training set; then, a C-class classification training set with noisy labels can be defined as:

[0016]

[0017] Preferably, in the two-stream network described in S2, the spectral stream branch based on 1D CNN includes a pooling layer, a convolutional layer, and a residual block, and incorporates a channel attention mechanism; the convolutional kernel size is set to 1×1×m, and a batch normalization layer and a ReLU activation function are applied after each convolution operation.

[0018] Preferably, the spatial stream branch in S3 extracts and stacks spatial features along the spectral dimension, and obtains output features with enhanced spatial features through a spatial weighting strategy, which specifically includes the following steps:

[0019] Perform double pooling and convolution on the input features based on 2D CNN. After the last convolution, use the Sigmoid activation function to obtain the image spatial weight information , multiply the input features by the spatial weight information to obtain the output features with enhanced spatial features , and the expression is as follows:

[0020]

[0021] Preferably, the CVT hybrid architecture described in S4 includes: a multi-head self-attention module MHSA, a multi-layer perceptron MLP, and two layer normalizations LN, which are used to model the depth relationship of semantic token features ; two residual connections are introduced before the MHSA module and the MLP layer to prevent information loss; the MHSA module uses multiple groups of weight matrices to map Q, K, and V through the same operation, and a module with the same architecture as the mapping layer is used to perform output projection on the features, and the final self-attention results are stacked together and output, and the expression is defined as follows:

[0022]

[0023] In the formula, h represents the number of heads, and W is a linear transformation matrix to ensure that the output feature channels are consistent.

[0024] Preferably, the fuzzy supervision function is:

[0025] The normalized cross-entropy loss (Normalization Cross Entropy, NCE) and the reverse cross-entropy loss (Reverse Cross Entropy, RCE) are combined to form a fuzzy supervision function. It is worth mentioning that the RCE loss exchanges the predicted value and the true value positions in the ordinary cross-entropy loss (Cross Entropy, CE). The expressions of NCE, RCE, and the fuzzy supervision function are as follows:

[0026]

[0027]

[0028]

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] (1) The robust HSI classification composite depth network construction method of the present invention incorporates skip connections in the 1D CNN and 2D CNN dual-branch structure, which can alleviate the problems of gradient disappearance and explosion in the training of deep networks, making it lighter and more efficient, and contributing to improving the enhancement and extraction capabilities of spatial and spectral features and the network robustness.

[0031] (2) The robust HSI classification composite deep network construction method of the present invention adopts a CVT cascading strategy for the design of the feature extraction backbone network, combining the advantages of CNN in local feature induction and the excellent performance of Transformer in long-range correlation representation, which has obvious advantages compared with the mainstream 3D CNN and can capture spatial and spectral features more effectively.

[0032] (3) The robust HSI classification composite deep network construction method of the present invention designs a fuzzy supervision function based on the cross-entropy loss function theory, which can effectively guide the network to learn and utilize prior knowledge, and can suppress noise samples while ensuring the extraction of deep semantic features. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic diagram of the method flow of the present invention;

[0034] Figure 2 is a flow chart of spectral feature extraction of the present invention;

[0035] Figure 3 is a flow chart of spatial feature extraction of the present invention;

[0036] Figure 4 is a schematic diagram of CVT feature depth representation of the present invention;

[0037] Figure 5 is a schematic diagram of the multi-head self-attention principle of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.

[0039] Referring to the attached Figure 1 , a robust HSI classification composite deep network construction method includes the following steps:

[0040] S1. Based on the inaccurate label introduction mechanism, define the proportion of noise labels to be introduced in each dataset, and modify the original labels according to the selected noise type and proportion to create a noise label training set.

[0041] Divide the training dataset and the test dataset for 4 benchmark HSI datasets. Each benchmark dataset has two training sets, that is, different numbers of noise labels are added based on the number of clean labels. Specifically, for each class, randomly select n samples with true labels to form an initial training set; then, uniformly select Samples from other classes are taken and their labels are marked as the current class; then, all samples in each class are mixed to ensure that the noisy samples can be fully integrated into the initial training set.

[0042] Then, a classification training set of class C with noisy labels can be defined as:

[0043] (1)

[0044] where r represents the training sample, represents the label of the training sample (including the noisy label). In addition, is introduced to represent the possible correct class label. Based on the independent and identically distributed assumption, we have:

[0045] (2)

[0046] Since all samples are randomly shuffled, each label has a probability of being incorrect. Therefore, the probability that the sample j from the i-th class has a potentially clean label can be expressed as:

[0047] (3)

[0048] In the formula, .

[0049] S2. In the two-stream network, a spectral weighting strategy is introduced in the spectral stream branch, allowing the network to dynamically assign weights to different bands for local spectral pattern extraction and discriminative spectral feature enhancement.

[0050] To not affect the spatial features of the original data and effectively extract spectral features, a spectral stream branch based on 1D CNN is designed. This channel contains a pooling layer, a convolutional layer, a residual block, and incorporates a channel attention mechanism. The convolutional kernel size is set to 1×1×m. After each convolution operation, a batch normalization layer (Batch normalization layer, BN) and a ReLU activation function are applied. In the first convolutional layer, 64 convolutional kernels of size 1×1×m and a sampling stride of (1, 1, m) are used to perform a convolution operation on the input w×h×m 3D cube to obtain w×h×m 3D cube features, eliminating redundant spectral features and better focusing on the crucial features in classification; then, a residual block composed of two consecutive convolutional layers is used to perform deep feature extraction and enhancement on the first-layer convolutional features, emphasizing the important regions of the image and enhancing spectral robustness when dealing with noisy labels. This residual block has 64 convolutional kernels of size 1×1×m; after the residual block, a channel attention mechanism is adopted to further perform discriminative extraction of spectral features, thereby maintaining the separability of spectral information between classes.

[0051] As shown in the appendix Figure 2 Assume that the input feature is , and global average pooling is performed to obtain a one-dimensional vector feature of size 1×1×m The expression is as follows:

[0052] (4)

[0053] Then, the above one-dimensional vector feature is successively operated on by a convolution kernel of size k and an activation function to obtain the feature weight information in the spectral dimension , and the calculation method is as follows:

[0054] (5)

[0055] Among them, and respectively represent the Sigmoid activation function and the one-dimensional convolution function, and k represents the size of the convolution kernel. In addition, in this embodiment, a non-linear mapping method is used to adaptively determine the size of k according to m, and the expression is as follows:

[0056] (6)

[0057] Among them, b and are empirically preset parameters.

[0058] Finally, the input feature is multiplied pointwise with the spectral weight information to obtain the output feature with enhanced spectral features , and the expression is as follows:

[0059] (7)

[0060] S3. The spatial stream branch receives the data processed spectroscopically, extracts and stacks its spatial features along the spectral dimension, and obtains the output feature with enhanced spatial features through a spatial weighting strategy.

[0061] First, average pooling and max pooling are respectively performed on the input feature along the spectral dimension, and feature stacking is performed; next, the first layer of convolution uses 128 convolution kernels of size 3×3 for convolution operations; then, the second layer of convolution uses two convolution kernels of the same size for consecutive convolutions, and then a spatial weighting strategy is used to operate on the output feature of this layer to retain and emphasize important spatial information.

[0062] As shown in the appendix Figure 3 Assume that the input feature is also . In this embodiment, average pooling and max pooling are respectively performed along the spectral dimension, and feature stacking is performed. The expression is as follows:

[0063] (8)

[0064] (9)

[0065] (10)

[0066] Among them, 、 and represent the average pooling operation, the max pooling operation, and the spectral dimension feature stacking operation, respectively.

[0067] Next, this embodiment uses a convolutional kernel of size 7×7 to perform convolution on the Fm3 feature, and at the same time uses the Sigmoid activation function to obtain the image spatial weight information , and the expression is as follows:

[0068] (11)

[0069] Finally, the input feature is multiplied pointwise with the spatial weight information to obtain the output feature with enhanced spatial features , and the expression is as follows:

[0070] (12)

[0071] S4. Input the feature extraction results of the two-stream network into the CVT (CNN-vision transformer) hybrid architecture for feature mapping and convolutional encoding, and model the depth relationship of the semantic label features. Use a module with the same architecture as the mapping layer to perform output projection on the features to obtain the category information of the image patches.

[0072] As shown in the appendix Figure 4 , the CVT hybrid architecture is used to perform deep fusion of the above-mentioned spectral and spatial features, and at the same time perform feature stitching along the spectral dimension, and perform local and global depth feature characterization on the stitched features. First, the spectral feature and the spatial feature are stitched and stacked along the spectral dimension; then, a feature mapping layer is constructed using one-dimensional convolution, the ReLU activation function, and layer normalization; secondly, the Transformer is used to fuse and represent the spectral and spatial features of the features after convolutional encoding; next, a module with the same architecture as the mapping layer is used to perform output projection on the features; finally, a fully connected layer is used to obtain the category information of the image patches.

[0073] Assume is the input feature of the CVT module. This embodiment first uses 32 convolutional kernels of size 1×1 to perform feature mapping on the input feature to obtain the encoded feature , and the expression is as follows:

[0074] (13)

[0075] To capture the spatial long-distance feature dependencies in the image patches, in this paper, position information is added to the encoded features as follows: The expression is as follows:

[0076] (14)

[0077] To learn deep semantic features, the processing flow in the multi-head self-attention module is shown in the appendix Figure 5 as follows. Three learnable weight matrices , and are defined, and the above three learnable weight matrices are used to linearly map to query information Q, key information K, and value information V respectively. Then, the attention scores are obtained based on the product of Q and K, and the weights of the scores are obtained through the Softmax function. The multi-head self-attention module (Multi-head self attention, MHSA) uses multiple groups of weight matrices to map Q, K, and V through the same operation, and a module with the same architecture as the mapping layer is used to perform output projection on the features. Therefore, the final self-attention results are stacked and output, and the expression is defined as follows:

[0078] (15)

[0079] In the formula, h represents the number of heads, and W is the linear transformation matrix to ensure that the output feature channels are consistent.

[0080] Finally, the features after global representation are sequentially input into a multi-layer perceptron (Multi-Layer Perceptron, MLP) and layer normalization (Layer Normalization, LN). It is worth mentioning that the MLP layer consists of two fully connected layers with an activation function in the middle. The fully connected layer is used to obtain the final class information, and layer normalization can prevent gradient explosion or disappearance and accelerate network convergence.

[0081] S5. Introduce a fuzzy supervision function to convert hard labels into soft labels, reduce the network's overconfidence in specific labels, and dynamically adjust the loss value according to the prediction probability in each training batch to suppress the expression of mislabeled sample information.

[0082] Combine the Normalization Cross Entropy (NCE) and the Reverse Cross Entropy (RCE) to form a Fuzzy supervision function (FSL). It is worth mentioning that the RCE loss exchanges the predicted value and the true value in the ordinary Cross Entropy (CE). The expressions of NCE, RCE, and FSL are as follows:

[0083] (16)

[0084] (17)

[0085] (18)

[0086] Specifically, assume that is the empirical risk when training with noisy labels, and is the empirical risk when training with true labels (i.e., and ). According to the law of total expectation, . By combining the chain rule and the independent assumption in (2), (16) can be transformed into:

[0087] (19)

[0088] According to the probability definition in (3), (19) can be rewritten as:

[0089] (20)

[0090] Since the NCE loss is a normalized version of the CE loss, the expectation . Therefore, can be simplified to:

[0091] (21)

[0092] Here, assume that is the global minimum of the empirical loss, then:

[0093] (22)

[0094] In addition, when the noise rate is , is also local minimum, and the NCE loss is robust to noisy labels. Replace with , a necessary condition for NCE to be robust to noisy labels is .

[0095] The noise processing process of the RCE loss is the same as the above-mentioned NCE. The combined use of the two loss functions can not only reduce the influence of noisy labels, but also encourage the learning and training process of the network. The NCE loss, as an active learning strategy, aims to maximize the probability of correct labels, while the RCE loss, as a passive learning strategy, aims to suppress the probability of at least one wrong label. This learning process can further improve the classification accuracy of noisy labels.

[0096] The foregoing description of specific exemplary embodiments of the invention has been presented for purposes of illustration and example. These descriptions are not intended to limit the invention to the precise forms disclosed, and it is apparent that, in light of the above teachings, many modifications and variations are possible. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the invention and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the invention, as well as various different selections and modifications. The scope of the invention is intended to be defined by the claims and their equivalents.

Claims

1. A robust HSI classification composite deep network construction method, characterized in that: The following steps are involved: S1. Based on the inaccurate label introduction mechanism, define the proportion of noise labels to be introduced in each data set, modify the original labels according to the selected noise type and proportion, and create a noise label training set; In S2, a spectral weighting strategy is introduced in the spectral stream branch in the two-stream network, allowing the network to dynamically assign weights to different bands to extract local spectral patterns and enhance discriminative spectral features; S3, the spatial stream branch receives the spectrally processed data, extracts and stacks the spatial features along the spectral dimension, and obtains the output features with spatial feature enhancement through the spatial weighting strategy; S4, input the feature extraction results of the two-stream network into the CVT hybrid architecture, perform feature mapping and convolutional coding, model the deep relationship between local-global semantic tag features, use the module with the same architecture as the mapping layer to output the features and obtain the category information of the image block; S5. Introduce a fuzzy supervision function to convert hard labels into soft labels, reduce the network's overconfidence in specific labels, and dynamically adjust the loss value according to the prediction probability in each training batch to suppress the expression of mislabeled sample information.

2. The robust HSI classification composite deep network construction method according to claim 1, characterized in that: The specific mechanism of inaccurate label introduction described in S1 is: First, for each class, n samples with true labels are randomly selected to form the initial training set; then, samples from other classes and label them as the current class; then, all samples in each class are mixed to ensure that the noise samples can be fully integrated into the initial training set; then, a C-class classification training set with noise labels can be defined as:

3. The robust HSI classification composite deep network construction method according to claim 1, characterized in that: In the two-stream network described in S2, the spectral stream branch based on 1D CNN contains a pooling layer, a convolution layer and a residual block, and integrates a channel attention mechanism; the convolution kernel size is set to 1×1×m, and a batch normalization layer BN and ReLU activation function are applied after each convolution operation.

4. The robust HSI classification composite deep network construction method according to claim 1, characterized in that: In S3, the spatial stream branch is used to extract and stack spatial features along the spectral dimension, and the output features of spatial feature enhancement are obtained through the spatial weighting strategy. Specifically, the following steps are included: Based on 2D CNN, double pooling and convolution are performed on the input features. After the last layer of convolution, the Sigmoid activation function is used to obtain the image spatial weight information. , multiply the input features by the spatial weight information to obtain the output features enhanced by the spatial features , the expression is as follows:

5. The robust HSI classification composite deep network construction method according to claim 1, characterized in that: The CVT hybrid architecture described in S4 includes a multi-head self-attention module MHSA, a multi-layer perceptron MLP and two layer normalization LNs for semantic labeling features. The deep relationship between Q, K, and V is modeled; two residual connections are introduced between the MHSA module and the MLP layer. The MHSA module uses multiple sets of weight matrices to map Q, K, and V through the same operation, and uses a module with the same architecture as the mapping layer to output the features. The final self-attention results are stacked together and output. The expression is defined as follows: Where h represents the number of heads and W is the linear transformation matrix.

6. The robust HSI classification composite deep network construction method according to claim 1, characterized in that: The fuzzy supervision function is formed by the combination of the normalized cross entropy loss NCE and the inverse cross entropy loss RCE; the RCE loss exchanges the predicted value for the ordinary cross entropy loss CE. and the true value The expressions of NCE, RCE and fuzzy supervision function FSF are as follows: 。

Citation Information

Patent Citations

  • Hyperspectral image classification method fusing CNN and ViT spatial spectrum features

    CN116824220A

  • Label noise-containing hyperspectral image crop classification based on 3DCNN combined with MLP

    CN119180984A

Cited By

  • Hyperspectral image classification method based on spectral feature reconstruction

    CN120356015A