Hyperspectral image classification method and system based on hierarchical bidirectional contrast learning
Through the method based on hierarchical two-way contrast learning, hyperspectral remote sensing images are preprocessed and data augmented, and the features of image data are extracted and calculated. The encoder network model is trained to realize self-supervised classification, which solves the problems of information redundancy and labeling consumption in hyperspectral image classification and improves classification accuracy.
Patent Information
- Application Number
- CN202411967756.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Information redundancy is prone to occur during hyperspectral image classification, and accurate annotation of hyperspectral images requires a lot of manpower and material resources.
A hyperspectral image classification method based on hierarchical bidirectional contrast learning is proposed. By pre-processing and data augmenting of hyperspectral remote sensing image data, online characterization and target characterization are extracted, positive and negative samples are calculated, and the encoder network model is trained to achieve self-supervised classification.
It effectively solves the problem of information redundancy in hyperspectral image classification, improves the accuracy of classification, and reduces the manpower and material consumption for precise labeling of hyperspectral images.
Smart Images

Figure CN120014317A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image classification, and in particular to a hyperspectral image classification method and system based on hierarchical bidirectional contrast learning. Background Art
[0002] Hyperspectral Image (HSI) classification refers to assigning images to certain ground object categories based on their spatial and spectral characteristics. It is one of the most active research directions in the field of hyperspectral remote sensing.
[0003] At present, in the early stage of the development of hyperspectral image classification, most methods focus on exploring the role of spectral information in classification. Many traditional machine learning methods have achieved good results in hyperspectral image classification, such as Support Vector Machine (SVM), Multinomial Logistic Regression (MLR), Dynamic or Random Subspace (DRS), etc. In addition, due to the existence of the "dimensionality curse" phenomenon, some methods are committed to designing effective feature extraction or dimensionality reduction techniques, such as Principal Component Analysis (PCA), Independent Component Analysis (ICA) and Linear Discriminant Analysis (LDA). Although these methods have achieved good results in HSI classification, they have a common defect: the method relies on prior knowledge and parameter settings, which seriously affects the robustness and generalization of traditional methods, resulting in poor performance under complex conditions.
[0004] In recent years, deep learning has made significant breakthroughs in computer vision tasks such as image classification. Under the premise of sufficient training samples, deep learning continuously fits and approximates the sample distribution of the target data set through the superposition of multiple linear layers, and then uses reverse gradient propagation to update the linear layer parameters layer by layer from shallow to deep during the training process, which can well extract data features, making the application scope of deep learning wider.
[0005] However, the distribution of these types of objects is obtained by annotating the pixels of hyperspectral images through image processing technology. Although hyperspectral images contain rich spatial spectral information of objects that can be used for classification, their spectral dimensions are geometrically increased compared to ordinary remote sensing images, so they are prone to problems such as information redundancy and dimensionality disaster; and with the development of technology, it is easier to obtain hyperspectral images, but achieving accurate annotation of hyperspectral images is a complex and laborious task. Summary of the invention
[0006] In order to overcome the problems that information redundancy is prone to occur in the classification process of existing hyperspectral image classification technology, and the accurate labeling of hyperspectral images requires a large amount of manpower and material resources, the purpose of the present invention is to propose a hyperspectral image classification method and system based on hierarchical bidirectional contrastive learning, which can effectively solve the information redundancy that occurs in the hyperspectral image classification process, improve the accuracy of hyperspectral image classification, and thus effectively reduce the consumption of large manpower and material resources for accurate labeling of hyperspectral images.
[0007] To achieve the purpose of the present invention, the present invention adopts the following technical solutions:
[0008] A hyperspectral image classification method based on hierarchical bidirectional contrast learning, the method comprising the following steps:
[0009] Acquiring hyperspectral remote sensing image data, and preprocessing the hyperspectral remote sensing image data;
[0010] Performing data enhancement on the preprocessed hyperspectral remote sensing image to construct a number of data enhancement samples, and extracting online representation and target representation of the image data based on the constructed data enhancement samples;
[0011] Calculate the positive and negative samples of the online representation based on the target representation, and train the encoder network model based on the online representation and the positive and negative samples;
[0012] The encoder network model is used to classify the hyperspectral remote sensing image data to be classified, and the classification result is output.
[0013] In the above technical scheme, preprocessing the acquired hyperspectral remote sensing image data can improve the practicality and reliability of the image data; by performing data enhancement on the preprocessed image data, the features of the image data can be effectively highlighted to effectively extract the online representation and target representation of the image data; positive and negative samples of the image data can be further extracted based on the extracted online representation and target representation, so as to more finely reflect the correlation and difference between different types of data in the image data, so as to improve the aggregation between the same type of objects in the image data and the separability between different types of objects, and effectively solve the information redundancy that occurs in the hyperspectral image classification process; an encoder network model is trained based on the extracted positive and negative samples to obtain an unsupervised image classification model that can effectively grasp the changes in the hyperspectral remote sensing image data, so as to use the model to perform self-supervised classification of the input hyperspectral remote sensing image data without label dependence, effectively improve the model's discriminability for hyperspectral difficult-to-distinguish samples and the accuracy of hyperspectral image classification, and thus effectively reduce the large amount of manpower and material resources consumed by the precise annotation of hyperspectral images.
[0014] Furthermore, the process of preprocessing the hyperspectral remote sensing image data includes:
[0015] The acquired hyperspectral remote sensing image data is subjected to denoising and normalization processing.
[0016] In the above technical solution, the acquired hyperspectral remote sensing image data is subjected to denoising and normalization processing, which can improve the practicality and reliability of the image data and reduce data redundancy.
[0017] Further, the extracted data augmentation samples are respectively input into a preset online encoder and a target encoder;
[0018] Extracting online representation of image data through online encoder;
[0019] The target representation of the image data is extracted through the target encoder.
[0020] Furthermore, the process of calculating the positive samples and negative samples of the online representation based on the target representation includes:
[0021] Design an encoder network based on hierarchical dictionary query, expressed as:
[0022]
[0023]
[0024] The encoder network based on hierarchical dictionary query maintains two dictionary structures at different depths of the network, updates the dictionary prototype online by receiving gradient updates, retrieves positive and negative samples of different depths of online representation by calculating the similarity between the dictionary prototype and the target representation, and uses the retrieved dictionary prototype with the maximum similarity value with the target representation as the positive sample, and uses the retrieved dictionary prototype with the minimum similarity value with the target representation as the negative sample;
[0025] Among them, D 1 represents the shallow dictionary of the encoder network, D 2 represents the deep dictionary of the encoder network, D 1 * Represents the parameters of the shallow dictionary, D 2 * Parameters representing the deep dictionary, Indicates the parameters of the online encoder, represents the positive sample prototype, Represents the remaining prototypes, L P (x i ) represents the prototype contrast loss function, z i Represents the online representation extracted by the online encoder.
[0026] Furthermore, the online representation and target representation of the shallow and deep layers are calculated at different depths of the encoder network, and the expressions are:
[0027]
[0028] in, and Respectively represent the online encoder φ q (·) and the target encoder φ k (·) is the s-th convolutional layer, Avg represents dynamic average pooling, g(·) represents Batch Normalization and ReLU activation function mapping, represents the online representation of the output of the sth convolutional layer, represents the target representation output by the sth convolutional layer, x q and x k Represent the enhanced images of the input online encoder and the target encoder respectively;
[0029] The expression for retrieving positive samples of online representations based on the similarity between the dictionary prototype and the target representation is:
[0030]
[0031] The expression for retrieving negative samples of online representations based on the similarity between the dictionary prototype and the target representation is:
[0032]
[0033] in, Represents the input sample x i The target representation, d j Represents a dictionary prototype.
[0034] In the above technical scheme, an encoder network based on hierarchical dictionary query is designed to store the learned difficult-to-distinguish positive and negative samples. At the same time, it is proposed to capture data representations of different scales through dictionary structures in the shallow and deep layers of the deep network for comparative supervision, so as to enhance the category perception ability of the network and better extract the positive and negative sample data features of hyperspectral remote sensing images, thereby improving the discriminability of hyperspectral difficult-to-distinguish samples and the accuracy of hyperspectral image classification.
[0035] Furthermore, the process of training the encoder network model based on the online representation and its positive and negative samples includes:
[0036] After removing labels from all samples of the preprocessed hyperspectral remote sensing image data, the samples are input into the encoder network model, and several rounds of iterative optimization training are set to optimize the parameters of the encoder network model;
[0037] Based on the hierarchical dictionary structure, positive and negative samples at different depths of the encoder network are extracted;
[0038] According to the extracted positive and negative samples of different depths, the prototype contrast loss is constructed for positive contrast learning, and the reverse contrast loss is constructed for reverse contrast learning according to the negative samples. The two losses are combined to form a bidirectional contrast loss.
[0039] The encoder network is used to learn the potential representation in the hyperspectral remote sensing image data samples, and the bidirectional contrast loss function is used to optimize the parameters of the encoder network during the iterative training process. When the preset training round ends or the loss function converges, the trained encoder network model parameters are obtained.
[0040] Initialize the new encoder and fully connected layer, load the pre-trained encoder network model parameters to the new encoder and freeze them, and divide the hyperspectral remote sensing image data into a training set and a test set;
[0041] The training set is used to train the fully connected layer, and the encoder loaded with the network model parameters and the trained fully connected layer are used to classify the hyperspectral images of the test set and output the classification results.
[0042] In the above technical scheme, all samples of the preprocessed hyperspectral remote sensing image data are delabeled and input into the encoder network model, and positive and negative samples of different depths are extracted through a hierarchical dictionary structure. Then, a prototype contrast loss is constructed according to the positive and negative samples of different depths for forward contrast learning, and a reverse contrast loss is used for reverse contrast learning. This can more finely reflect the correlation and difference between different types of data in the image data, so as to improve the aggregation between similar objects and the separability between different categories of objects in the image data. The encoder network is used to learn the potential representation of the hyperspectral remote sensing image data samples, and a bidirectional contrast loss function is used to optimize the parameters of the encoder network in the iterative training process. The trained encoder network model can effectively grasp the changes in hyperspectral remote sensing image data, thereby improving the aggregation between objects of the same category and the separability between objects of different categories under self-supervision conditions; initializing an encoder and a fully connected layer can load the parameters of the trained encoder network model, and realize self-supervised classification of hyperspectral remote sensing images in a division of labor with the encoder network model, effectively improving the model's discriminability for hyperspectral difficult-to-distinguish samples and the accuracy of hyperspectral image classification; and when optimizing the parameters of the existing model, there is no need to retrain the encoder, only the fully connected layer that processes downstream tasks needs to be trained, thereby effectively reducing the large amount of manpower and material resources consumed by the precise annotation of hyperspectral images.
[0043] Furthermore, the process of constructing reverse contrast loss and performing reverse contrast learning according to the extracted negative samples of different depths includes:
[0044] Based on the negative samples retrieved from the dictionary, a reverse contrast loss function is constructed, which is expressed as:
[0045]
[0046] in, is the negative sample prototype, and τ represents the temperature coefficient.
[0047] Furthermore, a hierarchical bidirectional contrast loss function is set to optimize the parameters of the encoder network model, and the expression is:
[0048]
[0049] Among them, L P (x i ) and L N (x i ) represent the prototype contrast loss function and the reverse contrast loss function, respectively, which are set in the shallow and deep layers of the encoder network at the same time, λ represents the weight term of the reverse contrast loss, and B represents the batch size.
[0050] Furthermore, the expression of the potential representation in the hyperspectral remote sensing image data sample learned by the encoder network is:
[0051]
[0052] The expression for hyperspectral image classification using the learned potential representation through the initialized fully connected layer is:
[0053] D=FC(L)
[0054] Among them, x represents the test set sample, φ(·) represents the encoder network, represents the model parameters of the encoder network, D represents the classification result, and FC represents the initialized fully connected layer.
[0055] In the above technical scheme, the set reverse contrast loss function and bidirectional contrast loss function can effectively optimize the parameters of the encoder network model in the iterative training process to improve the discriminability of the encoder network model for hyperspectral difficult-to-distinguish samples and the accuracy of hyperspectral image classification.
[0056] A hyperspectral image classification system based on hierarchical bidirectional contrast learning, the system comprising:
[0057] A data acquisition module, used to acquire hyperspectral remote sensing image data;
[0058] A data processing module, used for preprocessing the hyperspectral remote sensing image data;
[0059] A representation extraction module is used to perform data enhancement on the preprocessed hyperspectral remote sensing image to construct a number of data enhancement samples, and extract online representation and target representation of the image data according to the constructed data enhancement samples;
[0060] A model training module is used to calculate positive and negative samples of the online representation based on the target representation, and train the encoder network model according to the online representation and the positive and negative samples;
[0061] The image classification module is used to classify the hyperspectral remote sensing image data to be classified using the encoder network model and output the classification result.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] The present invention proposes a hyperspectral image classification method and system based on hierarchical bidirectional contrast learning. By preprocessing the acquired hyperspectral remote sensing image data, the practicability and reliability of the image data can be improved. By performing data enhancement on the preprocessed image data, the features of the image data can be effectively highlighted to effectively extract the online representation and target representation of the image data. According to the extracted online representation and target representation, positive and negative samples of the image data can be further extracted, so as to more finely reflect the correlation and difference between different types of data in the image data, so as to improve the aggregation between the same type of objects and the separability between different types of objects in the image data, and effectively solve the information redundancy in the hyperspectral image classification process. According to the extracted positive and negative samples, an encoder network model is trained to obtain an unsupervised image classification model that can effectively grasp the changes of the hyperspectral remote sensing image data, so as to use the model to perform self-supervised classification of the input hyperspectral remote sensing image data without label dependence, effectively improve the discriminability of the model for hyperspectral difficult-to-distinguish samples and the accuracy of hyperspectral image classification, and thus effectively reduce the large manpower and material resources consumed by the precise annotation of hyperspectral images. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 A flowchart of a method for hyperspectral image classification based on hierarchical bidirectional contrast learning provided in an embodiment of the present application;
[0065] Figure 2 A schematic diagram of the hierarchical bidirectional contrastive self-supervised network model structure provided in an embodiment of the present application;
[0066] Figure 3 A schematic diagram of a dictionary structure provided in an embodiment of the present application;
[0067] Figure 4 A diagram showing the structure of an encoder in a hierarchical bidirectional contrastive self-supervisory network provided in an embodiment of the present application;
[0068] Figure 5 A schematic diagram of a process for image classification of a given test sample provided in an embodiment of the present application;
[0069] Figure 6 A hyperspectral remote sensing image classification result diagram provided in an embodiment of the present application;
[0070] Figure 7 A schematic diagram of the structure of a hyperspectral image classification system based on hierarchical bidirectional contrast learning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Preferred embodiments of the present invention are provided in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive.
[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0073] Embodiment 1:
[0074] This embodiment provides a hyperspectral image classification method based on hierarchical bidirectional contrast learning. Figure 1 and Figure 2 , the method comprises the following steps:
[0075] Step S1: acquiring hyperspectral remote sensing image data, and preprocessing the hyperspectral remote sensing image data;
[0076] Step S2: performing data enhancement on the preprocessed hyperspectral remote sensing image to construct a number of data enhancement samples, and extracting online representation and target representation of the image data according to the constructed number of data enhancement samples;
[0077] Step S3: Calculate positive samples and negative samples of the online representation based on the target representation, and train the encoder network model according to the online representation and the positive and negative samples;
[0078] Step S4: classify the hyperspectral remote sensing image data to be classified using the encoder network model, and output the classification result.
[0079] In step S1, the process of preprocessing the hyperspectral remote sensing image data includes:
[0080] The acquired hyperspectral remote sensing image data is subjected to denoising and normalization processing.
[0081] In the above technical solution, the acquired hyperspectral remote sensing image data is subjected to denoising and normalization processing, which can improve the practicality and reliability of the image data and reduce data redundancy.
[0082] As a preferred embodiment, in step S2, the extracted data enhancement samples are respectively input into a preset online encoder and a target encoder;
[0083] Extracting online representation of image data through online encoder;
[0084] The target representation of the image data is extracted through the target encoder.
[0085] In this embodiment, preprocessing the acquired hyperspectral remote sensing image data can improve the practicality and reliability of the image data; by performing data enhancement on the preprocessed image, the features of the image data can be effectively highlighted to effectively extract the online representation and target representation of the image data; positive and negative samples of the image data can be further extracted based on the extracted online representation and target representation, thereby more finely reflecting the correlation and difference between different types of data in the image data, so as to improve the aggregation between similar objects and the separability between different categories of objects in the image data, and effectively solve the information redundancy that occurs in the hyperspectral image classification process; an encoder network model is trained based on the extracted positive and negative samples to obtain a model that can effectively grasp the changes in hyperspectral remote sensing image data and realize self-supervised image classification, so as to use the model to perform self-supervised classification on other hyperspectral remote sensing image data that are subsequently input, effectively improving the generalization of the model, the discriminability of hyperspectral difficult-to-distinguish samples, and the accuracy of hyperspectral image classification, thereby effectively reducing the large amount of manpower and material resources consumed by the precise annotation of hyperspectral images.
[0086] Embodiment 2:
[0087] This embodiment further describes steps S3 to S4 in the first embodiment, as follows:
[0088] As a preferred embodiment, in step S3, the process of calculating positive samples and negative samples of the online representation based on the target representation includes:
[0089] See also Figure 3 , design an encoder network based on hierarchical dictionary query, the expression is:
[0090]
[0091] The encoder network based on hierarchical dictionary query maintains two dictionary structures at different depths of the network, updates the dictionary prototype online by receiving gradient updates, retrieves positive and negative samples of the online representation by calculating the similarity between the dictionary prototype and the target representation, and uses the retrieved dictionary prototype with the maximum similarity value to the target representation as the positive sample, and uses the retrieved dictionary prototype with the minimum similarity value to the target representation as the negative sample;
[0092] Among them, D 1 represents the shallow dictionary of the encoder network, D 2 represents the deep dictionary of the encoder network, D 1 * Represents the parameters of the shallow dictionary, D2 * Parameters representing the deep dictionary, Indicates the parameters of the online encoder, represents the positive sample prototype, Represents the remaining prototypes, L P (x i ) represents the prototype contrast loss function, z i Represents the online representation extracted by the online encoder.
[0093] Specifically, the encoder network based on hierarchical dictionary query specifically includes:
[0094] A shallow dictionary D maintained in a way that receives gradient updates 1 and the deep dictionary D 2 ;
[0095] Input hyperspectral image sample x and its enhanced image x q and x k ;
[0096] Online encoder φ trained online q and momentum updated target encoder φ k ;
[0097] The enhanced image is output through two encoders to obtain the online representation and target representation of the image. The target representation is calculated with all the prototypes of the dictionary to obtain the similarity value. The dictionary prototype with the maximum similarity value is retrieved as the positive sample, and the dictionary prototype with the minimum similarity value is the negative sample.
[0098] Furthermore, the shallow dictionary D 1 and the deep dictionary D 2 The construction and acquisition of appropriate positive samples Specifically include:
[0099] The dictionary positive prototype is updated in the form of minimizing the contrast loss function, and the dictionary negative prototype is updated in the form of maximizing the contrast loss function. The dictionary is updated along with the encoder network as part of the network parameters. Since the representation is normalized on the unit hypersphere, the inner product operation between the target representation and the dictionary prototype represents the sample cosine similarity. The dictionary prototype with the highest similarity to the online representation is retrieved as the positive sample of the online representation.
[0100] As a preferred embodiment, the online representation and target representation of the shallow layer and the deep layer are calculated at different depths of the encoder network, respectively, and the expressions are:
[0101]
[0102] in, and Respectively represent the online encoder φq (·) and the target encoder φ k (·) is the s-th convolutional layer, Avg represents dynamic average pooling, g(·) represents Batch Normalization and ReLU activation function mapping, represents the online representation of the output of the sth convolutional layer, represents the target representation output by the sth convolutional layer (in this embodiment, s is 1 or 5, s is 1 to extract shallow representation, and s is 5 to extract deep representation), x q and x k Represent the enhanced images of the input online encoder and the target encoder respectively;
[0103] The expression for retrieving positive samples of online representations based on the similarity between the dictionary prototype and the target representation is:
[0104]
[0105] The expression for retrieving negative samples of online representations based on the similarity between the dictionary prototype and the target representation is:
[0106]
[0107] in, Represents the input sample x i The target representation, d j Represents a dictionary prototype.
[0108] It is understandable that, considering the difficulty in sampling positive and negative samples, an encoder network based on dictionary query is designed to store the learned difficult-to-distinguish positive and negative samples, so as to query the similarity between multi-batch samples and online representations, realize efficient positive and negative sample retrieval, decouple the model's dependence on large batches, and improve the generalization of the network; at the same time, it is proposed to capture data representations of different scales through dictionary structures in the shallow and deep layers of the encoder network for comparative supervision, which solves the problem of lack of attention to shallow semantic features in some contrastive learning, realizes more sufficient sample separability, and improves the network's ability to maintain the manifold structure and category perception of hyperspectral data; the bidirectional contrast loss combined with reverse contrastive learning alleviates the network's extreme uniform distribution through bidirectional information flow, maintains the potential manifold structure of the data, and further improves the network's semantic representation ability and generalization, thereby improving the model's ability to discriminate hyperspectral difficult samples and the accuracy of hyperspectral image classification.
[0109] As a preferred embodiment, in step S3, see Figure 2 and Figure 4 , the process of training the encoder network model based on the online representation and its positive and negative samples includes:
[0110] Step S31: removing labels from all samples of the preprocessed hyperspectral remote sensing image data and inputting them into the encoder network model, and setting several rounds of iterative optimization training to optimize the parameters of the encoder network model;
[0111] Step S32: extracting positive and negative samples at different network depths of the encoder based on the hierarchical dictionary structure;
[0112] Step S33: constructing a prototype contrast loss based on the extracted positive samples of different depths to perform positive contrast learning, constructing a reverse contrast loss based on the negative samples to perform reverse contrast learning, and combining the two losses to form a bidirectional contrast loss;
[0113] Step S34: learning the potential representation in the hyperspectral remote sensing image data sample through the encoder network, and optimizing the parameters of the encoder network in the iterative training process by using a bidirectional contrast loss function, and obtaining the trained encoder network model parameters when the preset training round ends or the loss function converges;
[0114] Step S35: Initialize a new encoder and a fully connected layer, load the pre-trained encoder network model parameters to the new encoder and freeze them, and divide the hyperspectral remote sensing image data into a training set and a test set;
[0115] Step S36: Use the training set to train the fully connected layer, and use the encoder loaded with the network model parameters and the trained fully connected layer to classify the hyperspectral images of the test set, and output the classification results.
[0116] Specifically, in step S33, the process of constructing reverse contrast loss and performing reverse contrast learning according to the extracted negative samples of different depths includes:
[0117] Based on the negative samples retrieved from the dictionary, a reverse contrast loss function is constructed, which is expressed as:
[0118]
[0119] in, is the negative sample prototype, and τ represents the temperature coefficient.
[0120] Specifically, in step S34, a bidirectional contrast loss function is set to optimize the parameters of the encoder network model, and the expression is:
[0121]
[0122] Among them, L P (x i ) and L N (x i) represent the prototype contrast loss function and the reverse contrast loss function, respectively, which are set in the shallow and deep layers of the encoder network at the same time, λ represents the weight term of the reverse contrast loss, and B represents the batch size.
[0123] The expression for learning the latent representation in the hyperspectral remote sensing image data sample by the encoder network is:
[0124]
[0125] In step S35 and step S36, see Figure 5 and Figure 6 , the expression of hyperspectral image classification using the learned potential representation through the initialized fully connected layer is:
[0126] D=FC(L)
[0127] Among them, x represents the test set sample, φ(·) represents the encoder network, represents the model parameters of the encoder network, D represents the classification result, and FC represents the initialized fully connected layer.
[0128] For example, construct Figure 2 The self-supervised network model shown in Figure 2, where the encoder structure is as follows: Figure 4 As shown, the dictionary structure is as follows Figure 3 The encoder consists of the initial Conv 3×3 convolution layer, the batch normalization layer BN, the activation layer ReLU, the MaxPool layer, and the repeated N 0 times (repeated twice in this embodiment, the difficulty of the task can be increased by N 0 The whole framework uses a twin network structure, including two encoder branches φ with the same structure. q and φ k , where φ q Update φ with slow momentum update k In the encoder, the convolution layer is used to extract local features, the BN layer is used to normalize the features, and the ReLU layer is used for nonlinear activation. Image classification is not performed in the network pre-training stage. After the pre-training stage is completed, a new initialized encoder and FC layer are created for classification tasks.
[0129] Without loss of generality, let For a hyperspectral sample, is the learned potential representation, then L can be expressed as:
[0130]
[0131] In the above formula, Represents the parameters of the pre-trained model. Initialize a fully connected layer FC. For a given input sample x, input it into the pre-trained model φ(·) and obtain the corresponding latent representation L, and classify the latent representation through the fully connected layer FC to obtain the final classification result D.
[0132] The final classification result D is specifically expressed as:
[0133] D=FC(L)
[0134] In this embodiment, refer to Figure 6 , gives the hyperspectral image classification result diagram, where Figure 6 (a) is a false color image composed of three bands of the original hyperspectral image. Figure 6 (b) is the ground truth. Figure 6 (c) is the classification result of Indian Pines hyperspectral remote sensing image.
[0135] It can be understood that the set reverse contrast loss function and the bidirectional contrast loss function combined with the prototype contrast loss can effectively optimize the parameters of the encoder network model in the iterative training process to improve the discriminability of the encoder network model for hyperspectral difficult samples and the accuracy of hyperspectral image classification.
[0136] In an embodiment, all samples of the preprocessed hyperspectral remote sensing image data are delabeled and input into an encoder network model, and positive and negative samples of different depths are extracted through a hierarchical dictionary structure, and then a prototype contrast loss is constructed according to the positive and negative samples of different depths for forward contrast learning, and a reverse contrast loss is used for reverse contrast learning, which can more finely reflect the correlation and difference between different types of data in the image data, so as to improve the aggregation between similar objects and the separability between different categories of objects in the image data; the potential representation in the hyperspectral remote sensing image data samples is learned through the encoder network, and the parameters of the encoder network in the iterative training process are optimized using a bidirectional contrast loss function, which can make The trained encoder network model can effectively grasp the changes in hyperspectral remote sensing image data, thereby improving the aggregation between objects of the same category and the separability between objects of different categories under self-supervision conditions; initializing an encoder and a fully connected layer can load the parameters of the trained encoder network model, and realize self-supervised classification of hyperspectral remote sensing images in a division of labor with the encoder network model, effectively improving the model's discriminability for hyperspectral difficult-to-distinguish samples and the accuracy of hyperspectral image classification, and when optimizing the parameters of the existing model, there is no need to retrain the encoder, only the fully connected layer that processes downstream tasks needs to be trained, thereby effectively reducing the large amount of manpower and material resources consumed by the accurate annotation of hyperspectral images.
[0137] Embodiment three:
[0138] This embodiment provides a hyperspectral image classification system based on hierarchical bidirectional contrast learning. Figure 7 , the system comprising:
[0139] A data acquisition module, used to acquire hyperspectral remote sensing image data;
[0140] A data processing module, used for preprocessing the hyperspectral remote sensing image data;
[0141] A representation extraction module is used to perform data enhancement on the preprocessed hyperspectral remote sensing image data to construct a number of data enhancement samples, and extract online representation and target representation of the image data based on the constructed data enhancement samples;
[0142] A model training module is used to calculate positive and negative samples of the online representation based on the target representation, and train the encoder network model according to the online representation and the positive and negative samples;
[0143] The image classification module is used to classify the hyperspectral remote sensing image data to be classified using the encoder network model and output the classification result.
[0144] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A hyperspectral image classification method based on hierarchical bidirectional contrast learning, characterized in that: The method comprises the following steps: Acquiring hyperspectral remote sensing image data, and preprocessing the hyperspectral remote sensing image data; Performing data enhancement on the preprocessed hyperspectral remote sensing image to construct a number of data enhancement samples, and extracting online representation and target representation of the image data based on the constructed data enhancement samples; Calculate the positive and negative samples of the online representation based on the target representation, and train the encoder network model based on the online representation and the positive and negative samples; The encoder network model is used to classify the hyperspectral remote sensing image data to be classified, and the classification result is output.
2. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 1 is characterized in that: The process of preprocessing the hyperspectral remote sensing image data includes: The acquired hyperspectral remote sensing image data is subjected to denoising and normalization processing.
3. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 1 is characterized in that: Input the constructed data augmentation samples into the preset online encoder and target encoder respectively; Extracting online representation of image data through online encoder; The target representation of the image data is extracted through the target encoder.
4. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 1 is characterized in that: The process of calculating the positive and negative samples of the online representation based on the target representation includes: Design an encoder network based on hierarchical dictionary query, expressed as: The encoder network based on hierarchical dictionary query maintains two dictionary structures at different depths of the network, updates the dictionary prototype online by receiving gradient updates, retrieves positive and negative samples of the online representation by calculating the similarity between the dictionary prototype and the target representation, and uses the retrieved dictionary prototype with the maximum similarity value to the target representation as the positive sample, and uses the retrieved dictionary prototype with the minimum similarity value to the target representation as the negative sample; Among them, D1 represents the shallow dictionary of the encoder network, D2 represents the deep dictionary of the encoder network, and D1 * Represents the parameters of the shallow dictionary, D2 * Parameters representing the deep dictionary, Indicates the parameters of the online encoder, represents the positive sample prototype, Represents the remaining prototypes, L P (x i ) represents the prototype contrast loss function, z i Represents the online representation extracted by the online encoder.
5. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 4 is characterized in that: At different depths of the encoder network, shallow and deep online representations and target representations are extracted respectively, and the expressions are: in, and Respectively represent the online encoder φ q (·) and the target encoder φ k (·) is the s-th convolutional layer, Avg represents dynamic average pooling, g(·) represents Batch Normalization and ReLU activation function mapping, represents the online representation of the output of the sth convolutional layer, represents the target representation output by the sth convolutional layer, x q and x k Represent the enhanced images of the input online encoder and the target encoder respectively; The expression for retrieving positive samples of online representations based on the similarity between the dictionary prototype and the target representation is: The expression for retrieving negative samples of online representations based on the similarity between the dictionary prototype and the target representation is: in, Represents the input sample x i The target representation, d j Represents a dictionary prototype.
6. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 4 is characterized in that: The process of training the encoder network model based on the online representation and its positive and negative samples includes: After removing labels from all samples of the preprocessed hyperspectral remote sensing image data, the samples are input into the encoder network model, and several rounds of iterative optimization training are set to optimize the parameters of the encoder network model; Based on the hierarchical dictionary structure, positive and negative samples at different depths of the encoder network are extracted; According to the extracted positive samples of different depths, the prototype contrast loss is constructed for positive contrast learning, and the reverse contrast loss is constructed for reverse contrast learning according to the negative samples. The two losses are combined to form a bidirectional contrast loss. The encoder network is used to learn the potential representation in the hyperspectral remote sensing image data samples, and the bidirectional contrast loss function is used to optimize the parameters of the encoder network during the iterative training process. When the preset training round ends or the loss function converges, the trained encoder network model parameters are obtained. Initialize the new encoder and fully connected layer, load the pre-trained encoder network model parameters to the new encoder and freeze them, and divide the hyperspectral remote sensing image data into a training set and a test set; The training set is used to train the fully connected layer, and the encoder loaded with the network model parameters and the trained fully connected layer are used to classify the hyperspectral images of the test set and output the classification results.
7. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 6 is characterized in that: The process of constructing reverse contrast loss based on the extracted negative samples of different depths for reverse contrast learning includes: Based on the negative samples retrieved from the dictionary, a reverse contrast loss function is constructed, which is expressed as: in, is the negative sample prototype, and τ represents the temperature coefficient.
8. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 6 is characterized in that: A hierarchical bidirectional contrast loss function is set to optimize the parameters of the encoder network model. The expression is: Among them, L P (x i ) and L N (x i ) represent the prototype contrast loss function and the reverse contrast loss function, respectively, which are set in the shallow and deep layers of the encoder network at the same time, λ represents the weight term of the reverse contrast loss, and B represents the batch size.
9. The hyperspectral image classification method based on hierarchical bidirectional contrast learning according to claim 6, characterized in that: The expression for learning the latent representation in the hyperspectral remote sensing image data sample by the encoder network is: The expression for hyperspectral image classification using the learned potential representation through the initialized fully connected layer is: D=FC(L) Among them, x represents the test set sample, φ(·) represents the encoder network, represents the model parameters of the encoder network, D represents the classification result, and FC represents the initialized fully connected layer.
10. A hyperspectral image classification system based on hierarchical bidirectional contrast learning, characterized in that: The system comprises: A data acquisition module, used to acquire hyperspectral remote sensing image data; A data processing module, used for preprocessing the hyperspectral remote sensing image data; A representation extraction module is used to perform data enhancement on the preprocessed hyperspectral remote sensing image to construct a number of data enhancement samples, and extract online representation and target representation of the image data according to the constructed data enhancement samples; A model training module is used to calculate positive and negative samples of the online representation based on the target representation, and train the encoder network model according to the online representation and the positive and negative samples; The image classification module is used to classify the hyperspectral remote sensing image data to be classified using the encoder network model and output the classification result.
Citation Information
Patent Citations
Multi-modal remote sensing data ground feature classification method based on meta-learning
CN116704330A
Sliding window weighted sparse expansion network method for hyperspectral unmixing
CN118887108A
Artificial intelligence based image caption creation systems and methods thereof
US10713830B1