Hyperspectral image classification method and system based on spectral domain perception and uncertainty regulation
The band importance weight is obtained through noise removal and spectral domain perception mechanisms, combined with dynamic focus and uncertainty evaluation units, an adaptive adjustment mechanism is built, which solves the problems of high-dimensional, complex features and data imbalance in hyperspectral image classification, and achieves high-precision and efficient classification effects.
Patent Information
- Application Number
- CN202510764121.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
When existing hyperspectral image classification methods deal with high-dimensional, complex features, data imbalance and noise interference, it is difficult to achieve efficient and accurate classification effects. In particular, traditional methods rely on artificial design features and deep learning models are prone to bias towards most categories during training, resulting in a decrease in the classification accuracy of a few categories.
The importance weight of the band is obtained through noise removal and spectral domain perception mechanisms, combined with dynamic focus units and uncertainty evaluation units, an adaptive adjustment mechanism is built, and a multi-layer perceptron is used for feature extraction and classification, and the loss function is dynamically regulated to adapt to hyperspectral image data of different scenarios.
It improves classification accuracy and computing efficiency, enhances the adaptability and generalization capabilities of the model, can maintain high performance under complex situations, and is suitable for various practical hyperspectral image classification tasks.
Smart Images

Figure CN120279428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and analysis, and particularly to a hyperspectral image classification method and system based on spectral domain perception and uncertainty regulation. Background Art
[0002] Hyperspectral images have demonstrated unique value and broad application prospects in numerous fields due to their rich spectral information, such as crop monitoring in agricultural production, mineral identification in geological exploration, ecological assessment in environmental science, and target detection in military reconnaissance. However, the crucial task of hyperspectral image classification faces numerous intractable problems, severely restricting the full realization of its application effectiveness.
[0003] Traditional machine learning methods, such as support vector machines (SVMs), decision trees, etc., have obvious limitations when dealing with hyperspectral images. These methods usually rely on manually designed features, and it is difficult to extract sufficiently comprehensive and discriminative features for data like hyperspectral images, which are high-dimensional, complex, and contain rich spectral details. For example, when facing large-scale hyperspectral datasets, the selection and parameter tuning of the kernel function of SVM become extremely complex, with high computational costs, and due to insufficient feature extraction, the classification accuracy is difficult to reach a satisfactory level. Decision tree methods are prone to overfitting, especially in high-dimensional data, being sensitive to noise, and the decision rules they generate often fail to accurately capture the complex relationship between spectral features and ground object categories, resulting in poor reliability and stability of the classification results.
[0004] In recent years, with the rise of deep learning techniques, some neural network-based methods have been applied to the field of hyperspectral image classification in an attempt to break through the bottlenecks of traditional methods. Recurrent neural networks (RNNs) and their variants, although having the ability to process sequential data, have difficulty effectively learning long-term dependencies when dealing with the long-sequence spectral data of hyperspectral images due to the problems of gradient vanishing or gradient explosion, and their training process is complex and time-consuming, with low processing efficiency for high-dimensional data. In addition, unsupervised learning methods such as autoencoders (AEs) and their derivative variational autoencoders (VAEs) can perform feature extraction and dimensionality reduction on hyperspectral data, but in classification tasks, they often need to be combined with other classifiers to achieve the final class judgment, increasing the complexity and uncertainty of the model, and their feature extraction is not highly targeted, with limited ability to capture key features for specific classification tasks, making it difficult to be directly and effectively applied to complex hyperspectral image classification scenarios.
[0005] Meanwhile, there is also a problem of data imbalance in hyperspectral images, that is, there are significant differences in the number of samples among different ground object categories. This makes traditional classification methods and some conventional deep learning models tend to be biased towards samples of the majority class during the training process, resulting in a serious decline in the classification accuracy of the minority class and a significant reduction in the overall classification performance. In addition, hyperspectral images are inevitably interfered by various noises during the acquisition and transmission processes, including sensor noise, atmospheric noise, etc. These noises will obscure the true spectral characteristics of ground objects, further increasing the difficulty of classification and reducing the accuracy and reliability of classification.
[0006] In summary, although various methods have been tried and applied to hyperspectral image classification, there is still a lack of a classification method that can efficiently and accurately handle problems such as the high-dimensionality, complex features, data imbalance, and noise interference of hyperspectral images. There is an urgent need for an innovative solution to meet the growing actual application requirements and promote the in-depth development and wide application of hyperspectral image technology in various fields. Summary of the Invention
[0007] In view of the deficiencies of the prior art, the purpose of the present invention is to provide a hyperspectral image classification method based on spectral domain perception and uncertainty regulation. Through noise removal and spectral domain perception mechanisms, spectral feature analysis is performed on each band of the hyperspectral image to obtain the information entropy of each band and the correlation coefficient with the class label. Based on this, the importance weight of the band is obtained, focusing on the key spectral bands and reducing the interference of redundant information such as the bands where noise is located. Moreover, the data imbalance of hyperspectral images is solved through an uncertainty evaluation unit and an adaptive adjustment mechanism.
[0008] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A hyperspectral image classification method based on spectral domain perception and uncertainty regulation, comprising: Establishing a hyperspectral image set including a number of hyperspectral images; Performing noise removal processing on each hyperspectral image to obtain a first image; Performing normalization processing on the first image to obtain a second image; Obtaining the importance weight of each band of the second image based on the spectral domain perception mechanism; Obtaining the feature tensor of the second image, constructing a dynamic focusing unit, using a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time, combining the attention mechanism to perform dynamic weighted summation on the projected feature tensor based on the importance weight to obtain a dynamic focusing feature tensor; Construct a feature enhancement sub-module, use a multi-layer perceptron to perform non-linear transformation and enhancement on the dynamic focusing feature tensor and conduct model training, construct an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron, and construct an adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron; Use the trained multi-layer perceptron to output the classification result.
[0009] Furthermore, the normalization method is as follows: Perform Min-Max linear operation on the first image, and map the pixel values of the first image to the interval [0, 1] through linear transformation.
[0010] Furthermore, the method for obtaining the importance weight of each band of the second image based on the spectral domain perception mechanism is as follows: Obtain the information entropy of each band, and the calculation formula is:
[0011] Among them, represents the number of different value types of pixel values in a certain band of the second image, is the index variable for traversing pixel values, is the pixel value, is the pixel value in the band is the probability of is the information entropy of each band, is the logarithmic function with base 10; Adopt the Pearson correlation coefficient calculation method to obtain the correlation coefficient of the class label of each band; Normalize the information entropy and the correlation coefficient of the class label of each band, and obtain the importance weight of each band through weighted combination of the normalized information entropy and the correlation coefficient of the class label.
[0012] Furthermore, construct a dynamic focusing unit. The method for using a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time combining the attention mechanism to perform dynamic weighted summation on the projected feature tensor based on the importance weight to obtain a dynamic focusing feature tensor is specifically as follows: Create a query projection layer, a key projection layer, and a value projection layer; Obtain the projected feature tensor by passing the feature tensor of the second image through the corresponding linear projection layer. The projected feature tensor includes a query tensor, a key tensor, and a value tensor. Among them, the feature tensor of the second image is [B, N, C], and the projected feature tensor is [B, N, 64]. B is the batch size, N is the length of the one-dimensional sequence obtained by expanding the two-dimensional spatial information of the second image, and C is the number of bands; Perform a weighted operation on the value tensor according to the importance weights to obtain a weighted value tensor; Obtain the attention energy, and the calculation formula is:
[0013] where, is the attention energy, which is used to measure the correlation between different features, is the batch matrix multiplication, is the transpose operation on the value tensor, represents the number of different value types of pixels in a certain band of the second image; Obtain the attention weights through the softmax function, and the calculation formula is:
[0014] where, is the attention weight, is to convert each element of to a probability value, is the last dimension of the projected feature tensor; Multiply the attention weights by the weighted value tensor to obtain the dynamic focusing feature tensor, and perform Token pruning operation on the dynamic focusing feature tensor according to the preset Token pruning ratio, select a certain proportion of bands with higher importance scores to retain, and fill the dynamic focusing feature tensor after Token pruning through zero-value filling operation.
[0015] Furthermore, the method for performing non-linear transformation and enhancement on the dynamic focusing feature tensor using a multi-layer perceptron and for model training is: Construct a three-layer multi-layer perceptron. The first layer is the input layer, and the number of nodes in the input layer is the same as the dimension of the dynamic focusing feature tensor; the second layer is the hidden layer and uses ReLU as the activation function; the third layer is the output layer.
[0016] Furthermore, the method for constructing an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron is: At initialization, initialize the relevant parameters according to the input prior probability, class weight, focus weight, loss function type, and warm-up rounds, and transfer the prior probability and class weight to the image processing unit; During the forward propagation process, determine the type of loss function to be used according to the current training round, specifically: When , use the BCE loss function to calculate the loss, and the calculation formula is:
[0017] where, is the value of the BCE loss function, is the first parameter, is the second parameter, represents the number of warm-up rounds, represents the current round number, is the logarithm function with base 10; For the positive sample loss, the calculation formula is:
[0018] where, is the predicted value of the multi-layer perceptron, is the positive sample mask, is the positive sample loss corresponding to the BCE loss function; For the negative sample loss, the calculation formula is:
[0019] where, is the negative sample mask, is the negative sample loss corresponding to the BCE loss function; For the unlabeled sample loss, the calculation formula is:
[0020] where, is the unlabeled sample mask, is the unlabeled sample loss corresponding to the BCE loss function; When is true, use the sigmoid loss function to calculate the loss; For the positive sample loss, the calculation formula is:
[0021] where, , is the model predicted value of the th positive sample, N is the number of positive samples, is the positive sample loss corresponding to the sigmoid loss function, is the natural exponential function; For the negative sample loss, the calculation formula is:
[0022] where, , is the model predicted value of the th negative sample, M is the number of negative samples, is the negative sample loss corresponding to the sigmoid loss function; For the unlabeled sample loss, the calculation formula is:
[0023] Among them, , is the model prediction value of the th unlabeled sample, K is the number of unlabeled samples, is the loss of the unlabeled sample corresponding to the sigmoid loss function; When calculating the final loss, when the loss of the unlabeled sample does not need to be considered, the calculation formula is:
[0024] Among them, is the positive sample loss, is the negative sample loss. When , is , is ; when , is , is ; When the influence of the loss of the unlabeled sample on the final loss needs to be considered according to requirements, it is added to the final loss through a certain weighting method:
[0025] Among them, is the weight coefficient of the loss of the unlabeled sample. When , is ; when , is , is the weight.
[0026] Furthermore, the method for adaptively adjusting the model parameters of the multi-layer perceptron by constructing an adaptive adjustment mechanism is: When it is monitored that the accuracy rate of the model on the validation set has not increased for several consecutive rounds or the loss value starts to rise, it is judged that the model has an overfitting trend, and the learning rate is multiplied by the decay factor, the weight of L2 regularization is increased, and at the same time, the sampling ratio of difficult samples is increased.
[0027] A hyperspectral image classification system based on spectral domain perception and uncertainty regulation, comprising: An image acquisition module, configured to establish a hyperspectral image set including several hyperspectral images; A first processing module, configured to perform noise removal processing on each hyperspectral image to obtain a first image; A second processing module for normalizing the first image to obtain a second image; A third processing module for obtaining the importance weights of each band of the second image based on a spectral domain perception mechanism; A fourth processing module for obtaining the feature tensor of the second image, constructing a dynamic focusing unit, using a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time combining an attention mechanism to perform dynamic weighted summation on the projected feature tensor based on the importance weights to obtain a dynamic focusing feature tensor; A model training module for constructing a feature enhancement sub-module, using a multi-layer perceptron to perform non-linear transformation and enhancement on the dynamic focusing feature tensor and performing model training, constructing an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron, and constructing an adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron; A result output module for outputting a classification result using the trained multi-layer perceptron.
[0028] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned hyperspectral image classification method based on spectral domain perception and uncertainty regulation.
[0029] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned hyperspectral image classification method based on spectral domain perception and uncertainty regulation.
[0030] Compared with the prior art, the present invention has the following advantages and beneficial effects: The present invention can improve the classification accuracy: the spectral domain perception mechanism can accurately focus on the key spectral bands, effectively extract discriminative features, reduce the interference of irrelevant spectral information, and at the same time the multi-layer perceptron further improves the feature expression ability, enabling the model to more accurately distinguish different ground object categories. The adaptive adjustment mechanism optimizes the model training process by reasonably evaluating and adjusting the sample risk, enhances the adaptability of the model to unbalanced data, and further improves the classification accuracy. Combining the advantages of both, this method can achieve a high accuracy rate in the hyperspectral image classification task, can more accurately identify various ground object types in practical applications, and provides a more reliable decision-making basis for related fields.
[0031] The present invention can improve the computing efficiency: The dynamic focusing unit effectively reduces the data dimension and computational amount through the projection and Token pruning operations on the feature tensor. Meanwhile, the multi-layer perceptron structure has a lower computational complexity compared to the complex convolutional neural network structure while ensuring the feature enhancement effect. The uncertainty control module avoids unnecessary computational overhead by dynamically adjusting the training parameters, improving the training efficiency. Therefore, the method of the present invention can significantly shorten the training time and inference time of the model when processing large-scale hyperspectral image data, improve the computing efficiency, and meet the requirements of real-time and fast processing in practical applications. For example, in scenarios such as real-time UAV monitoring and online environmental monitoring, accurate classification results can be quickly given to provide support for timely decision-making.
[0032] The present invention can enhance the model adaptability and generalization ability: During the model training process, through the efficient extraction and optimization of features by the dynamic focusing unit and the uncertainty evaluation unit, the model can better adapt to hyperspectral image data collected from different scenarios and different sensors, and has stronger robustness to the diversity and complexity of the data. Even in the face of complex situations such as changes in data distribution and large noise interference, the model of the present invention can still maintain a high classification performance, reduce the occurrence of overfitting phenomenon, has better generalization ability, and can be widely applied to various actual hyperspectral image classification tasks, providing strong technical support for research and applications in different fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a flowchart of a hyperspectral image classification method based on spectral domain perception and uncertainty control according to the present invention; Figure 2 is an overall flowchart diagram of a hyperspectral image classification method based on spectral domain perception dynamic focusing and uncertainty performance control according to the present invention; Figure 3 is a recognition result diagram of water in a specific embodiment of the present invention; Figure 4 is a recognition result diagram of strawberries in a specific embodiment of the present invention; Figure 5 is a recognition result diagram of a road in a specific embodiment of the present invention; Figure 6 is a recognition result diagram of soybeans in a specific embodiment of the present invention; Figure 7 is a recognition result diagram of a melon field in a specific embodiment of the present invention; Figure 8Recognition result diagram of water spinach in a specific embodiment of the present invention; Figure 9 Schematic diagram of a hyperspectral image classification system based on spectral domain perception and uncertainty regulation of the present invention. Specific implementation manner
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0035] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.
[0036] Embodiment 1 Embodiment 1 provides a hyperspectral image classification method based on spectral domain perception and uncertainty regulation, as Figure 1 shown, including: Step S1: Establish a hyperspectral image set including several hyperspectral images; Step S2: Perform noise removal processing on each hyperspectral image to obtain a first image; Step S3: Perform normalization processing on the first image to obtain a second image; Step S4: Obtain the importance weight of each band of the second image based on the spectral domain perception mechanism; Step S5: Obtain the feature tensor of the second image, construct a dynamic focusing unit, project the feature tensor of the second image using a linear projection layer to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time, based on the importance weight, perform dynamic weighted summation on the projected feature tensor through an attention mechanism to obtain a dynamic focusing feature tensor; Step S6: Construct a feature enhancement sub-module, perform non-linear transformation and enhancement on the dynamic focusing feature tensor using a multi-layer perceptron and perform model training, construct an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron, and construct an adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron; Step S7: Output the classification result using the trained multi-layer perceptron.
[0037] A hyperspectral image classification method based on spectral domain perception and uncertainty regulation provided by this embodiment effectively improves the performance of hyperspectral image classification through unique model design and training strategies, overcomes the deficiencies of the prior art, analyzes the spectral features of each band of the hyperspectral image through noise removal and spectral domain perception mechanisms, obtains the information entropy of each band and the correlation coefficient with the class label, and obtains the importance weight of the band based on this, focuses on the key spectral bands, reduces the interference of redundant information such as the bands where noise is located, and solves the data imbalance of hyperspectral images through the uncertainty evaluation unit and the adaptive adjustment mechanism. In step S2 of this embodiment, as Figure 2 shown, perform noise removal processing on the obtained hyperspectral image, adopt the median filtering method, traverse and filter the image with a 3x3 filtering window to remove the noise points in the image and reduce the interference of noise on image feature extraction and classification.
[0038] In step S3 of this embodiment, as Figure 2 shown, the normalization method is: perform Min-Max linear operation on the first image, map the pixel values of the first image to the interval [0, 1] through linear transformation, and the calculation formula is:
[0039] where, is the original pixel value, and are the minimum and maximum pixel values of this band in the image respectively, is the pixel value after normalization.
[0040] In step S4 of this embodiment, design a spectral domain perception mechanism (Spectral Perception Mechanism, SPM), calculate the information entropy of each band of the second image and the correlation coefficient with the class label by analyzing the spectral features of each band of the second image, and obtain the importance weight of the band based on this.
[0041] As Figure 2 shown, the method for obtaining the importance weight of each band of the second image based on the spectral domain perception mechanism is: Obtain the information entropy of each band, and the calculation formula is:
[0042] where, represents the number of different value types of pixel values in a certain band of the second image, is the index variable for traversing pixel values, is the pixel value, is the pixel value in the band probability, is the information entropy of each band, is the logarithmic function with base 10; The Pearson correlation coefficient calculation method is used to obtain the correlation coefficient of the class label for each band; The information entropy and the correlation coefficient of the class label for each band are normalized, and the normalized information entropy and correlation coefficient are combined by weighted combination to obtain the importance weight for each band. For example, in the experiment, the information entropy weight is 0.4 and the correlation coefficient weight is 0.6.
[0043] In step S5 of this embodiment, based on the above spectral domain perception results, the feature tensor of the second image is obtained, a Dynamic Focusing Unit (DFU) is constructed, and a linear projection layer (nn.Linear) is used to project the feature tensor of the second image, mapping it from the original dimension to a low-dimensional space (the embedding dimension is set to 64). At the same time, combined with the attention mechanism, the projected feature tensor is dynamically weighted and summed according to the importance weight of each band, enabling the model to focus on key spectral bands and feature regions during the feature extraction process, reducing the interference of redundant information, and improving the efficiency and accuracy of feature extraction.
[0044] The feature tensor of the second image is a spectral feature sequence after flattening the two-dimensional space information. The method for obtaining the feature tensor of the second image is as follows: The second image is represented as [B, H, W, C], where B is the batch size, H is the height, W is the width, and C is the number of bands. After flattening the spatial dimension, the feature tensor is obtained, and the feature tensor is represented as [B, N, C], where N = H × W, and N is the length of the one-dimensional sequence obtained by expanding the two-dimensional space information of the second image.
[0045] As Figure 2 shown, the method for constructing a dynamic focusing unit, using a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time, combined with the attention mechanism to dynamically weight and sum the projected feature tensor based on the importance weight to obtain a dynamic focusing feature tensor is specifically as follows: Create a query projection layer, a key projection layer, and a value projection layer. The query projection layer, the key projection layer, and the value projection layer are all linear projection layers (nn.Linear); The feature tensor of the second image passes through the corresponding linear projection layer to obtain a projected feature tensor, which includes a query tensor, a key tensor, and a value tensor. Among them, the feature tensor of the second image is [B, N, C], and the projected feature tensor is [B, N, 64]. B is the batch size, N is the length of unfolding the two-dimensional spatial information of the second image into a one-dimensional sequence, and C is the number of bands; Perform a weighted operation on the value tensor (V) according to the importance weight to obtain a weighted value tensor; Obtain the attention energy, and the calculation formula is:
[0046] Among them, is the attention energy, which is used to measure the correlation between different features, is the batch matrix multiplication, is the transpose operation on the value tensor; Obtain the attention weight through the softmax function, and the calculation formula is:
[0047] Among them, is the attention weight, is to convert each element of to a probability value, is the last dimension of the projected feature tensor; Multiply the attention weight by the weighted value tensor to obtain a dynamic focusing feature tensor, and perform a Token pruning operation on the dynamic focusing feature tensor according to a preset Token pruning ratio. Select 50% of the bands with higher importance scores to be retained, and fill the dynamic focusing feature tensor after Token pruning through zero-value filling operation.
[0048] In step S6 of this embodiment, as Figure 2 shown, construct a feature enhancement sub-module (FeatureEnhancementSubmodule, FES), and use a multi-layer perceptron (MLP) structure to perform non-linear transformation and enhancement on the features after dynamic focusing. The MLP consists of multiple fully connected layers and activation functions (ReLU). By performing multi-layer linear transformation and non-linear activation on the features, the expression ability and discriminability of the features are further improved, making it more conducive to subsequent classification tasks.
[0049] The method of using a multi-layer perceptron to perform non-linear transformation and enhancement on the dynamic focusing feature tensor and perform model training is: Construct a three-layer multi-layer perceptron. The first layer is the input layer, and the number of nodes in the input layer is the same as the dimension of the dynamic focusing feature tensor; the second layer is the hidden layer and uses ReLU as the activation function; the third layer is the output layer.
[0050] Specifically, construct a three-layer multi-layer perceptron (MLP). The first layer is the input layer, and the number of nodes is the same as the feature dimension after dynamic focusing (such as 64); the second layer is the hidden layer, and the number of nodes is set to 128, using ReLU as the activation function; the third layer is the output layer, and the number of nodes is also 64. The features after the non-linear transformation of the hidden layer are output as the enhanced feature tensor to further improve the expression ability and discriminability of the features.
[0051] In step S6 of this embodiment, as Figure 2 shown, design an uncertainty assessment unit (Risk Assessment Unit, RAU), and dynamically adjust the model training process of the multi-layer perceptron according to the set risk assessment strategy, combining prior knowledge and sample weights.
[0052] The method for constructing the uncertainty assessment unit to dynamically adjust the loss function of the multi-layer perceptron is as follows: During initialization, initialize the relevant parameters according to the input prior probability, class weight, focal weight, loss function type, and warm-up rounds, and transfer the prior probability and class weight to the image processing unit. For example, the prior probability (prior) is 0.3769, the class weight (class_weight) is 0.5, the focal weight (focal_weight) is 2, and the warm-up rounds (warm_up) is 50; During the forward propagation process, determine the type of loss function used according to the current training round, including: When it is the case, use the BCE (binary cross-entropy) loss function (binary cross-entropy) to calculate the loss, and the calculation formula is:
[0053] where is the value of the BCE loss function, is the first parameter, is the second parameter, represents the warm-up rounds, represents the current round; For the positive sample loss, the calculation formula is:
[0054] where is the prediction value of the multi-layer perceptron, is the positive sample mask, When it is 0, it indicates that it is not a positive sample. When it is 1, it indicates a positive sample. It is the positive sample loss corresponding to the BCE loss function. For the negative sample loss, the calculation formula is:
[0055] Where is the negative sample mask, is the negative sample loss corresponding to the BCE loss function; For the unlabeled sample loss, the calculation formula is:
[0056] Where is the unlabeled sample mask, is the unlabeled sample loss corresponding to the BCE loss function; When , the sigmoid loss function is used to calculate the loss; For the positive sample loss, the calculation formula is:
[0057] Where , is the model prediction value of the -th positive sample, N is the number of positive samples, is the positive sample loss corresponding to the sigmoid loss function, is the natural exponential function; For the negative sample loss, the calculation formula is:
[0058] Where , is the model prediction value of the -th negative sample, M is the number of negative samples, is the negative sample loss corresponding to the sigmoid loss function; For the unlabeled sample loss, the calculation formula is:
[0059] Where , is the model prediction value of the -th unlabeled sample, K is the number of unlabeled samples, is the unlabeled sample loss corresponding to the sigmoid loss function; When calculating the final loss, when the unlabeled sample loss does not need to be considered, the calculation formula is:
[0060] Among them, is the positive sample loss, is the negative sample loss. When at this time, is , is . When at this time, is , is , is the weight; When considering the impact of the unlabeled sample loss on the final loss as needed, it is added to the final loss through a certain weighting method:
[0061] Among them, is the weight coefficient of the unlabeled sample loss. When at this time, is . When at this time, is .
[0062] In step S6 of this embodiment, as Figure 2 shown, an Adaptive Adjustment Mechanism (AAM) is constructed, and according to the output result of the uncertainty evaluation unit, training parameters such as the learning rate and parameter update step size of the multi-layer perceptron model are adaptively adjusted.
[0063] The method for constructing the adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron is as follows: When it is monitored that the accuracy of the model on the validation set has not improved for several consecutive rounds or the loss value starts to rise, it is determined that the model has an overfitting trend. Multiply the learning rate by the decay factor, increase the weight of ridge regression (L2 regularization), and at the same time increase the sampling ratio of difficult samples.
[0064] Multiple monitoring metrics are set, such as the accuracy of the model on the validation set, the changing trend of the loss value, etc. When it is monitored that the accuracy of the model on the validation set has not improved for several consecutive rounds or the loss value starts to rise, it is judged that the model may have an overfitting trend. At this time, the adaptive adjustment mechanism takes the following measures: multiply the learning rate by a decay factor (such as 0.8), increase the weight of L2 regularization (such as increasing from 0.0005 to 0.001), and at the same time adjust the sampling strategy of the samples, for example, increase the sampling ratio of difficult samples (from 0.2 to 0.3), to optimize the training process of the model, ensure that the model can maintain good performance and stability at different training stages, avoid falling into local optima, and improve the robustness and generalization ability of the model.
[0065] To verify the effectiveness of the method of the present invention, based on the above technical solution, a simulation experiment was carried out in this embodiment, and the specific result analysis is as follows: 1. Experimental images In this embodiment, the proposed hyperspectral image classification method based on spectral domain perception and uncertainty regulation is experimented on the WHU-Hi-HanChuan public dataset to verify the effectiveness and reliability of the method.
[0066] The hyperspectral image of the HanChuan dataset has 274 spectral channels, and its wavelength span is from 400 nanometers to 1000 nanometers. The spatial resolution of the generated image is as high as 0.109 meters, and the size specification is 1217x303. In view of the more significant shadow situation, this dataset presents more prominent spectral diversity characteristics. When implementing the detection procedure, six ground objects were specifically selected in the present invention to ensure that there are no unclear markings, thereby ensuring the accuracy and reliability of the detection.
[0067] 2. Experimental methods and related parameter settings The random gradient descent algorithm was selected for the experiment, and the Adam (Adaptive Moment Estimation) optimizer was used, and the entire training process went through 1000 times. The learning rate was 0.001, and the batch size was set to 16. Among them, and The default values of were determined to be 0.3 and 0.1 respectively. And the proportion of all experimental test sets was always 0.9.
[0068] The main evaluation metrics include precision, recall, and Precision is mainly used to measure the accuracy of a classifier in identifying true positive samples among all the samples it classifies as positive. In contrast, recall is used to evaluate the ability of a classifier to detect all actual positive samples. The F1 score, which is obtained by balancing precision and recall, is a more appropriate metric for evaluating the performance of a single-class classification task and can comprehensively and objectively reflect the quality of the completion of the relevant task.
[0069] 3. Comparison of Experimental Results Table 1 - F1 Test Results of HanChuan Dataset
[0070] Table 2 - Precision / Recall Test Results of HanChuan Dataset
[0071] As Figures 3 - 8 shown, the recognition results of water, strawberries, roads, soybeans, melon fields, and water spinach are presented. As shown in Table 1, the average F1 value of this method for 6 types of ground objects on the HanChuan dataset reaches 92.91%, and the F1 value of the road classification task is the highest (95.71%), verifying the ability of the dynamic focusing unit to extract highly discriminative spectral features. Table 2 further shows that the average precision and recall of this method are 91.74% and 94.13% respectively, and it performs particularly well in the road (recall rate 98.52%) and water body (recall rate 95.31%) classification tasks, reflecting the dynamic weighting advantage of the spectral domain perception mechanism for key bands.
[0072] Example 2 Example 2 provides a hyperspectral image classification system based on spectral domain perception and uncertainty regulation, which is applied to the above-mentioned hyperspectral image classification method based on spectral domain perception and uncertainty regulation. As Figure 9 shown, it includes: An image acquisition module for establishing a hyperspectral image set including several hyperspectral images; A first processing module for performing noise removal processing on each hyperspectral image to obtain a first image; A second processing module for performing normalization processing on the first image to obtain a second image; A third processing module for obtaining the importance weights of each band of the second image based on the spectral domain perception mechanism; A fourth processing module for obtaining the feature tensor of the second image, constructing a dynamic focusing unit, using a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time combining the attention mechanism to perform dynamic weighted summation on the projected feature tensor based on the importance weights to obtain a dynamic focusing feature tensor; A model training module, configured to construct a feature enhancer sub-module, perform non-linear transformation and enhancement on the dynamic focusing feature tensor using a multi-layer perceptron and conduct model training, construct an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron, and construct an adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron; A result output module, configured to output a classification result using the trained multi-layer perceptron.
[0073] Embodiment 3 Embodiment 3 provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned hyperspectral image classification method based on spectral domain perception and uncertainty regulation.
[0074] Embodiment 4 Embodiment 4 provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned hyperspectral image classification method based on spectral domain perception and uncertainty regulation.
[0075] The memory in the embodiments of the present invention is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program for operating on the electronic device.
[0076] The hyperspectral image classification method based on spectral domain perception and uncertainty regulation disclosed in the embodiments of the present invention can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the hyperspectral image classification method based on spectral domain perception and uncertainty regulation can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention, it can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the hyperspectral image classification method based on spectral domain perception and uncertainty regulation provided in the embodiments of the present invention.
[0077] In an exemplary embodiment, the electronic device may be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components for performing the foregoing method.
[0078] It can be understood that the memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, RandomAccessMemory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, SynchronousDynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDRSDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random AccessMemory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random AccessMemory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory). The memory described in the embodiments of the present invention is intended to include but not be limited to these and any other suitable types of memory.
[0079] The above embodiments are merely illustrative examples of the technical solution of the present invention. The method involved in the present invention is not limited solely to the content described in the above embodiments, but is subject to the scope defined by the claims. Any modification, supplement, or equivalent replacement made by those skilled in the art to which the present invention pertains based on this embodiment falls within the scope protected by the claims of the present invention.
Claims
1. A hyperspectral image classification method based on spectral domain perception and uncertain regulation, characterized in that, Including: Construct a hyperspectral image set including several hyperspectral images; Perform noise removal processing on each hyperspectral image to obtain a first image; Perform normalization processing on the first image to obtain a second image; Obtain the importance weight of each band of the second image based on the spectral domain perception mechanism; Obtain the feature tensor of the second image, construct a dynamic focusing unit, use a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time combine the attention mechanism to perform dynamic weighted summation on the projected feature tensor based on the importance weight to obtain a dynamic focusing feature tensor; Construct a feature enhancement sub-module, use a multi-layer perceptron to perform non-linear transformation and enhancement on the dynamic focusing feature tensor and perform model training, construct an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron, and construct an adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron; Use the trained multi-layer perceptron to output the classification result.
2. The hyperspectral image classification method based on spectral domain perception and uncertainty regulation according to claim 1, wherein The method of normalization is: Perform Min-Max linear operation on the first image, and map the pixel values of the first image to the interval [0, 1] through linear transformation.
3. The hyperspectral image classification method based on spectral domain perception and uncertainty regulation according to claim 1, wherein, The method of obtaining the importance weight of each band of the second image based on the spectral domain perception mechanism is: Obtain the information entropy of each band, and the calculation formula is: Among them, represents the number of different value types of pixel values in a certain band of the second image, is an index variable for traversing pixel values, is the pixel value, is the pixel value in the band probability, is the information entropy of each band, is the logarithmic function with base 10; Use the Pearson correlation coefficient calculation method to obtain the correlation coefficient of the class label of each band; Normalize the information entropy and the correlation coefficient of the class label of each band, and obtain the importance weight of each band by weighted combination of the normalized information entropy and the correlation coefficient of the class label.
4. The hyperspectral image classification method based on spectral domain perception and uncertainty regulation according to claim 1, wherein The method of constructing a dynamic focusing unit, using a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time combining the attention mechanism to perform dynamic weighted summation on the projected feature tensor based on the importance weight to obtain a dynamic focusing feature tensor is specifically: Create a query projection layer, a key projection layer and a value projection layer; Obtain the projected feature tensor by passing the feature tensor of the second image through the corresponding linear projection layer. The projected feature tensor includes a query tensor, a key tensor and a value tensor. Among them, the feature tensor of the second image is [B, N, C], and the projected feature tensor is [B, N, 64]. B is the batch size, N is the length of expanding the two-dimensional spatial information of the second image into a one-dimensional sequence, and C is the number of bands; Perform a weighted operation on the value tensor according to the importance weight to obtain a weighted value tensor; Obtain the attention energy, and the calculation formula is: Among them, is the attention energy, which is used to measure the correlation between different features, is the batch matrix multiplication, is the transpose operation on the value tensor, represents the number of different value types of pixel values in a certain band of the second image; Obtain the attention weight through the softmax function, and the calculation formula is: Among them, is the attention weight, is to convert each element of to a probability value, where is the last dimension of the projected feature tensor; It should be noted that the original text seems a bit fragmented in its logical expression. The above translation tries to make sense of it based on the overall context. If there are inaccuracies, it may be due to the somewhat unclear nature of the original Chinese text. Multiply the attention weight by the weighted value tensor to obtain a dynamic focusing feature tensor, and perform Token pruning operation on the dynamic focusing feature tensor according to the preset Token pruning ratio, select a certain proportion of bands with higher importance scores to keep, and perform padding operation on the dynamic focusing feature tensor after Token pruning through zero-value filling.
5. The hyperspectral image classification method based on spectral domain perception and uncertainty regulation according to claim 1, characterized in that The method of using a multi-layer perceptron to perform non-linear transformation and enhancement on the dynamic focusing feature tensor and perform model training is: Construct a three-layer multi-layer perceptron. The first layer is the input layer, and the number of nodes in the input layer is the same as the dimension of the dynamic focusing feature tensor. The second layer is the hidden layer and uses ReLU as the activation function. The third layer is the output layer.
6. The hyperspectral image classification method based on spectral domain perception and uncertainty regulation according to claim 1, wherein The method for constructing an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron is as follows: During initialization, relevant parameters are initialized according to the input prior probability, class weight, focus weight, loss function type, and warm-up rounds, and the prior probability and class weight are transferred to the image processing unit. During the forward propagation process, the type of loss function to be used is determined according to the current training round, specifically: When the BCE loss function is used to calculate the loss, and the calculation formula is: Among them, is the BCE loss function value, is the first parameter, is the second parameter, represents the number of warm-up rounds, represents the current round number, is the logarithmic function with base 10; For the positive sample loss, the calculation formula is: Among them, is the predicted value of the multi-layer perceptron, is the positive sample mask, is the positive sample loss corresponding to the BCE loss function; For the negative sample loss, the calculation formula is: Among them, is the negative sample mask, is the negative sample loss corresponding to the BCE loss function; For the unlabeled sample loss, the calculation formula is: Among them, is the unlabeled sample mask, is the loss of the unlabeled sample corresponding to the BCE loss function; When the sigmoid loss function is used to calculate the loss; For the positive sample loss, the calculation formula is: Among them, , is the model prediction value of the th positive sample, N is the number of positive samples, is the loss of the positive sample corresponding to the sigmoid loss function, is the natural exponential function; For the negative sample loss, the calculation formula is: Among them, , is the model prediction value of the th negative sample, M is the number of negative samples, is the loss of the negative sample corresponding to the sigmoid loss function; For the unlabeled sample loss, the calculation formula is: Among them, , is the model prediction value of the -th unlabeled sample, K is the number of unlabeled samples, is the loss of the unlabeled sample corresponding to the sigmoid loss function; When calculating the final loss, when the unlabeled sample loss does not need to be considered, the calculation formula is: Among them, is the positive sample loss, is the negative sample loss. When holds, is , is . When holds, is , is , is the weight; When the influence of the unlabeled sample loss on the final loss needs to be considered according to the need, it is added to the final loss through a certain weighting method: Among them, is the weight coefficient of the loss of unlabeled samples. When holds,[[]] is . When holds,[[]] is .
7. The hyperspectral image classification method based on spectral domain perception and uncertainty regulation according to claim 1, characterized in that The method for constructing an adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron is as follows: When it is monitored that the accuracy of the model on the validation set has not improved for several consecutive rounds or the loss value begins to rise, it is judged that the model has an overfitting trend, and the learning rate is multiplied by the decay factor, the weight of ridge regression is increased, and at the same time, the sampling ratio of difficult samples is increased.
8. A hyperspectral image classification system based on spectral domain perception and uncertain regulation, characterized in that, Including: An image acquisition module for establishing a hyperspectral image set including several hyperspectral images; A first processing module for performing noise removal processing on each hyperspectral image to obtain a first image; A second processing module for performing normalization processing on the first image to obtain a second image; A third processing module for obtaining the importance weight of each band of the second image based on the spectral domain perception mechanism; A fourth processing module for obtaining the feature tensor of the second image, constructing a dynamic focusing unit, using a linear projection layer to project the feature tensor of the second image to map the original dimension to a low-dimensional space to obtain a projected feature tensor, and at the same time, combining the attention mechanism to perform dynamic weighted summation on the projected feature tensor based on the importance weight to obtain a dynamic focusing feature tensor; A model training module for constructing a feature enhancement sub-module, using a multi-layer perceptron to perform non-linear transformation and enhancement on the dynamic focusing feature tensor and performing model training, constructing an uncertainty evaluation unit to dynamically regulate the loss function of the multi-layer perceptron, and constructing an adaptive adjustment mechanism to adaptively adjust the model parameters of the multi-layer perceptron; A result output module for outputting classification results using the trained multi-layer perceptron.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the hyperspectral image classification method based on spectral domain perception and uncertainty regulation as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the hyperspectral image classification method based on spectral domain perception and uncertainty regulation as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Unmanned aerial vehicle hyperspectral vegetation species classification method and system based on deep learning
CN117935090A
Hyperspectral image dimension reduction method based on tensor graph embedding
CN118097324A
Remote sensing image small sample semantic segmentation method based on spectral super-resolution reconstruction
CN118154427A
Superobject information-based remote sensing image target extraction method, device, electronic apparatus, and medium
WO2020232905A1
Cited By
Self-adaptive fluorescent immune layer quantitative detection feature extraction method and system
CN120877013A
An Adaptive Feature Extraction Method and System for Quantitative Detection of Fluorescent Immunochromatograms
CN120877013B