Hyperspectral image classification method based on Retinex model
Features are extracted through Retinex decomposition and dual-branch network, and feature fusion is used to solve the problem of illumination changes and overfitting in hyperspectral image classification, achieving high-precision and low-complexity classification effects.
Patent Information
- Application Number
- CN202510103678.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Hyperspectral image classification faces problems such as light change influence, overfitting problems caused by finite labeled data, limitations of single feature extraction and high computational complexity.
Using a hyperspectral image classification method based on the Retinex model, the image is separated into inherent attribute components and illumination components through Retinex decomposition, features are extracted separately using a dual-branch network, and feature fusion is performed through interactive attention mechanism.
It effectively reduces the impact of lighting changes on classification performance, improves the generalization ability of the model, improves classification accuracy and robustness, and reduces the computational complexity.
Smart Images

Figure CN119942220A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and specifically relates to a hyperspectral image classification method based on a Retinex model. Background Art
[0002] Hyperspectral image classification is an important research direction in remote sensing science. It can accurately identify and classify objects on the ground by capturing the continuous reflection characteristics of objects on the ground in the entire electromagnetic spectrum. Hyperspectral images have the characteristics of many bands and large information redundancy, and can provide extremely detailed spectral information, providing unprecedented insights for scientific research and application practice.
[0003] However, hyperspectral image classification faces many challenges, especially in the case of changing lighting conditions and limited labeled data. The classification performance of existing technologies is often limited, mainly as follows: Effect of illumination changes: The classification performance of hyperspectral images is very sensitive to illumination conditions. Illumination changes can cause significant changes in the spectral characteristics of the image, thus affecting the accuracy of the classification model. Existing deep learning methods perform poorly when dealing with illumination changes, especially in the case of uneven illumination or complex illumination conditions, where the classification accuracy drops significantly.
[0004] Overfitting problem with limited labeled data: The acquisition of labeled data for hyperspectral images is costly and difficult, resulting in limited training data. Existing deep learning methods are prone to overfitting problems with limited labeled data, resulting in poor generalization of the model on unseen data and unsatisfactory classification results.
[0005] Limitations of single feature extraction: Existing hyperspectral image classification methods usually only focus on a single feature (such as spectral features or spatial features), while ignoring the complementary relationship between illumination features and inherent attribute features. This single feature extraction method limits the classification performance of the model, especially in complex scenes, and the classification accuracy is difficult to further improve.
[0006] High computational complexity: Existing deep learning methods usually require a lot of computing resources, especially when processing high-dimensional hyperspectral data, the computational complexity increases significantly, limiting its promotion in practical applications. Summary of the invention
[0007] In view of the above problems, the present invention provides a hyperspectral image classification method based on the Retinex model. The present invention can effectively extract spatial-spectral features in hyperspectral images, significantly improve the classification accuracy, has broad application prospects, and is suitable for hyperspectral image classification tasks under complex lighting conditions.
[0008] To achieve the above object, the technical solution adopted by the present invention is: A hyperspectral image classification method based on a Retinex model comprises the following steps: Perform Retinex decomposition on the hyperspectral image to obtain the intrinsic attribute component and the illumination component; The dual-branch network is used to extract features of the intrinsic attribute component and the illumination component respectively, wherein the intrinsic attribute component feature extractor is used to extract the material features and texture information of the hyperspectral image, and the illumination component feature extractor is used to extract the illumination features of the hyperspectral image; The feature fusion module based on the interactive attention mechanism is used to fuse the extracted intrinsic attribute features and illumination features to generate a fused feature representation. The fused features are classified to generate the classification results of the hyperspectral image.
[0009] The hyperspectral image classification method based on the Retinex model in this paper separates the hyperspectral image into the intrinsic attribute component and the illumination component through Retinex decomposition, extracts features respectively using a dual-branch network, and realizes feature fusion and classification through an interactive attention mechanism. It realizes the effective separation of illumination changes and intrinsic attributes in hyperspectral images, improves classification accuracy and robustness, and solves the impact of illumination changes on classification performance.
[0010] As a further improvement of the above solution, the Retinex decomposition step includes: Perform logarithmic transformation on each band of the hyperspectral image; The illumination component is obtained by filtering using Gaussian kernels of different scales, where the scale parameter σi of the Gaussian kernel is dynamically adjusted according to the local and global information of the hyperspectral image; Subtracting the illumination component from the logarithmically transformed image to obtain a reflectance image, wherein the reflectance image represents an intrinsic attribute component of the hyperspectral image; Perform weighted averaging on reflectance images of different scales to obtain the final intrinsic attribute component; Perform exponential transformation on the intrinsic attribute component and the illumination component and restore them to the linear space.
[0011] Technical effects of the above improvements: Retinex decomposition performs logarithmic transformation on each band of the hyperspectral image, uses multi-scale Gaussian kernel filtering to obtain the illumination component, then subtracts the illumination component from the logarithmic transformed image to obtain the reflectance image (intrinsic attribute component), and finally performs weighted averaging on the reflectance images of different scales and restores the linear space. Through multi-scale Gaussian kernel filtering and weighted averaging, the local and global information of the hyperspectral image are fully considered, the illumination component and the intrinsic attribute component are effectively separated, and the expression ability of the image features is enhanced.
[0012] As a further improvement of the above solution, the intrinsic attribute component feature extractor includes: The convolutional layer is used to reduce the dimension of the input intrinsic attribute components and generate the initial feature map; Local features are gradually extracted through multiple convolutional layers, where the convolutional layers include convolution kernels of different sizes; Capture global and local salient information through global pooling operations, and generate global average pooling vectors and global maximum pooling vectors; After concatenating the global average pooling vector and the global maximum pooling vector, the channel attention weight is generated through a fully connected layer and an activation function; The channel attention weights are multiplied element-wise with the original feature map to generate the enhanced feature map.
[0013] The technical effect of the above improvements is as follows: Through convolutional neural networks and channel attention mechanisms, the material features and texture information of hyperspectral images are extracted, the feature expression ability is enhanced, and the classification accuracy is improved.
[0014] As a further improvement of the above scheme, the illumination component feature extractor adopts the same deep learning architecture as the intrinsic attribute component feature extractor for feature extraction and is trained by a cross entropy loss function.
[0015] The technical effect of the above improvements is as follows: Through the same deep learning architecture, the illumination features of the hyperspectral image are extracted, ensuring the accuracy and consistency of the illumination feature extraction and improving the classification performance.
[0016] As a further improvement of the above scheme, the feature fusion module based on the interactive attention mechanism includes: The intrinsic attribute features and illumination features are converted into query vector Query, key vector Key and value vector Value through convolution layers, where the convolution kernel size is 3×3; Calculate the similarity between the query vector and the key vector through the dot product to generate the energy matrix; The energy matrix is normalized through the Softmax function to generate the attention weight matrix; Multiply the attention weight matrix by the value vector to generate a fused feature map; The fused feature map is added to the original feature map through residual connection, where the weight γ of the residual connection is a learnable parameter; The feature map after residual connection is weighted fused through two parallel multi-layer perceptrons (MLPs) to generate the final classification features.
[0017] The technical effect of the above improvements: The feature fusion module based on the interactive attention mechanism converts the inherent attribute features and illumination features into query vectors (Query), key vectors (Key) and value vectors (Value), respectively, calculates the similarity through dot product to generate an energy matrix, and then generates an attention weight matrix through the Softmax function. Finally, the attention weight matrix is multiplied with the value vector to generate a fused feature map, and the final classification features are generated through residual connection and weighted fusion. Through the interactive attention mechanism, the weights of the inherent attribute features and illumination features are dynamically adjusted to achieve information complementarity, improve the effect of feature fusion, and thus improve classification accuracy.
[0018] As a further improvement of the above scheme, in the feature fusion module of the interactive attention mechanism, the generation process of the query vector Query, the key vector Key and the value vector Value includes: Use 3×3 convolution kernels to perform convolution operations on the intrinsic attribute features and illumination features respectively to generate query vectors, key vectors, and value vectors; Calculate the similarity between the query vector and the key vector through the dot product to generate the energy matrix; The energy matrix is normalized through the Softmax function to generate the attention weight matrix; Multiply the attention weight matrix by the value vector to generate a fused feature map; The fused feature map is added to the original feature map through residual connection, where the weight γ of the residual connection is a learnable parameter with an initial value of 1.0; The feature map after residual connection is weighted fused through two parallel multi-layer perceptrons (MLP) to generate the final classification features.
[0019] The technical effect of the above improvements is as follows: through the feature fusion module of the interactive attention mechanism, the weights of the inherent attribute features and illumination features are dynamically adjusted to achieve information complementarity, which significantly improves the accuracy and robustness of hyperspectral image classification.
[0020] As a further improvement of the above scheme, the Retinex decomposition module uses a multi-scale Gaussian kernel for filtering, and the scale parameter σi of the Gaussian kernel is dynamically adjusted according to the local information and global information of the hyperspectral image to fully consider the local information and global information of the hyperspectral image.
[0021] The technical effect of the above improvements is as follows: through multi-scale Gaussian kernel filtering, the local and global information of the hyperspectral image are fully considered, the robustness of Retinex decomposition is enhanced, and the classification performance is improved.
[0022] As a further improvement of the above scheme, the intrinsic attribute component feature extractor and the illumination component feature extractor both use convolutional neural networks for feature extraction, and accelerate the training process through batch normalization and ReLU activation function.
[0023] The technical effects of the above improvements are as follows: Through batch normalization and ReLU activation function, the training process is accelerated, the convergence speed and stability of the model are improved, and the effect of feature extraction is enhanced.
[0024] As a further improvement of the above scheme, the feature fusion module based on the interactive attention mechanism achieves information complementarity by dynamically adjusting the weights of inherent attribute features and illumination features, thereby improving classification accuracy.
[0025] The technical effect of the above improvements is as follows: by dynamically adjusting the weights, the effective fusion of inherent attribute features and illumination features is achieved, the classification accuracy is improved, and the generalization ability of the model is enhanced.
[0026] As a further improvement of the above scheme, the method generates a final classification graph through a visualization network, which is convenient for users to intuitively understand the classification results.
[0027] The technical effect of the above improvements is: by visualizing the classification diagram, the interpretability and practicality of the method are enhanced, making it easier for users to intuitively understand the classification results.
[0028] Compared with the prior art, the overall technical effects produced by the present invention are: 1. Robustness to illumination changes: This paper separates the illumination component and the intrinsic attribute component of the hyperspectral image through Retinex decomposition, effectively reducing the impact of illumination changes on classification performance. By extracting illumination features and intrinsic attribute features respectively through a dual-branch network, the model can maintain high classification accuracy under different illumination conditions.
[0029] 2. Alleviation of overfitting problem: This invention introduces the interactive attention mechanism to make full use of the information in the limited labeled data and enhance the generalization ability of the model. By dynamically adjusting the weights of the inherent attribute features and illumination features, the model can better capture the intrinsic relationship of the data and reduce the risk of overfitting.
[0030] 3. Advantages of multi-feature fusion: This invention realizes the complementary fusion of illumination features and inherent attribute features through a dual-branch feature extractor and an interactive attention mechanism. This multi-feature fusion method significantly improves the classification performance of the model, especially in complex scenes, where the classification accuracy is significantly improved.
[0031] 4. Improvement of computational efficiency: This invention simplifies the processing of hyperspectral data and reduces computational complexity through Retinex decomposition and multi-scale filtering. At the same time, through the optimized design of convolutional neural network and attention mechanism, the computational efficiency is further improved, making this method more feasible in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is the overall framework diagram of the hyperspectral image classification method proposed in this invention.
[0033] Figure 2 This is a PU pseudo-color image.
[0034] Figure 3 This is the PU classification effect diagram.
[0035] Figure 4 is the PU ground truth label map.
[0036] Figure 5 Pseudo-color image of IP.
[0037] Figure 6 This is the IP classification effect diagram.
[0038] Figure 7 is the IP ground truth label map. DETAILED DESCRIPTION
[0039] In order to enable those skilled in the art to better understand the technical solution, the present invention is described in detail below in conjunction with embodiments. The description in this section is only exemplary and explanatory and should not have any limiting effect on the scope of protection of the present invention.
[0040] This paper applies the Retinex method to the field of hyperspectral image classification for the first time. In the traditional Retinex method, various information in the input RGB image is filtered to obtain a component reflecting the illumination change. Then, the illumination component is subtracted from the original image in the logarithmic domain to obtain a component representing the reflective property of the object. Hyperspectral images have the characteristics of a large number of bands and large information redundancy. One hyperspectral image is equivalent to hundreds of grayscale images. Therefore, the Retinex method is used for hyperspectral image processing. The weighted linear Retinex decomposition results of different scales are combined together to fully consider the local and global information of the important bands of the hyperspectral image. Then, the intrinsic attribute feature extractor and the illumination feature extractor are designed. Through deep feature embedding, the convolutional neural network is used to focus on the physical properties of the image and the illumination changes of the environment; in order to further ensure the separation of the intrinsic attribute component and the illumination component, a feature fusion module based on the intrinsic attribute and illumination interaction attention mechanism is introduced. The intrinsic attribute illumination interaction attention mechanism can complement the information of the two components, so that each component can dynamically obtain useful information from another component. The attention matrix is generated by calculating the similarity between the two component features, and the final feature fusion is adjusted accordingly to achieve feature complementarity and enhancement.
[0041] 1. Overall process: The present invention provides a hyperspectral image classification method based on the Retinex model, and the overall process thereof includes the following steps: 1) Retinex decomposition: Retinex decomposition is performed on the hyperspectral image to separate the intrinsic attribute component and the illumination component.
[0042] 2) Dual-branch feature extraction: The features of the intrinsic attribute component and the illumination component are extracted separately through a dual-branch network.
[0043] 3) Feature fusion: The inherent attribute features and illumination features are fused using a feature fusion module based on the interactive attention mechanism.
[0044] 4) Classification: Classify the fused features and generate the classification results of the hyperspectral image.
[0045] 2. Specific implementation of Retinex decomposition: In the Retinex decomposition step, the specific implementation is as follows: 1). Logarithmic transformation: Perform a logarithmic transformation on each band of the hyperspectral image to convert the image from linear space to logarithmic space. The formula for logarithmic transformation is:
[0046] in, : Input hyperspectral image, representing the position Place The pixel value of each band, is the reflectance image (intrinsic attribute component), is the light component, : Natural logarithmic transformation, converting the image from linear space to logarithmic space.
[0047] 2) Multi-scale Gaussian kernel filtering: Use Gaussian kernels of different scales (the scale parameter σi is dynamically adjusted according to the local and global information of the hyperspectral image) to filter the logarithmic transformed image to obtain the illumination component. The formula of the Gaussian kernel is:
[0048] : The scale parameter is Gaussian kernel is used to filter the image.
[0049] : The scale parameter of the Gaussian kernel indicates the scale of the filter, which is dynamically adjusted according to the local and global information of the image.
[0050] : Convolution operation, which means convolving the Gaussian kernel with the logarithmic transformation result of the image.
[0051] : In scale The light component obtained below.
[0052] 3). Calculation of reflectance image: Subtract the illumination component from the logarithmic transformed image to obtain the reflectance image (intrinsic attribute component). The formula is:
[0053] : In scale The reflectance image (intrinsic attribute component) obtained under
[0054] : Logarithmic transformation result of the input image.
[0055] : In scale The light component obtained below.
[0056] 4) Weighted average: Perform weighted average on reflectance images of different scales to obtain the final intrinsic attribute component. The formula is:
[0057] : Reflectance image after weighted average.
[0058] : The number of scales to use.
[0059] : In scale The reflectivity image obtained below.
[0060] 5). Exponential transformation: Perform exponential transformation on the intrinsic attribute component and the illumination component to restore them to linear space. The formula is:
[0061]
[0062] is the scale parameter of the multi-scale Gaussian kernel, which is used to capture local and global information of different scales of the image.
[0063] Reflectivity image and lighting component At different scales Calculated below.
[0064] : Reflectance image (intrinsic property component) after restoring the linear space.
[0065] : Lighting component after restoring linear space.
[0066] : Exponential transformation, restoring the result of logarithmic space to linear space.
[0067] Implementation of the intrinsic attribute component feature extractor: The convolutional layer is used to reduce the dimension of the input intrinsic attribute components and generate the initial feature map; Local features are gradually extracted through multiple convolutional layers, where the convolutional layers include convolution kernels of different sizes; Capture global and local salient information through global pooling operations, and generate global average pooling vectors and global maximum pooling vectors; After concatenating the global average pooling vector and the global maximum pooling vector, the channel attention weight is generated through a fully connected layer and an activation function; The channel attention weights are multiplied element-wise with the original feature map to generate the enhanced feature map.
[0068] Specific: In the two-branch network, the specific implementation of the intrinsic attribute component feature extractor is as follows: 1) 1×1×1 convolutional layer: Reduce the number of channels of the input band to 30 and generate the initial feature map:
[0069] 2) Multi-layer convolution layer: Local features are gradually extracted through 7×3×3, 5×3×3 and 3×3×3 convolution kernels to obtain the feature map after extracting local features:
[0070] 3). Reshaping operation: Reorganize the multi-channel feature map into a two-dimensional feature map for the next convolution operation.
[0071] 4). 3×3×3 convolutional layer: Reduce the number of channels from 576 to 64 and generate feature maps:
[0072] 5). Batch Normalization and ReLU Activation Function: Batch normalization layer and ReLU activation function are applied after each convolutional layer to speed up the training process.
[0073] 6) Dropout layer: Add the Dropout layer to reduce the model’s dependence on training data, improve the model’s generalization ability, and obtain the final feature map:
[0074] 7). Global pooling operation: Global Average Pooling:
[0075] : Feature vector after global average pooling.
[0076] : The height of the feature map.
[0077] : The width of the feature map.
[0078] : channel index of feature map.
[0079] Global max pooling:
[0080] : Feature vector after global maximum pooling.
[0081] : In the spatial dimension Take the maximum value.
[0082] 8) Channel Attention Mechanism: Concatenate the global average pooling and global maximum pooling vectors:
[0083] Generate intermediate features through the fully connected layer and ReLU activation function:
[0084] Generate channel attention weights through the Sigmoid activation function:
[0085] : Channel attention weight.
[0086] : Sigmoid activation function, which limits the output to [0,1].
[0087] : Weight parameters of the fully connected layer.
[0088] : Bias parameter of the fully connected layer.
[0089] : intermediate eigenvector.
[0090] 9) Feature enhancement: Multiply the channel attention weights by the original feature map element by element to generate an enhanced feature map:
[0091] 10) Cross Entropy Loss Function: During the training process, the feature extractor is trained using the cross entropy loss function:
[0092] 4. Implementation of illumination component feature extractor: The illumination component feature extractor uses the same deep learning architecture as the intrinsic attribute component feature extractor, and is implemented as follows: 1). Same architecture: Use the same convolutional neural network architecture as the intrinsic attribute component feature extractor.
[0093] 2). Cross entropy loss function: Training is performed through the cross entropy loss function to ensure the accuracy of illumination feature extraction. The formula is:
[0094] 5. Implementation of feature fusion module based on interactive attention mechanism: In the feature fusion step, the specific implementation of the feature fusion module based on the interactive attention mechanism is as follows: 1) Feature conversion: The intrinsic attribute features and illumination features are converted into query vector (Query), key vector (Key) and value vector (Value) through 3×3 convolution layers respectively. The formula is:
[0095]
[0096]
[0097] : query vector.
[0098] : key vector.
[0099] : value vector.
[0100] : weight parameter of the convolution kernel.
[0101] 2) Energy matrix generation: The energy matrix is generated by calculating the similarity between the query vector and the key vector through the dot product. The formula is:
[0102] : Energy matrix, representing the similarity between the query vector and the key vector.
[0103] 3) Attention weight matrix generation: Normalize the energy matrix through the Softmax function to generate the attention weight matrix. The formula is:
[0104] : Attention weight matrix, representing the The feature of the position The attention weight of the feature at each position.
[0105] 4) Feature fusion: Multiply the attention weight matrix and the value vector to generate a fused feature map. The formula is:
[0106] : The fused feature map.
[0107] : The transpose of the attention weight matrix.
[0108] 5) Residual connection: The fused feature map is added to the original feature map through the residual connection, where the weight γ of the residual connection is a learnable parameter. The formula is:
[0109] 6) Weighted fusion: The feature map after residual connection is weighted fused through two parallel multi-layer perceptrons (MLP) to generate the final classification features. The formula is:
[0110] 6. Generation process of feature fusion module of interactive attention mechanism: In the feature fusion module of the interactive attention mechanism, the generation process of the query vector, key vector, and value vector is as follows: 1) Convolution operation: Use a 3×3 convolution kernel to perform convolution operations on the intrinsic attribute features and illumination features respectively to generate query vectors, key vectors, and value vectors.
[0111] 2). Energy matrix generation: The energy matrix is generated by calculating the similarity between the query vector and the key vector through the dot product.
[0112] 3) Attention weight matrix generation: Normalize the energy matrix through the Softmax function to generate the attention weight matrix.
[0113] 4). Feature fusion: Multiply the attention weight matrix with the value vector to generate a fused feature map.
[0114] 5). Residual connection: The fused feature map is added to the original feature map through residual connection, where the weight γ of the residual connection is a learnable parameter with an initial value of 1.0.
[0115] 6) Weighted fusion: The feature map after residual connection is weighted fused through two parallel multi-layer perceptrons (MLP) to generate the final classification features.
[0116] 7). Multi-scale Gaussian kernel filtering of Retinex decomposition module: In the Retinex decomposition module, the specific implementation of multi-scale Gaussian kernel filtering is as follows: 1). Multi-scale Gaussian kernel: Use Gaussian kernels of different scales (the scale parameter σi is dynamically adjusted according to the local and global information of the hyperspectral image) for filtering.
[0117] 2) Local and global information: Through multi-scale Gaussian kernel filtering, the local and global information of the hyperspectral image is fully considered to enhance the robustness of Retinex decomposition.
[0118] 8. Convolutional Neural Network Architecture for Feature Extractor: Both the intrinsic attribute component feature extractor and the illumination component feature extractor use convolutional neural networks for feature extraction, and the specific implementation is as follows: 1). Convolutional neural network: extract local features through multiple convolutional layers.
[0119] 2). Batch normalization and ReLU activation function: Batch normalization and ReLU activation function are used to accelerate the training process and improve the convergence speed and stability of the model.
[0120] 9. Dynamic weight adjustment of the feature fusion module of the interactive attention mechanism: In the feature fusion module based on the interactive attention mechanism, the specific implementation of dynamic weight adjustment is as follows: 1) Dynamic weight adjustment: Dynamically adjust the weights of inherent attribute features and illumination features through the interactive attention mechanism.
[0121] 2). Information complementarity: Realize information complementarity between inherent attribute features and illumination features to improve classification accuracy.
[0122] 10. Visualize the network to generate classification graphs: After the classification step is completed, the final classification map is generated through the visualization network, which is implemented as follows: 1). Visualization network: Generate a classification graph through the classification results through the visualization network.
[0123] 2). Intuitive understanding: It is easy for users to intuitively understand the classification results and enhance the interpretability and practicality of the method. Specific embodiment: The overall framework of the hyperspectral image classification method proposed in this invention is as follows: Figure 1 shown.
[0125] First, the hyperspectral image is decomposed by Retinex; Secondly, the inherent attribute component and illumination component obtained by decomposition are respectively subjected to feature extraction through a dual-branch network; Then, the interactive attention-based feature fusion module is used to perform feature fusion and classification; Finally, the final classification graph is generated by visualizing the network.
[0126] 1. Hyperspectral Retinex Decomposition Module Retinex theory aims to explain how the human visual system perceives color and brightness, and how it can remain relatively constant even when lighting conditions vary greatly. The theory holds that the human visual system relies not only on the intensity of local pixels but also on the contextual information around them when perceiving color and brightness. Mathematically, the Retinex model can be expressed as:
[0127] In the formula is the pixel value of the input image. is a reflectance image, representing the intrinsic properties of the object surface. is a lighting image, which represents the change of ambient lighting. The goal is to The reflectance image is estimated from , because the reflectance image can better reflect the inherent characteristics of the object, while the illumination image includes the influence of ambient lighting.
[0128] In hyperspectral image processing, Retinex theory can also effectively enhance the contrast and details of the image while reducing the impact of illumination changes, but it needs to consider the information of multiple bands. Assume that the hyperspectral image is , where B represents the band. The pixels of the hyperspectral image can be decomposed into reflectance illumination The product of:
[0129] We use Gaussian kernels of different sizes. Filter the log-transformed image for each band:
[0130] By subtracting the logarithmic transformed image from the Gaussian filtering results at each scale, the reflectance images (intrinsic attribute components) at different scales are obtained:
[0131] The reflectivity images at different scales are weighted averaged to obtain the final reflectivity image:
[0132] Finally, the final reflectance image and the previously filtered illumination change image are exponentially transformed to restore to linear space:
[0133]
[0134] Through the above method, the present invention successfully applies the Retinex theory to the field of hyperspectral images and realizes the separation of inherent attributes and illumination changes in hyperspectral images.
[0135] 2. Dual-branch feature extractor 1) Intrinsic attribute component feature extractor The role of the intrinsic attribute component feature extractor is to effectively capture and represent features related to the intrinsic attributes of hyperspectral images so as to accurately identify and distinguish different land cover types in subsequent classification tasks. By deeply mining key features such as material characteristics and texture information in the intrinsic attribute components, it improves the representation ability and classification accuracy of the model.
[0136] First, the input layer receives the original intrinsic attribute components , where H and W represent the height and width of the image, respectively, and B represents the number of bands. In order to reduce the number of channels, simplify subsequent calculations, and retain important feature information, a 1×1×1 convolution layer is used to reduce the number of input bands to 30 channels to generate the initial feature map:
[0137] Next, through a series of multi-layer convolution layers, more complex local features are gradually extracted. Through convolution kernels of sizes 7×3×3, 5×3×3, and 3×3×3, the feature map after extracting local features is obtained:
[0138] Then, the multi-channel feature map is reshaped Reorganize into 2D feature map In order to carry out the next convolution operation. Finally, use a 3×3×3 convolution kernel to reduce the number of channels from 576 to 64 and generate a feature map:
[0139] In order to stabilize the input distribution of each layer, speed up the training process, and reduce internal covariate shift, a batch normalization layer and ReLU activation function are applied after each convolution layer. In order to reduce the model's dependence on training data and improve the model's generalization ability, a Dropout layer is added to obtain the final feature map:
[0140] In order to capture the feature map The global information of the feature map Fr is averagely pooled in the spatial dimension through the global average pooling operation. Specifically, for each channel c, its average value over all spatial locations is calculated:
[0141] In order to capture the local salient information of the feature map Fr, the feature map is pooled by global maximum pooling operation. Each channel of is max-pooled in the spatial dimension, and we get Specifically, for each channel c, calculate its maximum value over all spatial locations:
[0142] The two vectors obtained by global average pooling and global maximum pooling , Spliced together to form a new feature vector :
[0143] Through the fully connected layer (FC) and ReLU activation function, the feature vector Mapping to intermediate features The role of this layer is to compress the feature vector into a lower dimensional space while introducing nonlinearity so that it can better capture the relationship between features:
[0144] Then, through another fully connected layer and Sigmoid activation function, the channel attention weight ωr is generated. The Sigmoid activation function limits the output to between [0,1], so that the weight of each channel is within a reasonable range:
[0145] Finally, the channel attention weight With the original feature map Element-by-element multiplication to obtain the enhanced feature map The important channels are amplified and the unimportant channels are suppressed, thereby enhancing the expressiveness of the feature map:
[0146] During the training process, the feature extractor is trained using the cross entropy loss function to minimize the difference between the predicted value and the true label. The cross entropy loss function is defined as:
[0147] Where the predicted value is p, the true label is y, and C represents the number of categories.
[0148] 2) Light component feature extractor The role of the illumination component feature extractor is to effectively capture and represent illumination-related features in hyperspectral images so that different types of objects can be accurately identified and distinguished in subsequent classification tasks. It improves the model's representational capabilities and classification accuracy by focusing deeply on key features such as environmental changes. It is able to focus on extracting and enhancing features related to illumination conditions when processing hyperspectral images. These features are critical for accurately identifying and distinguishing different types of objects, especially when illumination conditions vary greatly. It uses the same deep learning architecture as the intrinsic attribute component feature extractor for feature extraction and is also trained using cross entropy loss.
[0149] 3. Feature Fusion Module Based on Interactive Attention Mechanism The extracted reflectance features and illumination features each contain useful information. In addition, they are complementary to each other. In order to make full use of this complementary information, we introduced the interactive attention mechanism for feature fusion. Through the interactive attention mechanism, the model can dynamically adjust the weights between the two features, so as to better capture the complementary information between them and extract richer feature representations, which helps to improve the overall classification performance of the model.
[0150] First, the input feature maps Fr1 and Fr2 are converted into query vectors (Query), key vectors (Key), and value vectors (Value), respectively:
[0151]
[0152]
[0153] In the formula is the convolution kernel weight of the query projection, is the convolution kernel weight of the key projection, is the convolution kernel weight for value projection.
[0154] The energy matrix is used to represent the similarity between the features of each position in Fr1 and the features of each position in Fr2. The similarity can be calculated by dot product, which can quantify the correlation between different positions. The energy matrix E is obtained by performing dot product operation on the query vector Q and the key vector K:
[0155] Softmax is used to ensure that the sum of the attention weights at each position is 1, so that the attention weights have the properties of probability distribution and can distribute the weights more effectively.
[0156]
[0157] In the formula Represents the attention weight of the feature at the i-th position to the feature at the j-th position.
[0158] By applying the attention weight matrix A to the value vector V, a new feature map O can be obtained, which fuses the complementary information between X1 and X2.
[0159]
[0160] In order to maintain the stability of features and prevent gradient disappearance, we introduce a learnable scaling parameter γ and perform a residual connection between the new feature map O1 and the original feature map X1:
[0161] Finally, the feature maps that pass through the interactive attention mechanism are passed through two parallel MLPs and then weighted fused and classified.
[0162] Classification effect: 1. Dataset Introduction 1) Pavia University: This dataset was collected in 2001 at the University of Pavia, Italy, using the Reflection Optical System Imaging Spectrometer (ROSIS). The dataset consists of 103 effective bands with a size of 610 × 340. The spatial resolution of the dataset is 1.3 m / pixel, and the wavelength range is 1.3 m / pixel, with a wavelength range of 0.43 to 0.86 μm. The dataset contains a total of 42,776 labeled pixels.
[0163] 2) IndianPines: This dataset is an image of the Indian Pines region in Indiana, USA, taken by the Airborne Visible Infrared Imaging Spectrometer (AVIRIS) in 1992. The dataset consists of 200 effective bands, with a size of 145×145, 16 categories, and a wavelength range of 0.4 to 2.5μm. The dataset contains a total of 10,249 labeled pixels.
[0164] 2. Classification Results 1) Hyperspectral image dataset Pavia University: Figure 2 This is a PU pseudo-color image.
[0165] Figure 3 This is the PU classification effect diagram.
[0166] Figure 4 is the PU ground truth label map.
[0167] 2) Hyperspectral image dataset Indian_Pines (IP): Figure 5 Pseudo-color image of IP.
[0168] Figure 6 This is the IP classification effect diagram.
[0169] Figure 7 is the IP ground truth label map.
[0170] The present invention conducts experiments on two widely used hyperspectral image datasets, where the division ratio of the training set and the test set is 0.1:99.9, and the classification accuracies reach 95.2% and 92.8% respectively, which are significantly better than the existing methods, proving the effectiveness and superiority of the proposed method.
[0171] It should be noted that, in this article, the terms: include, contain and any other variants are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Specific examples are used in this article to illustrate the principle and implementation of the technical solution of the present invention. The above examples are only used to help understand the method of the present invention and its core idea. The above is only a preferred implementation of the present invention. It should be pointed out that due to the limitations of textual expression and the objective existence of infinite specific structures, for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements, modifications or changes can be made, and the above technical features can also be combined in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the concept and technical solution of the present invention to other occasions without improvement, should be regarded as the protection scope of the present invention.
Claims
1. A hyperspectral image classification method based on Retinex model, characterized in that: The following steps are involved: Perform Retinex decomposition on the hyperspectral image to obtain the intrinsic attribute component and the illumination component; The dual-branch network is used to extract features of the intrinsic attribute component and the illumination component respectively, wherein the intrinsic attribute component feature extractor is used to extract the material features and texture information of the hyperspectral image, and the illumination component feature extractor is used to extract the illumination features of the hyperspectral image; The feature fusion module based on the interactive attention mechanism is used to fuse the extracted intrinsic attribute features and illumination features to generate a fused feature representation. The fused features are classified to generate the classification results of the hyperspectral image.
2. The hyperspectral image classification method based on the Retinex model according to claim 1 is characterized in that: The Retinex decomposition step comprises: Perform logarithmic transformation on each band of the hyperspectral image; The illumination component is obtained by filtering using Gaussian kernels of different scales, where the scale parameter σi of the Gaussian kernel is dynamically adjusted according to the local and global information of the hyperspectral image; Subtracting the illumination component from the logarithmically transformed image to obtain a reflectance image, wherein the reflectance image represents an intrinsic attribute component of the hyperspectral image; Perform weighted averaging on reflectance images of different scales to obtain the final intrinsic attribute component; Perform exponential transformation on the intrinsic attribute component and the illumination component and restore them to the linear space.
3. The hyperspectral image classification method based on the Retinex model according to claim 1 is characterized in that: The intrinsic attribute component feature extractor comprises: The convolutional layer is used to reduce the dimension of the input intrinsic attribute components and generate the initial feature map; Local features are gradually extracted through multiple convolutional layers, where the convolutional layers include convolution kernels of different sizes; Capture global and local salient information through global pooling operations, and generate global average pooling vectors and global maximum pooling vectors; After concatenating the global average pooling vector and the global maximum pooling vector, the channel attention weight is generated through a fully connected layer and an activation function; The channel attention weights are multiplied element-wise with the original feature map to generate the enhanced feature map.
4. The hyperspectral image classification method based on the Retinex model according to claim 1 is characterized in that: The illumination component feature extractor uses the same deep learning architecture as the intrinsic attribute component feature extractor for feature extraction and is trained using a cross entropy loss function.
5. The hyperspectral image classification method based on the Retinex model according to claim 1 is characterized in that: The feature fusion module based on the interactive attention mechanism includes: The intrinsic attribute features and illumination features are converted into query vector Query, key vector Key and value vector Value through convolution layers, where the convolution kernel size is 3×3; Calculate the similarity between the query vector and the key vector through the dot product to generate the energy matrix; The energy matrix is normalized through the Softmax function to generate the attention weight matrix; Multiply the attention weight matrix by the value vector to generate a fused feature map; The fused feature map is added to the original feature map through residual connection, where the weight γ of the residual connection is a learnable parameter; The feature map after residual connection is weighted fused through two parallel multi-layer perceptrons (MLPs) to generate the final classification features.
6. The hyperspectral image classification method based on the Retinex model according to claim 5 is characterized in that: In the feature fusion module of the interactive attention mechanism, the generation process of the query vector Query, the key vector Key and the value vector Value includes: Use 3×3 convolution kernels to perform convolution operations on the intrinsic attribute features and illumination features respectively to generate query vectors, key vectors, and value vectors; Calculate the similarity between the query vector and the key vector through the dot product to generate the energy matrix; The energy matrix is normalized through the Softmax function to generate the attention weight matrix; Multiply the attention weight matrix by the value vector to generate a fused feature map; The fused feature map is added to the original feature map through residual connection, where the weight γ of the residual connection is a learnable parameter with an initial value of 1.0; The feature map after residual connection is weighted fused through two parallel multi-layer perceptrons (MLP) to generate the final classification features.
7. The hyperspectral image classification method based on the Retinex model according to claim 1, characterized in that: The Retinex decomposition module uses a multi-scale Gaussian kernel for filtering, and the scale parameter σi of the Gaussian kernel is dynamically adjusted according to the local information and global information of the hyperspectral image to fully consider the local information and global information of the hyperspectral image.
8. The hyperspectral image classification method based on the Retinex model according to claim 1 is characterized in that: The intrinsic attribute component feature extractor and the illumination component feature extractor both use convolutional neural networks for feature extraction, and accelerate the training process through batch normalization and ReLU activation function.
9. The hyperspectral image classification method based on the Retinex model according to claim 1, characterized in that: The feature fusion module based on the interactive attention mechanism achieves information complementarity by dynamically adjusting the weights of inherent attribute features and illumination features, thereby improving classification accuracy.
10. The hyperspectral image classification method based on the Retinex model according to claim 1, characterized in that: The method generates a final classification graph through a visualization network, which is convenient for users to intuitively understand the classification results.
Citation Information
Patent Citations
Gesture image retrieval method based on multi-scale Retinex and improved VGGNet network
CN111695508A
Hyperspectral image classification method
CN116310471A
Low-illumination image enhancement method based on feature fusion and attention embedding
CN116797488A
System and method for radio-based sleep monitoring
EP4162865A1
Cited By
Method and system for extracting habitat area based on remote sensing image
CN121305374A
A habitat area extraction method and system based on remote sensing images
CN121305374B