Transform-based hyperspectral intrinsic decomposition device and method thereof
Through the Transformer-based hyperspectral eigendecomposition device, the reflectance spectral cross attention mechanism and the gated forward module are used to solve the problem of reflectance edge blurring of hyperspectral images, achieving more accurate reflectance edge decomposition, and improving the robustness and generalization of the algorithm.
Patent Information
- Application Number
- CN202411838934.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-13
Smart Images

Figure CN119992309A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a Transformer-based hyperspectral intrinsic decomposition device and method thereof. Background Art
[0002] Hyperspectral images record many clues about scene features. The information they contain is of great significance for scene perception and analysis, and they have natural advantages in many downstream computer vision tasks. With the popularization of spectral sensors in smartphones, hyperspectral data is expected to be widely promoted and applied in tasks such as material classification, face detection, light source estimation, and white balance. However, the reflectivity of objects in the scene is coupled with multiple physical quantities such as ambient lighting and geometry, resulting in hyperspectral data being sensitive to ambient geometry and light sources. The spectral information of two points in the same reflectivity area and the same point in different light sources show significant differences, which makes the robustness of directly applying hyperspectral image data very limited. Reflectivity can reflect material properties that are independent of the environment. Using scene reflectivity for downstream hyperspectral tasks can greatly improve the robustness and generalization of the algorithm. However, obtaining reflectivity through hyperspectral intrinsic decomposition is a highly complex and pathological problem with diverse degradation patterns. Therefore, further refining and extracting reflectivity that is independent of the environment has become a very valuable and challenging problem.
[0003] With the rapid development of deep learning technology, attempts have been made to use convolutional neural networks (CNNs) for intrinsic decomposition, and they have shown excellent performance. Although the convolution operation can efficiently extract image features due to its translation invariance, it is often constrained by the small receptive field and cannot fully mine global information, thus limiting the high-fidelity decomposition of hyperspectral images. In recent years, the application of Transformer models in computer vision tasks has achieved remarkable results. However, the proposed spatial attention mechanism will greatly increase the computational pressure. Therefore, in order to balance the information perception scale and computational overhead, the channel attention mechanism becomes an ideal alternative.
[0004] However, since non-Lambertian surfaces often exist in the acquisition scenes, compared with the Lambertian surface, the weak and widely distributed part of the additional specular reflection will blur the reflectivity edge, and the part with high intensity and local distribution will produce abnormal edges. The combination of the two makes it more difficult to accurately predict the reflectivity edge, further worsening the difficulty of reflectivity decomposition. Previous studies often did not pay attention to the contamination of the edge by specular reflection, resulting in unsatisfactory decomposed edge effects. Therefore, in order to further optimize the effect of hyperspectral intrinsic decomposition, it is necessary to design an algorithm to solve the impact of specular reflection on the reflectivity edge. Summary of the invention
[0005] In order to deal with the problem of blurred reflectivity edges caused by specular reflection and to better perform eigendecomposition on non-Lambertian surfaces, the present invention provides a Transformer-based hyperspectral eigendecomposition device and method.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A Transformer-based hyperspectral intrinsic decomposition device, the device comprising a CSR coding alignment module, an encoder, a first Transformer module and a decoder;
[0008] The CSR encoding alignment module is used to encode the cross-spectral gradient ratio CSR implicit feature and match the feature dimensions of the encoder and the decoder to achieve feature alignment, wherein the cross-spectral gradient ratio CSR implicit feature is derived from the original collected hyperspectral image and is only related to the reflectivity;
[0009] The encoder is used to perform feature encoding on the original acquired hyperspectral image and gradually generate hierarchical features with low spatial size and high number of channels;
[0010] The first Transformer module is used to further encode the features output by the encoder;
[0011] The decoder is used to decode the deep implicit features output by the first Transformer module, and gradually restore the features to a spatial size and channel size consistent with the corresponding level of the encoder, and finally decompose the reflectance image.
[0012] Furthermore, the CSR coding alignment module includes an embedding layer and two downsampling modules connected in sequence; the encoder includes an embedding layer and three groups of second Transformer modules and downsampling modules stacked in sequence; the decoder includes three groups of upsampling modules stacked in sequence and a third Transformer module and a mapping layer.
[0013] Furthermore, the first Transformer module includes a first normalization layer, a reflectance spectrum cross-attention module, a second normalization layer and a gated forward module connected in sequence; the reflectance spectrum cross-attention module is used to integrate the aligned cross-spectral gradient ratio CSR implicit features as the reflectance edge attention map into the channel dimension self-attention mechanism through Hadamard multiplication, so as to guide the attention to the reflectance edge decomposition; the gated forward module is used to divide the input features into two paths for linear projection to expand the feature channel, and then nonlinearly activate one of the features through an activation function and perform Hadamard multiplication with the other path to realize the feature gating mechanism.
[0014] The present invention also provides a decomposition method of a hyperspectral intrinsic decomposition device based on Transformer, the method comprising the following steps:
[0015] (1) Obtaining and preprocessing a hyperspectral image pair to obtain a data set including an original acquired image and a corresponding reflectance image;
[0016] (2) The original collected image is input into the encoder and the CSR coding alignment module respectively; the original collected image passes through the encoder, the first Transformer module and the decoder in sequence to decompose the reflectance of the original image size; after the original collected image passes through the CSR coding alignment module, the aligned cross spectral gradient ratio CSR implicit feature is obtained;
[0017] (3) The difference between the reflectivity decomposed in step (2) and the ground truth value of the reflectivity is obtained through the error function, and based on this, the CSR coding alignment module, the encoder, the first Transformer module and the decoder parameters are cyclically updated to reduce the error until convergence is finally achieved, thereby completing the training process of the decomposition device;
[0018] (4) The originally acquired hyperspectral image is input into the trained decomposition device to obtain the decomposed reflectance.
[0019] The above technical solution of the present invention has the following advantages compared with the prior art:
[0020] The present invention provides a Transformer-based hyperspectral intrinsic decomposition device and method thereof. The original collected hyperspectral image is input into the trained intrinsic decomposition model. The CSRGA module derives the prior data CSR based on the input image, and drives the network to focus on the reflectivity edge through Hadamard multiplication and multi-head self-attention block; the GDFN module further extracts deep features by convolution and groups features, so that one part is activated and then Hadamard multiplied with another part to achieve gating, so as to achieve feature screening, retention and forgetting; then, encoding and decoding are performed in combination with multi-scale information to decompose the reflectivity image. The present invention overcomes the shortcomings of the prior art that the contamination mode of the reflectivity edge caused by the mirror reflection is not paid attention to, resulting in difficulty in edge decomposition, and better realizes the intrinsic decomposition of the reflectivity of the hyperspectral image. The reflectivity obtained by the present invention can be used as prior knowledge to improve the task performance of mobile phone imaging such as local dark light enhancement, high dynamic range imaging and re-illumination. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 Schematic diagram of the structure of a hyperspectral intrinsic decomposition device based on Transformer in an embodiment of the present invention;
[0022] In the figure: 101-originally collected hyperspectral image, 102 cross spectral gradient ratio CSR (Cross Spectral-gradient Ratio), 103-CSR encoding alignment module, 104-encoder, 105-Transformer module, 106-decoder, 107-embedding layer, 108-cross spectral gradient ratio guided attention module CSRGA (CSR-GuidedAttention), 109-normalization layer, 110 gated forward module GDFN (Gated-Dconv Feed-Forward), 111-downsampling module, 112-upsampling module, 113-mapping layer, 114-reflectance image.
[0023] Figure 2 Schematic diagram of the structure of the cross-spectral gradient ratio guided attention module CSRGA in an embodiment of the present invention;
[0024] In the figure: 201-matrix dimension change, 202-CSR encoding features, 203-Hadamard multiplication, 204-matrix multiplication, 205-Sigmoid activation function, 206-matrix addition, 207-position encoding.
[0025] Figure 3 Schematic diagram of the structure of the gated forward module GDFN in an embodiment of the present invention;
[0026] In the figure: 301-3×3 convolution, 302-1×1 convolution.
[0027] Figure 4 Flow chart of the Transformer-based hyperspectral eigendecomposition method in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The present invention is further described below in conjunction with the accompanying drawings.
[0029] A hyperspectral intrinsic decomposition device based on Transformer in this embodiment, such as Figure 1As shown, it specifically includes a CSR coding alignment module 103, an encoder 104, a first Transformer module 105, and a decoder 106. The CSR coding alignment module 103 includes an embedding layer 107 and two sequentially connected downsampling modules 111. The CSR 102 is derived from the original acquired hyperspectral image 101 and is only related to the reflectivity. The CSR coding alignment module 103 encodes the CSR implicit features and matches the feature dimensions of the corresponding CSRGA modules 108 in the encoder 104 and the decoder 106 to achieve feature alignment. The encoder 104 includes an embedding layer 107 and three groups of sequentially stacked second Transformer modules 105 and downsampling modules 111. The decoder 106 includes three groups of sequentially stacked upsampling modules 112 and a third Transformer module 105 and a mapping layer 113. The hyperspectral image 101 is feature encoded by the embedding layer 107 through the encoder 104 and gradually generates hierarchical features with low spatial size and high number of channels. It is then input into the first Transformer module 105 to further encode implicit features, and then input into the decoder 106 to decode the deep implicit features, gradually restoring the features to the spatial size and channel size consistent with the corresponding level of the encoder, and finally decomposing the reflectance image 114 through the mapping layer 113.
[0030] The Transformer module 105 includes a normalization layer 109, a CSRGA module 108, a normalization layer 109 and a GDFN module 110 which are connected in sequence. The structure of the CSRGA module 108 is as follows: Figure 2 As shown, the aligned CSR features input by the CSR encoding alignment module 103 are used as the reflectivity edge attention map and the channel dimension self-attention mechanism is integrated through Hadamard multiplication to guide the network to focus on the reflectivity edge decomposition. The structure of the GDFN module 110 is shown in Figure 3 As shown, the module divides the input features into two paths for linear projection to expand the feature channels, and then activates one of the features through Sigmoid nonlinearity and performs Hadamard multiplication with the other path to implement the feature gating mechanism. By screening and controlling the forward information flow, the GDFN module 110 allows different levels of the codec to focus on detailed features that complement other levels and obtain high-quality reflectivity features.
[0031] The reflectivity in the original collected hyperspectral image 101 is modulated by the ambient light source and the scene geometry, and the specular reflection on the non-Lambertian surface further blurs the reflectivity edge. After the original collected hyperspectral image 101 passes through the hyperspectral intrinsic decomposition device of this embodiment, the reflectivity image 114 can be decomposed. The specific implementation process is as follows:
[0032] The original acquired hyperspectral image 101 is input into the CSR encoding alignment module 103 to obtain the CSR 102 related only to the reflectance and further obtain the multi-level aligned implicit features. The original acquired hyperspectral image 101 is also input into the encoder 104, and is subjected to feature encoding and channel expansion by the embedding layer 107. The multi-level and multi-scale features are further extracted by the multi-level downsampling module 111 and further encoded by the first Transformer module 105. Then, it is input into the decoder 106, and the reflectance implicit features are decomposed by the symmetrical upsampling module 112 and the spatial scale is gradually enlarged. Finally, the reflectance image 111 of the original image size is output by the mapping layer 113. Among them, the multiple Transformer modules 105 composed of the CSRGA module 108, the normalization layer 109 and the GDFN module 110 are respectively combined with each level of the downsampling module 111 and the upsampling module 112 to form a codec. The normalization layer 109 usually refers to batch normalization, and can also be layer normalization.
[0033] This embodiment also provides a decomposition method based on the above-mentioned hyperspectral intrinsic decomposition device, comprising the following steps:
[0034] (1) Obtain image pairs taken by a hyperspectral camera, preprocess the image pairs as a dataset, and obtain a dataset of original acquired images and corresponding reflectance ground truth values for model training.
[0035] (2) The original collected image is input into the encoder 104 and the CSR coding alignment module 103 respectively. The original collected image passes through the encoder 104, the first Transformer module 105 and the decoder 106 in sequence to decompose the reflectance of the original image size; after the original collected image passes through the CSR coding alignment module 103, the aligned cross spectral gradient ratio CSR implicit feature is obtained.
[0036] (3) The difference between the reflectivity decomposed in step (2) and the ground truth value of the reflectivity is obtained through the error function, and based on this, the parameters of the CSR coding alignment module 103, the encoder 104, the first Transformer module 105 and the decoder 106 are cyclically updated to reduce the error until convergence is finally achieved, thereby completing the model training.
[0037] (4) The original captured image is input into the model trained in step (3) to obtain the decomposed reflectance, and then the decomposition quality is quantitatively evaluated by the standard root mean square error (MSE) and structural similarity (SSIM) in combination with the corresponding ground truth value of the reflectance map.
[0038] In one embodiment of the present invention, the CSRGA module 108 uses the CSR coding features aligned with the CSR coding alignment module 103 as the reflectivity edge attention map, and then uses Hadamard multiplication to fuse it with the channel self-attention mechanism to guide the multi-head channel attention of the hierarchical features to focus on the reflectivity edge features.
[0039] In one embodiment of the present invention, the CSR is:
[0040]
[0041] λ∈1,…,C-2
[0042] i∈1,…,H-2
[0043] j∈1,…,W-2
[0044] Where C is the number of spectral channels of the hyperspectral image, H and W reflect the height and width of the spatial dimension of the hyperspectral image, respectively, and I λ,i,j represents the pixel value of the i-th row and j-th column of the λ-th channel of the reflectance ground truth image, CSR λ,i,j Represents the pixel value of the i-th row and j-th column of the λ-th channel of the derived prior guidance image.
[0045] In one embodiment of the present invention, in step (2), the GDFN module 110 first performs feature dimensioning through 3×3 convolution; then further performs feature mining through 1×1 convolution and groups the obtained implicit features, and performs a gating operation on half of the features through a Sigmoid activation function and the other half through Hadamard multiplication to achieve feature retention or forgetting, and finally uses a 3×3 convolution to restore the features to the input dimension size.
[0046] In one embodiment of the present invention, in step (3), the error function is an L2 function:
[0047]
[0048] Where C is the number of spectral channels of the hyperspectral image, H and W reflect the height and width of the spatial dimension of the hyperspectral image, respectively, and I λ,i,j represents the pixel value of the i-th row and j-th column of the λ-th channel of the reflectance ground truth image, Represents the pixel value of the i-th row and j-th column of the λ-th channel of the image decomposed by the model.
[0049] In one embodiment of the present invention, in step (4), the standard root mean square error MSE calculation formula is:
[0050]
[0051] Among them I i,jrepresents the pixel value of the i-th row and j-th column of the λ-th channel of the reflectance ground truth image, Represents the pixel value of the i-th row and j-th column of the λ-th channel of the image decomposed by the model.
[0052] In one embodiment of the present invention, in step (4), the structural similarity SSIM calculation formula is:
[0053]
[0054] Where I represents the reflectance ground truth image, represents the image decomposed by the model, and They represent the contrast differences of brightness, contrast, and structure respectively; the structural similarity evaluation adjusts the corresponding weights of the three parts through α, β, and γ, which are usually set to α=β=γ=1.
[0055] In one embodiment of the present invention, the for:
[0056]
[0057] d1=(K1L) 2
[0058] where μ I and Represent the ground truth reflectance and the average brightness of the model decomposition image. d1 is a small constant to prevent possible This causes the formula to be unstable. K1 is generally taken as 0.01. L is the dynamic range of grayscale, which is related to the type of image data.
[0059] In one embodiment of the present invention, the for:
[0060]
[0061] d2=(K2L) 2
[0062] where σ I and are the mean square error of the ground truth reflectivity and the model decomposition image. d2 is a small constant to prevent possible This causes the formula to be unstable, so K2 is generally taken as 0.03.
[0063] In one embodiment of the present invention, the for:
[0064]
[0065] d3=(K3L) 2
[0066] where σ I and Describe the mean square error between the ground truth reflectivity and the model decomposition image, Represents the covariance of two hyperspectral images. d3 is a small constant to prevent possible This causes the formula to be unstable, so K3 is generally taken as 0.015.
[0067] The eigendecomposition simulation experiment was conducted on the same dataset as other existing methods. The test experiment included ten scenes under five light sources. The experimental results are shown in Table 1.
[0068] Table 1 Comparison of experimental performance of each method
[0069] method Standardized root mean square error MSE Structural Similarity SSIM Krebs 0.0442 0.00556 U-Net 0.01948 0.71368 MST++ 0.00856 0.79257 The present invention 0.00556 0.81721
[0070] It can be found from Table 1 that the present invention has achieved satisfactory results in both standard root mean square error and structural similarity. Therefore, the present invention uses the photometric invariant prior to guide the network to focus on the transition area of different reflectivity, overcomes the semantic ambiguity caused by specular reflection, and has a better decomposition effect for non-Lambertian objects.
Claims
1. A hyperspectral eigendecomposition device based on Transformer, characterized in that: The device includes a CSR coding alignment module, an encoder, a first Transformer module and a decoder; The CSR encoding alignment module is used to encode the cross-spectral gradient ratio CSR implicit feature and match the feature dimensions of the encoder and the decoder to achieve feature alignment, wherein the cross-spectral gradient ratio CSR implicit feature is derived from the original collected hyperspectral image and is only related to the reflectivity; The encoder is used to perform feature encoding on the original acquired hyperspectral image and gradually generate hierarchical features with low spatial size and high number of channels; The first Transformer module is used to further encode the features output by the encoder; The decoder is used to decode the deep implicit features output by the first Transformer module, and gradually restore the features to a spatial size and channel size consistent with the corresponding level of the encoder, and finally decompose the reflectance image.
2. The hyperspectral intrinsic decomposition device based on Transformer according to claim 1, characterized in that: The CSR coding alignment module includes an embedding layer and two downsampling modules connected in sequence; the encoder includes an embedding layer and three groups of second Transformer modules and downsampling modules stacked in sequence; the decoder includes three groups of upsampling modules stacked in sequence and a third Transformer module and a mapping layer.
3. The hyperspectral intrinsic decomposition device based on Transformer according to claim 1, characterized in that: The first Transformer module includes a first normalization layer, a reflectance spectrum cross-attention module, a second normalization layer and a gated forward module connected in sequence; the reflectance spectrum cross-attention module is used to integrate the aligned cross-spectral gradient ratio CSR implicit feature as a reflectance edge attention map into the channel dimension self-attention mechanism through Hadamard multiplication, and guide the attention to focus on the reflectance edge decomposition; the gated forward module is used to divide the input features into two paths for linear projection to expand the feature channel, and then nonlinearly activate one of the features through an activation function and perform Hadamard multiplication with the other path to realize the feature gating mechanism.
4. The decomposition method using the Transformer-based hyperspectral eigendecomposition device as claimed in claim 1, characterized in that: The method comprises the following steps: (1) Obtaining and preprocessing a hyperspectral image pair to obtain a data set including an original acquired image and a corresponding reflectance image; (2) Input the original captured image into the encoder and CSR encoding alignment module respectively; The original collected image passes through the encoder, the first Transformer module and the decoder in sequence to decompose the reflectance of the original image size; after the original collected image passes through the CSR encoding alignment module, the aligned cross-spectral gradient ratio CSR implicit feature is obtained; (3) The difference between the reflectivity decomposed in step (2) and the ground truth value of the reflectivity is obtained through the error function, and based on this, the CSR coding alignment module, the encoder, the first Transformer module and the decoder parameters are cyclically updated to reduce the error until convergence is finally achieved, thereby completing the training process of the decomposition device; (4) The originally acquired hyperspectral image is input into the trained decomposition device to obtain the decomposed reflectance.
5. The decomposition method according to claim 4, characterized in that: In step (2), the cross spectral gradient ratio CSR is: λ∈1,…,C-2 i∈1,…,H-2 j∈1,…,W-2 Where C is the number of spectral channels of the hyperspectral image, H and W reflect the height and width of the spatial dimension of the hyperspectral image, respectively, and I λ,i,j represents the pixel value of the i-th row and j-th column of the λ-th channel of the reflectance ground truth image, CSR λ,i,j Represents the pixel value of the i-th row and j-th column of the λ-th channel of the derived prior guidance image.
6. The decomposition method according to claim 4, characterized in that: In step (3), the error function is an L2 function: Where C is the number of spectral channels of the hyperspectral image, H and W reflect the height and width of the spatial dimension of the hyperspectral image, respectively, and I λ,i,j represents the pixel value of the i-th row and j-th column of the λ-th channel of the reflectance ground truth image, Represents the pixel value of the i-th row and j-th column of the decomposed image in the λ-th channel.
7. The decomposition method according to claim 4, characterized in that: In step (4), after obtaining the decomposed reflectivity, the decomposition quality is quantitatively evaluated by the standard root mean square error MSE and the structural similarity SSIM, where the SSIM calculation formula is: Where I represents the reflectance ground truth image, represents the decomposed image, and They represent the contrast differences of brightness, contrast, and structure respectively; the structural similarity evaluation adjusts the corresponding weights of the three parts through α, β, and γ.
8. The decomposition method according to claim 7, characterized in that: In the SSIM calculation formula, Where C is the number of spectral channels of the hyperspectral image, H and W reflect the height and width of the spatial dimension of the hyperspectral image, respectively, and μ I and Represent the ground truth reflectance and the average brightness of the decomposed image, I λ,i,j represents the pixel value of the i-th row and j-th column of the λ-th channel of the reflectance ground truth image, and d1 is a constant.
9. The decomposition method according to claim 8, characterized in that: In the SSIM calculation formula, where σ I and are the mean square error between the ground truth reflectivity and the decomposed image, and d2 is a constant.
10. The decomposition method according to claim 9, characterized in that: In the SSIM calculation formula, where σ I and are the mean square error between the ground truth reflectivity and the decomposed image, represents the covariance of two hyperspectral images, It represents the pixel value of the i-th row and j-th column of the λ-th channel of the decomposed image, and d3 is a constant.