A spectral image fusion super-resolution method and system based on misaligned multi-source data

Through structural features guided multi-spectral and hyperspectral image alignment and fusion networks, the problem of poor alignment effect during super-segment of misaligned image fusion is solved, and efficient image fusion and super resolution are achieved, especially in high-magnification tasks.

CN118967450BActive Publication Date: 2025-05-06BEIJING INST OF TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411452810.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-05-06
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

The prior art has poor alignment effect when unaligned hyperspectral and multispectral images are fused with supersegment, and performance is significantly reduced in high-magnification supersegment tasks, and optical flow-based methods can lead to the loss of spectral information.

Method used

A structural feature-guided multi-spectral and hyperspectral image alignment and fusion network is proposed. Features are extracted through gradient calculation, texture encoder and structure encoder, and aligned and fusion using structural attention-guided feature fusion alignment module, and finally a high-resolution hyperspectral image is generated using the decoder network.

Benefits of technology

More robust and efficient image alignment and fusion are achieved, and super-resolution effects are improved, especially in high-magnification super-score tasks, and avoiding the loss of spectral information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967450B_ABST
    Figure CN118967450B_ABST
Patent Text Reader

Abstract

The present invention relates to a spectral image fusion super-resolution method and system for non-aligned multi-source data based on structural information, and belongs to the field of computational photography technology. The present invention first uses gradient calculation to extract gradient maps from input hyperspectral and multispectral images, and then uses a texture encoder and a structure encoder to encode the image and the gradient map respectively to obtain a texture feature pyramid and a structure feature pyramid, and then uses a feature fusion alignment module guided by structural attention to perform feature fusion at each level to obtain alignment features, and finally uses a decoder network to decode the alignment features to generate a high-resolution hyperspectral image. The present invention makes the feature alignment effect more robust and the super-resolution effect better. It achieves good results in super-resolution tasks of various magnifications of real data sets and simulation data sets, has advantages in high-magnification super-resolution tasks, and is easy to promote.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a spectral image fusion super-resolution method and system for non-aligned multi-source data based on structural information, belonging to the technical field of computational imaging. Background Art

[0002] Hyperspectral images are currently widely used in various fields such as earth observation, agriculture, environmental monitoring, and medical imaging. However, due to equipment limitations, high-resolution hyperspectral images are difficult to obtain directly. The fusion-based hyperspectral image super-resolution method can fuse low-resolution hyperspectral images (LR HSI) with high-resolution multispectral images to generate high-resolution hyperspectral images (HR MSI). Compared with single image super-resolution methods, fusion-based methods significantly improve performance. However, existing fusion-based methods rely on strictly aligned hyperspectral and multispectral images, which is difficult to achieve in practical application scenarios. As a result, the problem of misalignment in the fusion of multispectral and hyperspectral images arises, and LR HSI is super-resolved with reference to the misaligned HR MSI.

[0003] In recent years, researchers have focused on the fusion of unaligned multispectral and hyperspectral images, using methods such as displacement field, mutual information, and coarse and fine optical flow to align HR MSI and LR HIS. Among them, the optical flow-based method HSIFN achieved the best performance. Although various methods have been proposed, there are still two challenges in the field of super-resolution that have not been resolved. First, the misalignment between images is relatively complex, including distortion, blur, stretching, occlusion, etc., making it difficult to achieve effective alignment. Second, in actual situations, there is a large gap in magnification between LRHSI and HR MSI. While most existing methods focus on low-magnification super-resolution, their performance decreases significantly when the super-resolution magnification increases. Third, optical flow-based methods require mapping the HSI image to the same band as the MSI and aligning it using a spectral response matrix, which results in the loss of spectral information. Summary of the invention

[0004] The purpose of the present invention is to address the problems and shortcomings of the prior art, and to solve the technical problems such as poor alignment effect when fusion super-resolution of unaligned hyperspectral and multispectral images, and creatively propose a spectral image fusion super-resolution method and system based on unaligned multi-source data.

[0005] The innovation of the present invention includes: proposing a multi-spectral and hyperspectral image alignment and fusion network guided by structural features. Under the guidance of structural features, the features of each spectral channel are aligned separately, making the feature alignment effect more robust, thereby achieving better super-resolution effect.

[0006] Firstly, gradient calculation is used to extract the gradient maps of low-resolution hyperspectral image LR HSI and high-resolution hyperspectral image HRMSI. Then, the texture encoder and structure encoder are used to encode the image and gradient map respectively to obtain the texture feature pyramid and structure feature pyramid. Then, the feature fusion alignment module guided by structure attention is used to perform feature fusion at each level to obtain the alignment features. Finally, the decoder network is used to decode the alignment features to generate a high-resolution hyperspectral image (HR HSI).

[0007] The present invention is implemented by adopting the following technical solutions.

[0008] A spectral image fusion super-resolution method based on unaligned multi-source data includes the following steps:

[0009] Step 1: Extract the gradient maps of the low-resolution hyperspectral image LR HSI and the high-resolution hyperspectral image HR MSI.

[0010] Specifically, for each pixel in the image, its gradient is calculated using the following formula and used as the gradient value of the pixel:

[0011]

[0012] in, Indicates that the image is at point The gradient value of Indicates that the image is at point The pixel value of It means partial derivative.

[0013] Step 2: Use the texture encoder and structure encoder networks to encode the images and gradient maps of LR HSI and HR MSI respectively to obtain the texture feature pyramid and structure feature pyramid. Specifically, the texture encoder is composed of a series of convolutional networks, and the structure encoder is composed of a gradient enhancement module and a convolutional network.

[0014] Step 3: Use structured attention-guided feature fusion alignment to fuse hyperspectral and multispectral features at each level.

[0015] Among them, feature fusion alignment includes feature fusion alignment and structural attention guidance. Feature fusion alignment first uses channel convolution to reduce the dimension of each image's structural features and texture features, and then splices them together and passes them into the residual block to generate preliminary alignment features; structural attention guidance is to splice the structural features of hyperspectral and multispectral images together, first perform channel convolution to reduce the dimension, and then pass them together into the residual block to generate a gradient attention mask. Finally, the preliminary alignment feature is multiplied by the attention mask feature to obtain the alignment feature.

[0016] Step 4: Use the decoder network to fuse and decode the hyperspectral features with the alignment features to generate HR HSI.

[0017] On the other hand, to achieve the above objectives, the present invention further proposes a spectral image fusion super-resolution system based on non-aligned multi-source data, including a gradient encoding module, a feature encoding module, a feature fusion alignment module, a decoder module, an input module and an output module.

[0018] The input module is used to input HR MSI and LR HSI.

[0019] A gradient encoding module is used to encode the input HR MSI and LR HSI images to obtain their corresponding gradient maps;

[0020] The feature encoding module is used to encode the input image and gradient map to obtain the texture feature pyramid and the structural feature pyramid;

[0021] The feature fusion and alignment module is used to input texture features and gradient features, fuse and align texture features guided by structural information, and output alignment features.

[0022] The decoder module decodes the hyperspectral texture features step by step with the aid of alignment features to generate high-resolution hyperspectral images.

[0023] The output module is used to output high-resolution hyperspectral images.

[0024] The connection relationship of the above modules is as follows:

[0025] The output end of the input module is connected to the input end of the gradient encoding module. The output end of the gradient encoding module is connected to the input end of the feature encoding module. The output end of the feature encoding module is connected to the input end of the feature fusion alignment module. The output end of the feature fusion alignment module is connected to the input end of the decoder module, and the output end of the decoder module is connected to the input end of the output module.

[0026] Beneficial Effects

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] 1. This paper proposes a novel structural feature-guided multispectral and hyperspectral image alignment and fusion network. The network introduces structural features into the alignment and fusion process, thus providing a more robust and efficient method to handle misaligned images.

[0029] 2. The present invention introduces a structural attention-guided feature fusion and alignment module, which utilizes the structural features of HR HSI and LRHSI to guide the alignment and fusion of their texture features, achieving superior performance compared with existing methods.

[0030] 3. The present invention has achieved the best results in super-resolution tasks of various magnifications for real data sets and simulated data sets, especially in high-magnification super-resolution tasks, and has a high promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flow chart of the method of the present invention;

[0032] Figure 2 It is the overall structure diagram of the deep neural network of the method of the present invention;

[0033] Figure 3 It is a schematic diagram of the alignment fusion module and gradient extraction of the method of the present invention;

[0034] Figure 4 It is a schematic diagram of the structure of the system of the present invention. DETAILED DESCRIPTION

[0035] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0036] Example

[0037] like Figure 1 As shown, a spectral image fusion super-resolution method based on unaligned multi-source data includes the following steps:

[0038] Step 1: Calculate the gradient for each pixel of the hyperspectral and multispectral images to obtain the gradient map.

[0039] The gradient map is used as the representation of the gradient information. In order to obtain the gradient information, the gradient map of the input HR MSI and LR HSI needs to be calculated first.

[0040] Specifically, for each point in the image, its gradient in the x and y directions is calculated respectively, and then the square sum of the gradients in the two directions is calculated and the square root is taken to obtain the gradient of the point.

[0041] Step 2: Encode the image and gradient map respectively to obtain the texture feature pyramid and structure feature pyramid.

[0042] The encoder network is divided into two parts: texture encoder and structure encoder, which process images to extract texture features and gradient maps to extract structural features respectively. Figure 2 .

[0043] The texture encoder inputs an image and outputs its corresponding texture feature pyramid. The texture encoder consists of five consecutive Convolution-LeakyReLU layers. The convolution step size is set to 1,1,2,2,2. Starting from the second layer, the output of each layer is used as the level feature. The feature size of the next level is 1 / 2 of the previous level.

[0044] The structure encoder inputs the gradient image in step 1 and outputs its corresponding structure feature pyramid. The structure encoder consists of a gradient enhancement module and a structure feature encoding module. The directly extracted gradient image cannot represent the structural information of the image well. To this end, a gradient enhancement module is needed to enhance the gradient image. This module adopts an encoder-decoder structure. For the input features, the encoder is first used to encode the features, then the cyclic residual module is used to refine the intermediate features, then the mask module is used to generate the mask, and the mask is multiplied with the intermediate features, and finally the decoding module is used to generate the enhanced gradient image. Except for the network channels, the structure feature encoding module is similar to the texture encoder network. It inputs the enhanced gradient features and outputs the corresponding structure feature pyramid.

[0045] Step 3: Use the structure-guided feature alignment fusion module to perform feature fusion and alignment at each level.

[0046] Using the texture feature pyramid and structural feature pyramid extracted in step 2, the gradient attention-guided feature fusion alignment module is used at each level to generate the alignment features of that level. See the specific structure for details. Figure 3 .

[0047] The feature fusion alignment module is divided into a texture feature fusion alignment module and a gradient feature guidance module. The texture feature fusion alignment module inputs texture features and structural features. First, each texture feature and its corresponding texture feature single image are concatenated and fused using point-by-point convolution. Then, the three sets of fused features are passed to the residual block for further fusion to obtain preliminary fused features. The gradient feature guidance module inputs various structural features, concatenates them together and passes them to the convolution-residual-convolution network to generate attention weights based on the structural information. Finally, the structural attention is multiplied by the preliminary alignment feature to obtain the final alignment feature.

[0048] Step 4: Use the decoder network to decode the hyperspectral features level by level with reference to the aligned features.

[0049] The texture features of the hyperspectral image extracted in step 2 and the aligned images of each level extracted in step 3 are passed into the decoder network, and the texture features are decoded by referring to the aligned features step by step to obtain the HR MSI.

[0050] At this point, this method obtains the final high-resolution hyperspectral image.

[0051] like Figure 4 As shown, it is a structural schematic diagram of a spectral image fusion super-resolution system based on non-aligned multi-source data according to an embodiment of the present invention. The system includes: an input module 100, a gradient encoding module 200, a feature encoding module 300, a feature fusion alignment module 400, a decoder module 500 and an output module 600.

[0052] The input module 100 starts from the physical imaging process of the sensor, and transmits the HR MSI and LR HSI images to the gradient encoding module 200.

[0053] The gradient encoding module 200 is composed of a gradient encoder, which performs gradient encoding on the input HR MSI and LR HSI images to obtain their corresponding gradient maps.

[0054] The feature encoding module 300 is composed of a texture feature encoder and a structural feature encoder. The texture feature encoder encodes the input image to obtain a texture feature pyramid, and the structural feature encoder first enhances the gradient map and then encodes it to obtain a structural feature pyramid.

[0055] The feature fusion alignment module 400 uses structural features to guide texture features for fusion and alignment at each level of the feature pyramid. It first uses the feature fusion alignment module to generate a preliminary fusion alignment feature, and then uses the structural attention guidance module to generate a structural attention weight, which is multiplied by the preliminary fusion alignment feature to obtain the final alignment feature.

[0056] The decoder module 500 uses the HSI features encoded by the module 300 to perform layer-by-layer decoding with reference to the alignment features generated by the module 400 to obtain a high-resolution hyperspectral image.

[0057] Finally, the output module 600 transmits the output result of the decoder module 500 to the display.

[0058] The connection relationship of the above modules is as follows:

[0059] The output end of the input module 100 is connected to the input end of the gradient encoding module 200. The output end of the gradient encoding module 200 is connected to the input end of the feature encoding module 300. The output end of the feature encoding module 300 is connected to the input end of the feature fusion alignment module 400. The output end of the feature fusion alignment module 400 is connected to the input end of the decoder module 500, and the output end of the decoder module 500 is connected to the input end of the output module 600.

[0060] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A spectral image fusion super-resolution method based on unaligned multi-source data, characterized in that: The following steps are involved: Step 1: Use gradient calculation to extract the gradient map of the low-resolution hyperspectral image LR HSI and the high-resolution multispectral image HR MSI; Step 2: Encode the low-resolution hyperspectral image and the high-resolution multispectral image respectively to obtain the texture feature pyramid and the structural feature pyramid; Step 3: Use structured attention-guided feature fusion alignment to fuse hyperspectral and multispectral features at each level; Feature fusion alignment It includes two parts: feature fusion alignment and structural attention guidance; Among them, feature fusion alignment first uses channel convolution to reduce the dimension of each image's structural features and texture features, and then splices them together and passes them into the residual block to generate preliminary alignment features; Structural attention guidance is to stitch together the structural features of hyperspectral and multispectral images, first perform channel convolution to reduce the dimension, and then pass them together into the residual block to generate a gradient attention mask; Finally, the preliminary alignment feature is multiplied by the gradient attention mask to obtain the alignment feature; Step 4: Use the decoder network to fuse and decode the hyperspectral features and alignment features to generate a high-resolution hyperspectral image (HR HSI).

2. The spectral image fusion super-resolution method based on non-aligned multi-source data according to claim 1, characterized in that: In step 1, for each pixel in the image, its gradient is calculated using the following formula and used as the gradient value of the pixel: in, Indicates that the image is at point The gradient value of Indicates that the image is at point The pixel value of It means partial derivative.

3. The spectral image fusion super-resolution method based on non-aligned multi-source data according to claim 1, characterized in that: The texture encoder consists of a series of convolutional networks, and the structure encoder consists of a gradient enhancement module and a convolutional network; The texture encoder inputs an image and outputs its corresponding texture feature pyramid; The structure encoder inputs the gradient map in step 1 and outputs its corresponding structure feature pyramid.

4. The spectral image fusion super-resolution method based on non-aligned multi-source data according to claim 3, characterized in that: The texture encoder consists of five consecutive Convolution-LeakyReLU layers.

5. The spectral image fusion super-resolution method based on non-aligned multi-source data according to claim 4, characterized in that: The convolution stride is set to 1, 1, 2, 2, 2, and starting from the second layer, the output of each layer is used as the level feature.

6. The spectral image fusion super-resolution method based on non-aligned multi-source data according to claim 5, characterized in that: The feature size of the next level is 1 / 2 of the previous level.

7. A spectral image fusion super-resolution system based on non-aligned multi-source data, characterized in that: It includes a gradient encoding module, a feature encoding module, a feature fusion alignment module, a decoder module, an input module and an output module; Among them, the input module is used to input the gradient map of the low-resolution hyperspectral image LR HSI and the high-resolution multispectral image HR MSI; A gradient encoding module is used to encode the input HR MSI and LR HSI images to obtain their corresponding gradient maps; The feature encoding module is used to encode the input image and gradient map to obtain the texture feature pyramid and the structural feature pyramid; The feature fusion alignment module is used to input texture features and gradient features, fuse and align texture features guided by structural information, and output alignment features; feature fusion alignment includes two parts: feature fusion alignment and structural attention guidance; among them, feature fusion alignment first uses channel convolution to reduce the dimension of each image's structural features and texture features, and then splices them together and passes them into the residual block to generate preliminary alignment features; structural attention guidance is to splice the structural features of hyperspectral and multispectral images together, first perform channel convolution to reduce the dimension, and then pass them together into the residual block to generate a gradient attention mask; finally, the preliminary alignment feature is multiplied by the gradient attention mask to obtain the alignment feature; The decoder module decodes the hyperspectral texture features step by step with the aid of alignment features to generate high-resolution hyperspectral images; The output module is used to output high-resolution hyperspectral images; The connection relationship of the above modules is as follows: The output end of the input module is connected to the input end of the gradient encoding module; the output end of the gradient encoding module is connected to the input end of the feature encoding module; the output end of the feature encoding module is connected to the input end of the feature fusion alignment module; the output end of the feature fusion alignment module is connected to the input end of the decoder module, and the output end of the decoder module is connected to the input end of the output module.

Citation Information

Patent Citations

  • Hyperspectral super-resolution fusion method and system based on non-aligned RGB image

    CN114972022A

  • Hyperspectral imaging method and system based on double RGB image fusion, and medium

    WO2024027095A1