A multispectral image and panchromatic image fusion method and system based on full-spectrum space
By combining spatial and frequency domain feature processing in remote sensing image fusion, the improved Two-Stream Fusion Network (TSFNet) solves the problems of imperfect feature extraction and low computational efficiency in existing technologies, achieving high-quality, high-resolution multispectral image reconstruction and improving the fusion effect of remote sensing images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2026-03-27
AI Technical Summary
Existing remote sensing image fusion technologies suffer from imperfect feature extraction, lack of flexibility in fusion strategies, and low computational efficiency. They are unable to maintain the accuracy of spectral information while improving spatial resolution, and are prone to inaccurate detail restoration and frequency domain information distortion, especially when dealing with complex scenes.
An improved dual-stream fusion network (TSFNet) is adopted to process image features simultaneously in the spatial and frequency domains. Through residual fusion modules and multi-head attention mechanisms, combined with encoder-decoder architecture design, high-resolution reconstruction of panchromatic and multispectral images is achieved. The network training is optimized by using a multi-scale residual learning framework and progressive strategies, combined with multi-task loss functions.
Effectively balancing spatial-spectral information improves the fusion quality of images, maintains the accuracy of spectral information and spatial details, enhances computational efficiency and adaptability, and enables the reconstruction of high-quality, high-resolution multispectral images.
Smart Images

Figure CN120070195B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multispectral image and panchromatic image fusion reconstruction for obtaining high-resolution multispectral images, and in particular to a multispectral image and panchromatic image fusion method and system based on a full-spectrum space. BACKGROUND
[0002] With the rapid development of space technology, remote sensing technology plays an increasingly important role in earth observation, environmental monitoring, urban planning and other fields. These advanced remote sensing platforms not only significantly improve the spatial resolution of multispectral sensors, but also greatly improve the data acquisition frequency, providing massive high-quality multispectral and panchromatic image data for earth observation.
[0003] Corresponding to the progress of remote sensor hardware technology, data processing requirements are also constantly improving. First, the requirement for real-time data processing is increasingly urgent. With the increase in remote sensing data acquisition frequency, traditional offline processing methods have been unable to meet application requirements. Second, the processing accuracy requirement is constantly improving. In fine application scenarios such as urban planning and agricultural monitoring, higher requirements are placed on image fusion quality. In addition, the demand for automated processing has increased significantly. In the face of massive remote sensing data, manual processing methods have become unsustainable, and there is an urgent need to develop efficient automated processing algorithms.
[0004] At the algorithmic level, the rapid development of deep learning has brought new opportunities for remote sensing image processing. Deep learning models represented by CNN (Convolutional Neural Network), Transformer, etc. have made major breakthroughs in computer vision, demonstrating strong feature extraction and expression capabilities. At the same time, the widespread use of attention mechanisms has greatly improved model performance. In addition, the significant improvement in hardware computing power such as GPU has provided strong support for the training and deployment of complex models.
[0005] Traditional fusion methods play an important role in the field of remote sensing image fusion. Among them, the IHS (Intensity-Hue-Saturation) transformation method converts the RGB space to the IHS space for fusion, which has the advantage of simple calculation, but often leads to severe spectral distortion. The Gram-Schmidt orthogonal transformation method has improved in spectral preservation, but has low computational efficiency when processing high-resolution images, and the spatial detail preservation effect is not ideal. These traditional methods generally have problems such as severe spectral distortion, insufficient detail preservation, and low processing efficiency.
[0006] With the development of deep learning technology, fusion methods based on neural networks have gradually become a research hotspot. Early CNN-based models extract image features through multiple convolutional operations, which improves the fusion effect to some extent. The introduction of the Transformer architecture later uses self-attention mechanisms to capture long-range dependencies in images, further improving fusion performance. The application of GAN (Generative Adversarial Networks) has made significant progress in improving the realism of fused images. However, these methods still have problems such as incomplete feature extraction, unstable training process, and high computational resource consumption.
[0007] Current remote sensing image fusion technology faces three main challenges: First, the problem of incomplete feature extraction. Existing methods mainly focus on the extraction of spatial domain features, and the use of frequency domain information is severely insufficient. This single feature extraction method limits the model's feature expression ability, leading to obvious deficiencies in spectral fidelity and spatial detail preservation in the fusion results. Especially when dealing with complex scenes, due to the lack of effective fusion of multi-scale features, there are often problems such as inaccurate detail restoration and distortion of frequency domain information. Second, the lack of flexibility in fusion strategy. Existing fusion algorithms generally lack effective adaptive mechanisms and cannot dynamically adjust the fusion strategy according to the feature characteristics of different image regions. This rigid fusion method leads to inconsistent performance of the model when dealing with different types of images, especially in complex scenes or with heavy noise interference, the fusion effect is often not satisfactory. In addition, limited feature selection ability and poor regional adaptability also seriously restrict the actual application effect of the algorithm. Third, there are many difficulties in engineering implementation. In actual application, existing algorithms generally have low computational efficiency, large memory occupation, unstable training, and other problems. These problems directly lead to high deployment costs and poor real-time performance, severely limiting the application scenarios of the algorithm. Especially when dealing with large-size, high-resolution images, the problem of computational resource consumption is more prominent.
[0008] For example: the invention application with application number 202111322487.1 discloses a panchromatic multispectral image fusion method and device. The adaptive filtering network used in this application can dynamically generate filters according to the content of the input image, enhancing the adaptability of the filter to the content and better implementing fitting. The fusion result has achieved good visual effect without obvious spectral distortion and spatial distortion, and has also improved in quantitative evaluation indicators. The invention application with application number 202410947686.9 discloses a high-resolution multispectral image reconstruction method based on a dynamic edge-guided network. This application scheme can generate higher resolution multispectral images, while adaptively utilizing image edge priors without increasing additional computational burden.
[0009] But the above scheme also exists: space-spectral information balance problem, can not improve the spatial resolution while maintaining the accuracy of spectral information, can not solve the contradiction between spatial detail enhancement and spectral distortion, can not effectively avoid the appearance of spatial artifacts and spectral distortion in the fusion result; At the same time, the image feature extraction is not sufficient, mainly focusing on the spatial domain feature, ignoring the spectral information, resulting in lack of effective multi-scale feature extraction mechanism problem, the feature expression ability is limited, it is difficult to describe complex image structure; The above scheme also has the problem of low feature fusion efficiency, lack of adaptive feature selection mechanism, resulting in large information loss in the fusion process, the model generalization ability is insufficient, resulting in poor adaptability to images obtained by different sensors. SUMMARY
[0010] In view of the above problems, the purpose of the present application is to provide a multi-spectral image and panchromatic image fusion method and system based on full spectrum space, which fully utilizes spatial-spectral information to reconstruct high-quality high-resolution multi-spectral images.
[0011] The embodiment of the present application provides a multi-spectral image and panchromatic image fusion method and system based on full spectrum space.
[0012] The first aspect is a multi-spectral image and panchromatic image fusion method based on full spectrum space, which uses a trained improved double-flow fusion network (TSFNet) to perform high-resolution multi-spectral image reconstruction, and the steps include:
[0013] S1, obtaining panchromatic image and multi-spectral image, performing image preprocessing and normalization operation;
[0014] S2, simultaneously processing panchromatic image and multi-spectral image in spatial domain and frequency domain, capturing spatial features and frequency domain features of panchromatic image and multi-spectral image respectively;
[0015] S3, splicing the spatial features of panchromatic image and multi-spectral image to obtain image spatial features, and splicing the frequency domain features of panchromatic image and multi-spectral image to obtain image frequency domain features;
[0016] S4, using a residual fusion module (RFB) to perform residual fusion on the image spatial features and the image frequency domain features;
[0017] S5, reconstructing high-resolution multi-spectral image based on the fused features, performing quality evaluation, and outputting the qualified image.
[0018] Further, the image preprocessing and normalization operation in S1 includes:
[0019] For panchromatic image, the original spatial resolution is maintained; for multi-spectral image, bicubic interpolation is used for up-sampling to make it have the same spatial resolution as the panchromatic image.
[0020] Furthermore, step S2, which simultaneously processes panchromatic and multispectral images in both the spatial and frequency domains, includes:
[0021] In the spatial domain, spatial structural features of the image are extracted using a convolutional network;
[0022] In the frequency domain, the input image is first subjected to block-based discrete cosine transform (DCT), and then the frequency structure features of the image are extracted through a convolutional network.
[0023] Furthermore, the input image is subjected to block-based discrete cosine transform (DCT) in the frequency domain, including: performing block-based DCT transform on the input image, obtaining DCT transform coefficients, and performing block-based reconstruction and normalization processing based on the DCT transform coefficients.
[0024] Furthermore, step S4 includes the following steps:
[0025] S41. Perform spatial convolution processing on the image spatial features to obtain the convolutional spatial features.
[0026] S42. Simultaneously perform frequency domain convolution and upsampling on the image frequency domain features to obtain the convolutioned frequency domain features.
[0027] S43. Concatenate the convolutional spatial features and frequency domain features to obtain concatenated features;
[0028] S44. Input the stitched features into the EMSA module to output the fused features, and perform residual fusion with the image spatial features.
[0029] Furthermore, in step S44, the spliced features are input into the EMSA module, and the output fused features include:
[0030] S44a. The concatenated features are processed in parallel by multiple attention heads. Each attention head generates three sets of feature vectors: query, key, and value through linear transformation. Then, each attention head calculates the dot product of the query vector and the key vector, and performs scaling and softmax operations to obtain the attention weight. The attention weight is multiplied by the value vector to obtain the attention output.
[0031] S44b: Convolve the spliced features, extract the spatial features, and obtain the spatial feature output;
[0032] S44c: The attention output and spatial feature output are fused, normalized, and then the fused feature is output.
[0033] Furthermore, training the improved dual-stream fusion network (TSFNet) includes: calculating the mean squared error loss of the reconstructed image and the target image, calculating the high-frequency information preservation loss, calculating the peak signal-to-noise ratio loss, and then performing dual-stream fusion network training optimization after combining the losses.
[0034] The second aspect: A multispectral image and panchromatic image fusion system based on the full spectrum space, comprising:
[0035] The feature extraction module is used to extract image spatial features and image frequency domain features from the input panchromatic and multispectral images;
[0036] The feature fusion module performs residual fusion of image spatial features and image frequency domain features;
[0037] The image reconstruction module reconstructs high-resolution multispectral images based on residual fusion features.
[0038] Third aspect: An electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, performs the steps of the method provided in the first aspect.
[0039] Fourth aspect: A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect.
[0040] The beneficial effects of this invention are:
[0041] 1. This invention utilizes an improved dual-stream fusion network (TSFNet) with an encoder-decoder structure. At the input end, the network receives panchromatic and multispectral images separately, ensuring that the multispectral and panchromatic images have the same spatial resolution. During preprocessing, the input image is processed simultaneously in both the spatial and frequency domains. This dual-path feature extraction mechanism comprehensively captures different levels of image features; effectively balances spatial and spectral information, maintaining the accuracy of spectral information while improving spatial resolution, thus resolving the contradiction between spatial detail enhancement and spectral distortion in traditional methods. This avoids spatial artifacts and spectral distortion in the fusion results; it fully extracts image features, focusing on both spatial and frequency domain information, enabling the characterization of complex image structures; and it employs an adaptive feature selection mechanism, resulting in minimal information loss during fusion, high feature fusion efficiency, strong adaptability to images acquired from different sensors, and sensitivity to noise and imaging condition changes, enabling the reconstruction of high-quality, high-resolution multispectral images.
[0042] 2. The EMSA module of this invention employs a multi-head attention mechanism, dividing the input features into multiple attention heads for parallel processing. Each attention head can independently focus on different feature patterns, thereby achieving a comprehensive understanding of the input features. To enhance the effectiveness of the attention mechanism, this invention also introduces a learnable scaling factor. The scaling factor can adaptively adjust the importance of different attention heads, enabling the network to better adapt to different image content. Furthermore, this invention designs a relative position encoding mechanism, encoding positional information into the feature representation, enhancing the model's ability to perceive spatial location. In terms of feature fusion, the EMSA module adopts a dual-branch structure. The Epo branch enhances the expressive power of local features through grouped convolution operations, while the Epr branch utilizes a self-attention mechanism to capture long-range dependencies between features. The outputs of these two branches are fused through adaptive weights, preserving local details while establishing global semantic connections.
[0043] 3. In terms of frequency domain feature extraction, this invention performs block-based DCT transformation on the input image, effectively reducing the differences between different images. In further processing of frequency domain features, this invention employs a pixel unshuffle operation to rearrange the DCT coefficients. This rearrangement operation organizes adjacent frequency components together, facilitating subsequent extraction of frequency domain features by the convolutional network. This invention also designs a dedicated high-frequency component preservation module, which highlights important high-frequency information through an attention mechanism, ensuring that detailed features are not lost during feature extraction.
[0044] 4. The multi-scale residual learning framework constructed in this invention includes two branches for each residual block: spatial feature processing and frequency domain feature processing. The spatial feature branch extracts and enhances spatial domain features through multi-layer convolution operations, while the frequency domain feature branch processes and integrates frequency domain information. The features from both branches are upsampled through transposed convolution to ensure spatial scale matching of the feature maps. During feature fusion, an attention module is introduced to adaptively weight the features, highlighting important feature channels.
[0045] 5. This invention employs a progressive strategy for feature reconstruction; the network fuses features at different scales and transmits low-level detailed information through skip connections. This multi-scale processing approach can simultaneously consider global semantic information and local detailed features. To further improve the reconstruction effect, this invention introduces multiple global residual connections into the network. These connections can effectively reduce information loss during feature transmission and also accelerate the network's convergence speed.
[0046] 6. During network training, this invention employs a multi-component composite loss function that comprehensively considers image reconstruction quality, spectral fidelity, and spatial detail. Regarding reconstruction loss, it not only calculates the mean square error between the predicted result and the target image but also introduces a perception-based loss term. This perception loss extracts features through a pre-trained deep network and compares the differences between the predicted result and the target image in the feature space, thus better preserving the visual quality of the image. For spectral fidelity, this invention designs a dedicated spectral loss term. This term first downsamples the fusion result to the resolution of the original multispectral image and then calculates the angular difference between the spectral vectors. This design ensures accurate preservation of frequency domain information during fusion. Simultaneously, by calculating the correlation of spectral features, the fidelity of frequency domain information is further constrained. To preserve spatial detail, this invention adds a spatial detail loss term to the loss function. This term extracts edge information from the image using a designed high-pass filter and then compares the matching degree of the fusion result with the panchromatic image in terms of edge features. This design effectively ensures that the fusion result accurately reconstructs spatial details while preserving frequency domain information. This invention directly incorporates the PSNR metric into the loss function, setting the target PSNR value as the optimization objective. By dynamically adjusting the weights of each loss term, the fusion effect is gradually improved during training. This PSNR-based optimization strategy can more directly guide the network to generate high-quality fusion results. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating the structure of the system of the present invention;
[0048] Figure 2 This is a schematic flowchart of the method of the present invention;
[0049] Figure 3 This is a schematic diagram of the system of the present invention;
[0050] Figure 4 This is a schematic diagram of the feature fusion process of the present invention;
[0051] Figure 5 This is a schematic diagram of the EMSA module flow of the present invention;
[0052] Figure 6 This is a structural diagram of the electronic device of the present invention. Detailed Implementation
[0053] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar symbols denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0054] Current remote sensing image fusion technology still faces several challenges in practical applications: image fusion methods primarily focus on spatial domain feature extraction, with insufficient utilization of spectral information; deep learning fusion models lack effective feature selection mechanisms and cannot adaptively fuse based on the features of different image regions; and existing fusion model algorithms suffer from low computational efficiency and training instability when processing high-resolution, large-size images.
[0055] To address the aforementioned problems, this invention provides a method for fusing multispectral and panchromatic images based on the full spectrum space. It utilizes a trained two-stream fusion network (TSFNet) for high-resolution multispectral image reconstruction. Figure 2 This is a flowchart illustrating the fusion method of the present invention, which includes the following steps:
[0056] S1. Acquire panchromatic and multispectral images, and perform image preprocessing and normalization operations.
[0057] like Figure 1 As shown, this invention is based on an improved dual-stream fusion network (TSFNet), which employs an encoder-decoder architecture. At the input end, the TSFNet network receives panchromatic and multispectral images respectively. For the panchromatic image, its original spatial resolution is maintained, serving as a source of information providing high spatial frequency details; for the multispectral image, it is first upsampled using bicubic interpolation to achieve the same spatial resolution as the panchromatic image, laying the foundation for subsequent feature extraction and fusion.
[0058] Specifically, the system parameters and operating environment are initialized first, the GPU device and memory resources are configured, the pre-trained model weights are loaded, and the processing parameters (such as batch size, image size, etc.) are set to obtain panchromatic (PAN) images and multispectral (MS) images.
[0059] PAN images are high spatial resolution panchromatic images in 3-channel RGB format, while MS images are low spatial resolution multispectral images, also in 3-channel RGB format. Both PAN and MS images support common formats such as PNG and JPEG.
[0060] The PIL library can be used to read PAN and MS images, convert them to RGB mode, and record the original resolution of the PAN image (usually 256×256) and the original resolution of the MS image (usually 64×64).
[0061] Image preprocessing and normalization operations: Preprocessing includes image scaling and data type conversion. Image scaling refers to upsampling the MS image to the PAN image resolution using a bicubic interpolation algorithm. Data type conversion refers to converting the image data to floating-point type and normalizing the range to [0,1]. Specifically, image scaling can be performed using transforms.Resize, and data conversion can be performed using transforms.ToTensor.
[0062] Normalization refers to the standardization process of an image, including normalizing the image mean by subtracting the channel mean; normalizing the standard deviation by dividing by the channel standard deviation; and adjusting the image value range to [-1, 1]. Specifically, it can be implemented using `transforms.Normalize`, with the mean set to [0.5, 0.5, 0.5] and the standard deviation set to [0.5, 0.5, 0.5].
[0063] S2. Simultaneously process panchromatic and multispectral images in the spatial and frequency domains to capture the spatial and frequency domain features of each panchromatic and multispectral image.
[0064] like Figure 3 As shown, parallel processing is performed on the spatial and frequency domains of panchromatic and multispectral images. In the spatial domain, spatial structural features of the image are extracted through a convolutional network. In the frequency domain, the input image is first subjected to block discrete cosine transform (DCT), and then the frequency structural features of the image are extracted through a convolutional network. Complementary features are extracted through different processing branches to maintain the integrity and independence of the features.
[0065] Specifically, spatial domain feature extraction includes convolution using the ConvPReLUConv module, with input channels of 3 and output channels of 32 (PAN) / 96 (MS), feature normalization using BatchNorm2d, and PReLU as the activation function, with residual connections to maintain gradient flow.
[0066] The specific implementation code can be:
[0067] classConvPReLUConv(nn.Module):
[0068] def__init__(self,in_channels,out_channels):
[0069] self.conv1=nn.Conv2d(in_channels,out_channels,kernel_size=3,padding=1)
[0070] self.bn1=nn.BatchNorm2d(out_channels)
[0071] self.conv2=nn.Conv2d(out_channels, out_channels, kernel_size=3, padding=1)
[0072] self.bn2=nn.BatchNorm2d(out_channels)
[0073] self.prelu = nn.PReLU()
[0074] Frequency domain feature extraction includes performing block-based discrete cosine transform (DCT) on the input image in the frequency domain. First, the input image is subjected to block-based DCT transform to obtain the DCT transform coefficients. Then, based on the DCT transform coefficients, block-based recombination and normalization processing is performed to extract frequency domain features.
[0075] Block DCT transform refers to dividing the image into blocks (e.g., 4×4 size) and then performing orthogonal DCT transform on each block to preserve the integrity of the frequency domain information. Specifically, it can be implemented using the function `def apply_dct(self, img, block_size = 4)` to perform block DCT transform on each channel separately.
[0076] Block reorganization uses the Pixel Unshuffle operation to expand the channels to: C->C*(block_size^2), maintaining spatial correlation and facilitating subsequent convolution processing. Specifically, it can be implemented using def pixel_unshuffle(self, img, block_size) to reorganize the feature map, reduce spatial resolution, and increase the number of channels.
[0077] For frequency domain feature extraction, PAN images can be output with 48 channels, and MS images with 192 channels. BatchNorm is used for normalization, and nonlinear activation is used to enhance feature representation.
[0078] S3. The spatial features of the panchromatic image and the multispectral image are stitched together to obtain the image spatial features, and the frequency domain features of the panchromatic image and the multispectral image are stitched together to obtain the image frequency domain features.
[0079] Spatial and frequency domain features are concatenated using a channel attention mechanism. Adaptive weight learning is employed to evaluate feature importance. Torch.cat is used for feature concatenation, and the feature weights are adjusted through the attention mechanism to obtain the spatial and frequency domain features of the concatenated image.
[0080] S4. The residual fusion module (RFB) is used to perform residual fusion on the spatial features and frequency domain features of the image.
[0081] like Figure 4 As shown, the residual fusion module (RFB) performs residual fusion of image spatial features and image frequency domain features, including the following steps:
[0082] S41. Perform spatial convolution processing on the image spatial features to obtain the convolutional spatial features.
[0083] S42. Simultaneously perform frequency domain convolution and upsampling on the image frequency domain features to obtain the convolutioned frequency domain features.
[0084] S43. Concatenate the convolutional spatial features and frequency domain features to obtain concatenated features;
[0085] S44. Input the stitched features into the EMSA module to output the fused features, and perform residual fusion with the image spatial features.
[0086] Specifically, the Residual Fusion Block (RFB) mainly includes spatial branching, frequency domain branching, an EMSA module, and a residual connection module. This structure preserves the original feature information while improving gradient flow. A concrete implementation can use the class `ResidualFusionBlock(nn.Module):def__init__(self,Cs,Cf)` to implement the spatial convolution and frequency domain processing of the residual fusion block.
[0087] like Figure 5 As shown, the EMSA module employs a multi-head self-attention mechanism for local-global feature interaction. The EMSA module uses a dual-branch structure, including an Epo (Enhanced Position-wise Operation) branch and an Epr (Enhanced Pairwise Relation) branch. The Epo branch enhances the expressive power of local features through grouped convolutional operations, while the Epr branch utilizes a self-attention mechanism to capture long-range dependencies between features. The outputs of these two branches are fused using adaptive weights, preserving local details while establishing global semantic connections. The steps include:
[0088] S44a. The concatenated features are processed in parallel by multiple attention heads. Each attention head generates three sets of feature vectors: query, key, and value through linear transformation. Then, each attention head calculates the dot product of the query vector and the key vector, and performs scaling and softmax operations to obtain the attention weight. The attention weight is multiplied by the value vector to obtain the attention output. Specifically, the class ESMSA(nn.Module) and def__init__(self, C, Cs, h = 8, b = 16) can be used to implement enhanced spatial multi-head self-attention.
[0089] S44b: Convolve the spliced features, extract the spatial features, and obtain the spatial feature output;
[0090] S44c: The attention output and spatial feature output are fused, normalized, and then the fused feature is output.
[0091] S5. Reconstruct high-resolution multispectral images based on the fused features, perform quality assessment, and output images that meet the standards.
[0092] A specific implementation can be achieved using:
[0093] self.output_conv=nn.Sequential(
[0094] nn.Conv2d(128,64,kernel_size=3,padding=1),
[0095] nn.BatchNorm2d(64),
[0096] nn.PReLU(),
[0097] nn.Conv2d(64,ms_channels,kernel_size=3,padding=1),
[0098] nn.Tanh()
[0099] Channel dimensionality reduction and feature reconstruction are performed, and the final output is obtained using PReLU and Tanh functions.
[0100] Training the two-stream fusion network (TSFNet) involves: calculating the mean squared error loss of the reconstructed image and the target image, calculating the high-frequency information preservation loss, calculating the peak signal-to-noise ratio loss, and then optimizing the training of the two-stream fusion network after combining the losses.
[0101] For spectral preservation, a dedicated spectral loss term is employed. This term first downsamples the fused result to the resolution of the original multispectral image, and then calculates the angular difference between the spectral vectors of the two. This ensures the accurate preservation of frequency domain information during the fusion process. Furthermore, the fidelity of the frequency domain information is further constrained by calculating the correlation of spectral features.
[0102] To preserve spatial details, a spatial detail loss term was added to the loss function. This term extracts edge information from the image using a designed high-pass filter, and then compares the fused result with the panchromatic image in terms of edge feature matching. This design effectively ensures that the fused result accurately reconstructs spatial details while preserving spectral information.
[0103] By directly incorporating the PSNR metric into the loss function and setting a target PSNR value as the optimization objective, the fusion performance is gradually improved during training by dynamically adjusting the weights of each loss term. Structural similarity is evaluated using PSNR calculation, with a PSNR threshold of 47.0. A threshold check of the loss function is performed, and the ImprovedLoss class is used for evaluation. This PSNR-based optimization strategy, based on a comprehensive judgment of multiple metrics, can more directly guide the network to generate high-quality fusion results.
[0104] like Figure 3 As shown, based on the above method, the present invention also discloses an image fusion system, comprising:
[0105] The feature extraction module includes a spatial feature extraction unit and a frequency domain feature extraction unit. The spatial feature extraction unit contains a multi-layer convolutional neural network and a batch normalization layer; the frequency domain feature extraction unit contains a discrete cosine transform layer and a convolutional layer. The feature extraction module is used to extract spatial and frequency domain features from the panchromatic and multispectral images of the input layer.
[0106] The feature fusion module includes a spatial processing unit, a frequency domain processing unit, an EMSA unit, and a residual connection unit. The spatial processing unit is used to process spatial features, the frequency domain processing unit is used to process frequency domain features, the EMSA unit is used for feature fusion, and the residual connection unit is used for information transfer. The feature fusion module is used to perform residual fusion of image spatial features and image frequency domain features.
[0107] The EMSA unit includes a matrix construction unit, a weight calculation unit, and a feature weighting unit. The matrix construction unit is used to construct the query matrix, key matrix, and value matrix. The weight calculation unit is used to calculate the attention weights. The feature weighting unit is used to perform feature weighted fusion.
[0108] The image reconstruction module reconstructs high-resolution multispectral images based on residual fusion features.
[0109] The training of the improved dual-stream fusion network (TSFNet) can utilize a loss calculation unit and a network optimization unit. The loss calculation unit is used to calculate the reconstruction loss, high-frequency information preservation loss, and peak signal-to-noise ratio loss; the network optimization unit is used to optimize network parameters based on the loss function.
[0110] This invention utilizes a feature extraction module to acquire panchromatic and multispectral image data from the input layer. The panchromatic image data is convolved to obtain panchromatic image spatial features, and the multispectral image data is convolved to obtain multispectral image spatial features. Discrete cosine transform is performed on the panchromatic and multispectral image data to obtain frequency domain data, and convolution is then performed to obtain panchromatic and multispectral image frequency domain features. The panchromatic and multispectral image spatial features are concatenated to obtain image spatial features and image frequency domain features. A feature fusion module constructs a query matrix, a key matrix, and a value matrix. Attention weights are calculated based on the query and key matrices and weighted with the value matrix to obtain the attention output. The spatial feature output and the attention output are fused, and residual fusion is performed using the fused features and the image spatial features. Finally, an image reconstruction module reconstructs a high-resolution multispectral image based on the fused features.
[0111] Specific implementation examples are provided below:
[0112] An improved TSFNet network structure is adopted, as follows: Figure 1 As shown, the network structure in this embodiment mainly includes the following core components:
[0113] The network structure of the input layer receives four inputs: panchromatic image (PAN): 3 channels, size 256×256; multispectral image (MS): 3 channels, original size 64×64, upsampled to 256×256; DCT transformation result of PAN image; DCT transformation result of MS image;
[0114] Then, feature extraction is performed using the feature extraction module, which contains the following key component code:
[0115] Python
[0116] Copy
[0117] self.pan_input_conv=ConvPReLUConv(pan_channels=3,out_channels=32)
[0118] self.ms_input_conv=ConvPReLUConv(ms_channels=3,out_channels=96)
[0119] self.pan_dct_conv=ConvPReLUConv(pan_channels*4,out_channels=48)
[0120] self.ms_dct_conv=ConvPReLUConv(ms_channels*4,out_channels=192)
[0121] The specific implementation of the ConvPReLUConv module is as follows:
[0122] Python
[0123] Copy
[0124] classConvPReLUConv(nn.Module):
[0125] def__init__(self,in_channels,out_channels):
[0126] super(ConvPReLUConv,self).__init__()
[0127] self.conv1=nn.Conv2d(in_channels,out_channels,kernel_size=3,padding=1)
[0128] self.bn1=nn.BatchNorm2d(out_channels)
[0129] self.conv2=nn.Conv2d(out_channels, out_channels, kernel_size=3, padding=1)
[0130] self.bn2=nn.BatchNorm2d(out_channels)
[0131] self.prelu = nn.PReLU()
[0132] Then, feature fusion is performed using a feature fusion module, which employs an improved attention mechanism, ESMSA (Enhanced Spatial-Spectral Multi-head Self-Attention). Its core implementation includes:
[0133] Multi-head self-attention calculation:
[0134] python
[0135] Copy
[0136] Q = Q.view(B, -1, self.h, C / / self.h).permute(0, 2, 3, 1)
[0137] K = K.view(B, -1, self.h, C / / self.h).permute(0, 2, 3, 1)
[0138] V = V.view(B, -1, self.h, C / / self.h).permute(0, 2, 1, 3)
[0139] attn = torch.matmul(Q, K.transpose(-2, -1))
[0140] attn = attn * self.sigma.view(1, -1, 1, 1) / ((C / / self.h)**0.5 + 1e - 8)
[0141] attn = torch.softmax(attn, dim=-1)
[0142] Residual Fusion Block:
[0143] python
[0144] Copy
[0145] class ResidualFusionBlock(nn.Module):
[0146] def __init__(self, Cs, Cf):
[0147] self.spatial_conv = ConvPReLUConv(Cs, Cs)
[0148] self.freq_conv = ConvPReLUConv(Cf, Cf)
[0149] self.t_conv = nn.Sequential(nn.ConvTranspose2d(Cf, Cf, kernel_size = 4, stride = 2, padding = 1), nn.PReLU())
[0150] self.e_smsa=ESMSA(Cs+Cf,Cs,h=8,b=16)
[0151] The output reconstruction module employs a progressive reconstruction strategy, including:
[0152] Python
[0153] Copy
[0154] self.output_conv=nn.Sequential(nn.Conv2d(128,64,kernel_size=3,padding=1),
[0155] nn.BatchNorm2d(64),nn.PReLU(),nn.Conv2d(64,ms_channels,kernel_size=3,pa dding=1),nn.Tanh())
[0156] The training strategy for the network includes: designing a loss function; this embodiment uses an improved multi-task loss function.
[0157] Python
[0158] Copy
[0159] classImprovedLoss(nn.Module):
[0160] def__init__(self, alpha=5.0, beta=0.0, gamma=0.0, psnr_weight=0.5, target_psnr=47.0):
[0161] super(ImprovedLoss,self).__init__()
[0162] self.mse = nn.MSELoss()
[0163] self.alpha = alpha
[0164] self.beta = beta
[0165] self.gamma = gamma
[0166] self.psnr_weight=psnr_weight
[0167] self.target_psnr=target_psnr
[0168] The loss function consists of four parts:
[0169] Reconstruction loss (L_hrms): Measures the pixel-level difference between the generated image and the target image.
[0170] Multispectral consistency loss (L_ms): Ensures preservation of frequency domain information
[0171] Spatial structure loss (L_spatial): Preserving spatial detail information
[0172] PSNR-guided loss: Optimizes overall reconstruction quality
[0173] To optimize the network, the AdamW optimizer was used, and a warmup and step-wise learning rate adjustment strategy were designed:
[0174] Python
[0175] Copy
[0176] optimizer=AdamW(model.parameters(),lr=2e-4,weight_decay=0.05, betas=(0.9,0.999),eps=1e-8)
[0177] scheduler=StepLRScheduler(optimizer,init_lr=2e-4,milestones=[50,100,150,200,250,300],gamma=0.5,warmup_epochs=5,warmup_start_lr=1e-6)
[0178] The main parameters for the training process are set as follows: batch size: 8; number of training epochs: 2000; gradient accumulation steps: 4; validation interval: 5 epochs; early stopping mechanism: stop if there is no improvement after 20 epochs; checkpoint saving interval: 20 epochs.
[0179] This invention features a meticulously designed and optimized network structure and training parameters. Regarding the number of feature channels, spatial features are progressively increased from an initial 32 channels to 128 channels, while frequency domain features are increased from 48 channels to 240 channels. This gradual channel expansion strategy ensures sufficient feature extraction while avoiding excessive consumption of computational resources.
[0180] In terms of training strategy, this invention employs mixed-precision training to improve training efficiency. The batch size is set to 8, and gradient accumulation is used, updating parameters after accumulating four steps. This setup ensures training stability while also considering hardware resource limitations. The learning rate adopts a OneCycleLR strategy, with an initial value set to 4e-4, dynamically adjusted during training according to cosine annealing rules. This learning rate scheduling scheme effectively improves the model's convergence speed and performance.
[0181] Regarding the parameter configuration of the loss function, the reconstruction loss weight is set to 1.0 as the main optimization objective; the weights of spectral loss and spatial loss are set to 0.1 and 0.01 respectively to balance the reconstruction effects of different aspects; the weight of PSNR loss is set to 0.1, and the target PSNR value is set to 40.0. These parameter settings have been verified by a large number of experiments and can achieve good fusion results.
[0182] This invention utilizes an improved dual-stream fusion network (TSFNet) with an encoder-decoder architecture. At the input, the network receives both panchromatic and multispectral images. For the panchromatic image, its original spatial resolution is maintained, serving as a source of high spatial frequency detail. For the multispectral image, bicubic interpolation is used for upsampling to achieve the same spatial resolution as the panchromatic image. In the preprocessing stage, this invention processes the input image simultaneously in the spatial and frequency domains. In the spatial domain, spatial structural features are directly extracted from the image using a multi-layer convolutional network. In the frequency domain, the input image is first subjected to a block-based discrete cosine transform (DCT), and then frequency domain information is obtained through a specially designed frequency domain feature extraction network. This dual-path feature extraction mechanism comprehensively captures different levels of image features, providing a sufficient information foundation for subsequent high-quality fusion.
[0183] The present invention also provides an electronic device, Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, the electronic device may include a processor, a communications interface, memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory, for example, to execute the following method:
[0184] S1. Acquire panchromatic and multispectral images, and perform image preprocessing and normalization operations;
[0185] S2. Simultaneously process panchromatic and multispectral images in the spatial and frequency domains to capture the spatial and frequency domain features of each panchromatic and multispectral image.
[0186] S3. The spatial features of the panchromatic image and the multispectral image are stitched together to obtain the image spatial features, and the frequency domain features of the panchromatic image and the multispectral image are stitched together to obtain the image frequency domain features.
[0187] S4. Residual fusion module (RFB) is used to perform residual fusion of image spatial features and image frequency domain features;
[0188] S5. Reconstruct high-resolution multispectral images based on the fused features, perform quality assessment, and output images that meet the standards.
[0189] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0190] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments, including, for example:
[0191] S1. Acquire panchromatic and multispectral images, and perform image preprocessing and normalization operations;
[0192] S2. Simultaneously process panchromatic and multispectral images in the spatial and frequency domains to capture the spatial and frequency domain features of each panchromatic and multispectral image.
[0193] S3. The spatial features of the panchromatic image and the multispectral image are stitched together to obtain the image spatial features, and the frequency domain features of the panchromatic image and the multispectral image are stitched together to obtain the image frequency domain features.
[0194] S4. Residual fusion module (RFB) is used to perform residual fusion of image spatial features and image frequency domain features;
[0195] S5. Reconstruct high-resolution multispectral images based on the fused features, perform quality assessment, and output images that meet the standards.
[0196] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fusing multispectral and panchromatic images based on an attention mechanism, characterized in that, High-resolution multispectral image reconstruction is performed using a trained improved dual-stream fusion network (TSFNet). The steps include: S1. Acquire panchromatic and multispectral images, and perform image preprocessing and normalization operations; S2. Simultaneously process panchromatic and multispectral images in the spatial and frequency domains to capture the spatial and frequency domain features of each panchromatic and multispectral image. S3. The spatial features of the panchromatic image and the multispectral image are stitched together to obtain the image spatial features, and the frequency domain features of the panchromatic image and the multispectral image are stitched together to obtain the image frequency domain features. S4. The residual fusion module RFB is used to perform residual fusion of image spatial features and image frequency domain features; S5. Reconstruct high-resolution multispectral images based on the fused features, perform quality assessment, and output images that meet the standards. The process of simultaneously processing panchromatic and multispectral images in the spatial and frequency domains in S2 includes: In the spatial domain, spatial structural features of the image are extracted using a convolutional network; In the frequency domain, the input image is first subjected to block-based discrete cosine transform (DCT), and then the frequency structure features of the image are extracted through a convolutional network. Performing block-based discrete cosine transform (DCT) on the input image in the frequency domain includes: performing block-based DCT transform on the input image, obtaining DCT transform coefficients, and performing block-based reconstruction and normalization processing based on the DCT transform coefficients. S4 includes the following steps: S41. Perform spatial convolution processing on the image spatial features to obtain the convolutional spatial features. S42. Simultaneously perform frequency domain convolution and upsampling on the image frequency domain features to obtain the convolutioned frequency domain features. S43. Concatenate the convolutional spatial features and frequency domain features to obtain concatenated features; S44. Input the stitched features into the EMSA module to output the fused features, and perform residual fusion with the image spatial features; A progressive strategy is adopted for feature reconstruction; the network fuses features at different scales and passes low-level detailed information through skip connections. This multi-scale processing approach can simultaneously take into account global semantic information and local detailed features.
2. The fusion method according to claim 1, characterized in that, The image preprocessing and normalization operations in S1 include: For panchromatic images, the original spatial resolution is maintained; for multispectral images, upsampling is performed using bicubic interpolation to give them the same spatial resolution as the panchromatic images.
3. The fusion method according to claim 1, characterized in that, In step S44, the spliced features are input into the EMSA module, and the output fused features include: S44a. The concatenated features are processed in parallel by multiple attention heads. Each attention head generates three sets of feature vectors: query, key, and value through linear transformation. Then, each attention head calculates the dot product of the query vector and the key vector, and performs scaling and softmax operations to obtain the attention weight. The attention weight is multiplied by the value vector to obtain the attention output. S44b: Convolve the spliced features, extract the spatial features, and obtain the spatial feature output; S44c: The attention output and spatial feature output are fused, normalized, and then the fused feature is output.
4. The fusion method according to claim 1, characterized in that, The improved dual-stream fusion network (TSFNet) is trained by: calculating the mean squared error loss of the reconstructed image and the target image, calculating the high-frequency information preservation loss, calculating the peak signal-to-noise ratio loss, and then training and optimizing the dual-stream fusion network after combining the losses.
5. An image fusion system based on the method of any one of claims 1 to 4, characterized in that, include: The feature extraction module is used to extract image spatial features and image frequency domain features from the input panchromatic and multispectral images; The feature fusion module performs residual fusion of image spatial features and image frequency domain features; The image reconstruction module reconstructs high-resolution multispectral images based on residual fusion features.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as claimed in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as claimed in any one of claims 1 to 4.
Citation Information
Patent Citations
Panchromatic multispectral image fusion method and device
CN114331930A
High-resolution multispectral image reconstruction method based on dynamic edge-guided network
CN118505509B
Hyperspectral and multispectral image fusion method based on attention mechanism
CN117474781A
Multispectral remote sensing image enhancement method based on frequency domain-space double-domain learning
CN118587097A