Multi-spectral image and panchromatic image fusion method and system based on full-spectrum space

By adopting an improved dual-stream fusion network (TSFNet) based on full spectrum space in remote sensing image fusion technology, and simultaneously processing image features in the spatial domain and frequency domain, the problem of imperfect feature extraction and lack of flexibility in the fusion strategy in the prior art is solved, and high-quality high-resolution multi-spectral image reconstruction is achieved.

CN120070195AActive Publication Date: 2025-05-30ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD

Patent Information

Application Number
CN202510009323.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-30
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The existing remote sensing image fusion technology has problems such as imperfect feature extraction, lack of flexibility in fusion strategies and difficulty in engineering implementation, resulting in insufficient spectral fidelity and spatial detail maintenance, low computational efficiency, and unstable training.

Method used

Using an improved dual-stream fusion network (TSFNet) based on full spectrum space, image features are processed simultaneously in the spatial domain and frequency domain, and feature stitching and fusion is performed through the residual fusion module and the ESMSA module to reconstruct high-resolution multispectral images.

Benefits of technology

Effectively balance spatial-spectral information, improve spatial resolution while maintaining the accuracy of spectral information, solve the contradiction between spatial detail enhancement and spectral distortion, and realize high-quality high-resolution multi-spectral image reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070195A_ABST
    Figure CN120070195A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral image and panchromatic image fusion method and system based on a full-spectrum space, high-resolution multispectral image reconstruction is performed by using a trained improved double-flow fusion network (TSFNet), and the method comprises the following steps: acquiring a panchromatic image and a multispectral image, and capturing respective spatial features and frequency domain features; splicing the frequency domain features of the panchromatic image and the multispectral image to obtain image frequency domain features; carrying out residual fusion on the image spatial features and the image frequency domain features by adopting a residual fusion module (RFB); and reconstructing a high-resolution multispectral image based on the fused features, and outputting an up-to-standard image after quality evaluation. According to the method, the improved double-flow fusion network (TSFNet) is utilized to effectively balance space-spectrum information, the accuracy of frequency domain information is kept while the spatial resolution is improved, the information loss in the fusion process is small, and a high-quality high-resolution multispectral image can be reconstructed and obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi - spectral image and pan - chromatic image fusion and reconstruction to obtain high - resolution multi - spectral images, and in particular, to a multi - spectral image and pan - chromatic image fusion method and system based on the full - spectrum space. Background Art

[0002] With the rapid development of space technology, remote sensing technology plays an increasingly important role in fields such as earth observation, environmental monitoring, and urban planning. These advanced remote sensing platforms have not only significantly improved the spatial resolution of multi - spectral sensors but also greatly increased the data acquisition frequency, providing a vast amount of high - quality multi - spectral and pan - chromatic image data for earth observation.

[0003] Corresponding to the progress of remote sensor hardware technology, the data processing requirements are also continuously increasing. First, the requirement for real - time data processing is becoming increasingly urgent. With the increase in the remote sensing data acquisition frequency, the traditional offline processing method can no longer meet the application requirements. Second, the processing accuracy requirements are continuously improving. In refined application scenarios such as urban planning and agricultural monitoring, higher requirements are put forward for the image fusion quality. In addition, the demand for automated processing has increased significantly. Facing the vast amount of remote sensing data, the manual processing method is no longer sustainable, and there is an urgent need to develop efficient automated processing algorithms.

[0004] At the algorithm technology level, the booming development of deep learning has brought new opportunities to remote sensing image processing. Deep learning models represented by CNN (Convolutional Neural Network), Transformer, etc. have made significant breakthroughs in the field of computer vision, demonstrating powerful feature extraction and expression capabilities. At the same time, the wide application of the attention mechanism has greatly improved the performance of the model. In addition, the significant improvement in the hardware computing power of GPUs and the like provides strong support for the training and deployment of complex models.

[0005] In the field of remote sensing image fusion, traditional fusion methods, methods based on the transform domain occupy an important position. Among them, the IHS (Intensity - Hue - Saturation) transform method performs fusion by converting the RGB space to the IHS space, which has the advantage of simple calculation but often leads to serious spectral distortion. Although the Gram - Schmidt orthogonal transform method has improved in spectral preservation, its computational efficiency is low when processing high - resolution images, and the effect of maintaining spatial details is not ideal enough. These traditional methods generally have problems such as serious spectral distortion, insufficient detail preservation, and low processing efficiency.

[0006] With the development of deep learning technology, neural network-based fusion methods have gradually become a research hotspot. Early CNN-based models extracted image features through multi-layer convolution operations, improving the fusion effect to a certain extent. The later introduced Transformer architecture utilized the self-attention mechanism to capture the long-range dependencies of images, further enhancing the fusion performance. The application of GAN (Generative Adversarial Network) has made significant progress in improving the authenticity of fused images. However, these methods still have problems such as incomplete feature extraction, unstable training processes, and high computational resource consumption.

[0007] Current remote sensing image fusion technology faces challenges in three main aspects: Firstly, there is the problem of imperfect feature extraction. Existing methods mainly focus on the extraction of spatial domain features and severely lack the utilization of frequency domain information. This single feature extraction method limits the feature expression ability of the model, resulting in obvious deficiencies in spectral fidelity and spatial detail preservation in the fusion results. Especially when dealing with complex scenes, due to the lack of effective fusion of multi-scale features, problems such as inaccurate detail restoration and frequency domain information distortion often occur. Secondly, the fusion strategy lacks flexibility. Existing fusion algorithms generally lack effective adaptive mechanisms and cannot dynamically adjust the fusion strategy according to the feature characteristics of different image regions. This rigid fusion method leads to inconsistent performance of the model when dealing with different types of images. Especially in complex scenes or cases with high noise interference, the fusion effect is often unsatisfactory. In addition, limited feature selection ability and poor regional adaptability also severely restrict the practical application effect of the algorithm. Thirdly, there are many difficulties in engineering implementation. In practical applications, existing algorithms generally have problems such as low computational efficiency, large memory occupancy, and unstable training. These problems directly lead to high deployment costs and poor real-time performance, severely limiting the application scenarios of the algorithm. Especially when dealing with large-size, high-resolution images, the problem of computational resource consumption is more prominent.

[0008] For example: The invention application with the application number 202111322487.1 discloses a panchromatic multi-spectral image fusion method and device. The proposed scheme can dynamically generate filters according to the input image content by using an adaptive filtering network, enhancing the adaptability of the filter to the content and enabling better fitting. The fusion result has achieved good results visually, without obvious spectral distortion and spatial distortion phenomena, and has also been improved in quantitative evaluation indicators. The invention application with the application number 202410947686.9 discloses a high-resolution multi-spectral image reconstruction method based on a dynamic edge guidance network. The proposed scheme can generate higher-resolution multi-spectral images, and at the same time can adaptively utilize the image edge prior without increasing additional computational burdens.

[0009] However, the above solutions also have the following problems: the problem of spatial-spectral information balance, which cannot improve the spatial resolution while maintaining the accuracy of spectral information, cannot solve the contradiction between spatial detail enhancement and spectral distortion, and cannot effectively avoid spatial artifacts and spectral distortion in the fusion result; at the same time, there is insufficient extraction of image features, mainly focusing on spatial domain features and ignoring spectral information, resulting in the lack of an effective multi-scale feature extraction mechanism, with limited feature expression ability and difficulty in depicting complex image structures; the above solutions also have the problem of low feature fusion efficiency, lacking an adaptive feature selection mechanism, resulting in large information loss during the fusion process, and insufficient model generalization ability, resulting in poor adaptability to images obtained by different sensors. Summary of the Invention

[0010] In view of the above problems, the object of the present invention is to provide a multi-spectral image and panchromatic image fusion method and system based on the full-spectrum space, which fully utilizes spatial-spectral information to reconstruct and obtain high-quality high-resolution multi-spectral images.

[0011] An embodiment of the present invention provides a multi-spectral image and panchromatic image fusion method and system based on the full-spectrum space.

[0012] First aspect: A multi-spectral image and panchromatic image fusion method based on the full-spectrum space, which uses a trained improved two-stream fusion network (TSFNet) to reconstruct high-resolution multi-spectral images. The steps include:

[0013] S1. Obtain a panchromatic image and a multi-spectral image, and perform image preprocessing and normalization operations.

[0014] S2. Process the panchromatic image and the multi-spectral image in both the spatial domain and the frequency domain simultaneously to capture the spatial features and frequency domain features of the panchromatic image and the multi-spectral image respectively.

[0015] S3. Concatenate the spatial features of the panchromatic image and the multi-spectral image to obtain the image spatial features, and concatenate the frequency domain features of the panchromatic image and the multi-spectral image to obtain the image frequency domain features.

[0016] S4. Use a residual fusion module (RFB) to perform residual fusion on the image spatial features and the image frequency domain features.

[0017] S5. Reconstruct a high-resolution multi-spectral image based on the fused features, and after quality evaluation, output a qualified image.

[0018] Further, the image preprocessing and normalization operations in S1 include:

[0019] For the panchromatic image, maintain the original spatial resolution; for the multi-spectral image, perform upsampling through bicubic interpolation to make it have the same spatial resolution as the panchromatic image.

[0020] Further, in S2, the panchromatic image and the multispectral image are processed simultaneously in the spatial domain and the frequency domain, including:

[0021] In the spatial domain, the spatial structure features of the image are extracted through a convolutional network;

[0022] In the frequency domain, first, the input image is subjected to a block discrete cosine DCT transform, and then the frequency structure features of the image are extracted through a convolutional network.

[0023] Further, performing a block discrete cosine DCT transform on the input image in the frequency domain includes: performing a block DCT transform on the input image to obtain DCT transform coefficients, and performing block recombination normalization processing based on the DCT transform coefficients.

[0024] Further, S4 includes the steps of:

[0025] S41. Performing spatial convolution processing on the spatial features of the image to obtain the spatial features after convolution

[0026] S42. Simultaneously performing frequency domain convolution and upsampling processing on the frequency domain features of the image to obtain the frequency domain features after convolution;

[0027] S43. Concatenating the spatial features and the frequency domain features after convolution to obtain concatenated features;

[0028] S44. Inputting the concatenated features into the ESMSA module to output fused features, and performing residual fusion on the fused features and the spatial features of the image.

[0029] Further, inputting the concatenated features into the ESMSA module to output fused features in S44 includes:

[0030] S44a. Processing the concatenated features in multiple attention heads in parallel; each attention head generates three groups of feature vectors of query, key, and value through linear transformation; then each attention head calculates the dot product of the query vector and the key vector, and after scaling and softmax operations, obtains attention weights, and multiplies the attention weights by the value vector to obtain the attention output;

[0031] S44b. Convolving the concatenated features, and after extracting the spatial features, obtaining the spatial feature output;

[0032] S44c. Fusing, normalizing, and outputting the attention output and the spatial feature output to obtain the fused features.

[0033] Further, the training of the improved two-stream fusion network (TSFNet) includes: calculating the mean square error loss between the reconstructed image and the target image, calculating the high-frequency information preservation loss, calculating the peak signal-to-noise ratio loss, and optimizing the two-stream fusion network training after synthesizing the losses.

[0034] Second aspect: A multi-spectral image and panchromatic image fusion system based on the full-spectrum space, comprising:

[0035] A feature extraction module, configured to extract image spatial features and image frequency-domain features from the input panchromatic image and multi-spectral image;

[0036] A feature fusion module, which performs residual fusion on the image spatial features and image frequency-domain features;

[0037] An image reconstruction module, which reconstructs a high-resolution multi-spectral image based on the residual fusion features.

[0038] Third aspect: An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the method provided in the first aspect are implemented.

[0039] Fourth aspect: A non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.

[0040] Advantages of the present invention:

[0041] 1. The present invention utilizes the improved two-stream fusion network (TSFNet) and adopts an encoder-decoder structure design. At the input end, the network receives the panchromatic image and the multi-spectral image respectively; enables the multi-spectral image and the panchromatic image to have the same spatial resolution; processes the input images in both the spatial domain and the frequency domain during the preprocessing stage. This dual-path feature extraction mechanism can comprehensively capture different levels of features of the image; effectively balance the spatial-spectral information, maintain the accuracy of spectral information while improving the spatial resolution, solve the contradiction between spatial detail enhancement and spectral distortion in traditional methods, and can avoid spatial artifacts and spectral distortion in the fusion result; fully extract image features, pay attention to both spatial-domain features and frequency-domain information, and can depict complex image structures; adopt an adaptive feature selection mechanism, with small information loss during the fusion process, high feature fusion efficiency, strong adaptability to images obtained by different sensors, sensitive to noise and imaging condition changes, and can reconstruct and obtain high-quality high-resolution multi-spectral images.

[0042] 2. The ESMSA module of the present invention adopts a multi-head attention mechanism, which divides the input features into multiple attention heads for parallel processing; each attention head can independently focus on different feature patterns, thereby achieving a comprehensive understanding of the input features; in order to improve the effect of the attention mechanism, the present invention also introduces a learnable scaling factor. The scaling factor can adaptively adjust the importance of different attention heads, enabling the network to better adapt to different image contents. In addition, the present invention also designs a relative position encoding mechanism to encode position information into the feature representation, enhancing the model's ability to perceive spatial positions; in terms of feature fusion, the ESMSA module adopts a dual-branch structure. The Epo branch enhances the expression ability of local features through grouped convolution operations, while the Epr branch uses self-attention mechanisms to capture long-range dependencies between features; the outputs of these two branches are fused through adaptive weights, which not only preserves local details but also establishes global semantic connections.

[0043] 3. In terms of frequency-domain feature extraction of the present invention, the input image is subjected to a block DCT transformation, effectively reducing the differences between different images; in the further processing of frequency-domain features, the present invention adopts a Pixel Unshuffle operation to rearrange the DCT coefficients. This rearrangement operation organizes adjacent frequency components together, facilitating the subsequent convolutional network to extract frequency-domain features; the present invention designs a dedicated high-frequency component retention module to highlight important high-frequency information through an attention mechanism, ensuring that detailed features are not lost during the feature extraction process.

[0044] 4. The multi-scale residual learning framework constructed by the present invention, in this framework, each residual block contains two branches for spatial feature processing and frequency-domain feature processing. The spatial feature branch extracts and enhances spatial-domain features through multiple convolutional operations, while the frequency-domain feature branch is responsible for processing and integrating frequency-domain information. The features of the two branches are upsampled through transposed convolutions to ensure that the spatial scales of the feature maps match. During the feature fusion process, an attention module is introduced to adaptively weight the features, highlighting important feature channels.

[0045] 5. The present invention adopts a progressive strategy in feature reconstruction; the network performs feature fusion at different scales and passes low-level detailed information through skip connections. This multi-scale processing method can simultaneously take into account global semantic information and local detailed features. To further improve the reconstruction effect, the present invention introduces multiple global residual connections in the network, which can effectively reduce information loss during feature transmission and also accelerate the convergence speed of the network.

[0046] 6. When the present invention performs network training, the loss function is a multi-component composite loss, comprehensively considering three aspects: image reconstruction quality, spectral fidelity, and spatial details. In terms of the reconstruction loss, not only the mean square error between the prediction result and the target image is calculated, but also a perception-based loss term is introduced. This perception loss extracts features through a pre-trained deep network, compares the differences between the prediction result and the target image in the feature space, and can better maintain the visual quality of the image. In terms of spectral fidelity, the present invention designs a special spectral loss term. This loss term first downsamples the fusion result to the resolution of the original multi-spectral image, and then calculates the spectral vector angle difference between the two. This design ensures the accurate preservation of frequency domain information during the fusion process. At the same time, by calculating the correlation of spectral features, the fidelity of frequency domain information is further constrained. To maintain spatial details, the present invention adds a spatial detail loss term to the loss function. This loss term extracts the edge information of the image through a designed high-pass filter, and then compares the matching degree of the edge features between the fusion result and the panchromatic image. This design effectively ensures that while the fusion result maintains frequency domain information, it can also accurately reconstruct spatial details. The present invention directly introduces the PSNR index into the loss function, sets the target PSNR value as the optimization target, and gradually improves the fusion effect during the training process by dynamically adjusting the weights of each loss term. This PSNR-based optimization strategy can more directly guide the network to generate high-quality fusion results. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the structural flowchart of the system of the present invention;

[0048] Figure 2 is the schematic flowchart of the method of the present invention;

[0049] Figure 3 is the structural schematic diagram of the system of the present invention;

[0050] Figure 4 is the schematic flowchart of the feature fusion process of the present invention;

[0051] Figure 5 is the schematic flowchart of the ESMSA module process of the present invention;

[0052] Figure 6 is the structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0054] Current remote sensing image fusion technologies still have the following problems in practical applications: Image fusion methods mainly focus on spatial domain feature extraction and insufficient utilization of spectral information; deep learning fusion models lack effective feature selection mechanisms and cannot perform adaptive fusion according to the features of different image regions; and existing fusion model algorithms are inefficient in processing high-resolution and large-size images, with problems of training instability.

[0055] To address the above problems, the present invention provides a method for fusing multispectral images and panchromatic images based on the full-spectrum space, using a trained two-stream fusion network (TSFNet) to reconstruct high-resolution multispectral images. Figure 2 The flowchart of the fusion method of the present invention is shown as follows. The steps of the method include:

[0056] S1. Obtain the panchromatic image and the multispectral image, and perform image preprocessing and normalization operations.

[0057] As Figure 1 shown, the present invention is based on an improved two-stream fusion network (TSFNet), and TSFNet adopts an encoder-decoder structure design. At the input end, the TSFNet network receives the panchromatic image and the multispectral image respectively. For the panchromatic image, its original spatial resolution is maintained as the information source providing high spatial frequency details; for the multispectral image, it is first upsampled by bicubic interpolation to have the same spatial resolution as the panchromatic image, laying a foundation for subsequent feature extraction and fusion.

[0058] Specifically, first initialize the system parameters and operating environment, configure the GPU device and memory resources, load the pre-trained model weights, and set the processing parameters (such as batch size, image size, etc.) to obtain the panchromatic (PAN) image and the multispectral (MS) image.

[0059] The PAN image is a high-spatial-resolution panchromatic image in 3-channel RGB format, and the MS image is a low-spatial-resolution multispectral image also in 3-channel RGB format. The image formats of the PAN image and the MS image support common formats such as PNG and JPEG.

[0060] The PIL library can be used to read the PAN image and the MS image, convert them to RGB mode, record the original resolution of the PAN image (usually 256×256), and record the original resolution of the MS image (usually 64×64).

[0061] Preprocess and normalize the image: The preprocessing includes image scaling and data type conversion. Image scaling means using the bicubic interpolation algorithm to upsample the MS image to the resolution of the PAN image. Data type conversion means converting the image data to floating point type and normalizing the range to [0, 1]. Specifically, transforms.Resize can be used for image scaling and transforms.ToTensor can be used for data conversion.

[0062] Normalization refers to standardizing the image, including normalizing the image mean by subtracting the channel mean, normalizing the standard deviation by dividing the channel standard deviation, and adjusting the image value range to [-1, 1]. Specifically, transforms.Normalize can be used, with the mean set to [0.5, 0.5, 0.5] and the standard deviation set to [0.5, 0.5, 0.5].

[0063] S2. At the same time, process the panchromatic image and the multispectral image in the spatial domain and the frequency domain to capture the respective spatial features and frequency domain features of the panchromatic image and the multispectral image.

[0064] As Figure 3 shown, parallelly process the two feature domains of the spatial domain and the frequency domain of the panchromatic image and the multispectral image. In the spatial domain, extract the spatial structure features of the image through a convolutional network. In the frequency domain, first perform a block discrete cosine transform (DCT) on the input image, and then extract the image frequency structure features through a convolutional network. Extract complementary features through different processing branches to maintain the integrity and independence of the features.

[0065] Specifically, the extraction of spatial domain features includes using the ConvPReLUConv module for convolution. The input channels can be 3, and the output channels can be 32 (PAN) / 96 (MS). Use BatchNorm2d for feature normalization, use PReLU as the activation function, and perform residual connection to maintain the gradient flow.

[0066] The specific implementation code can be:

[0067] class ConvPReLUConv(nn.Module):

[0068] def __init__(self, in_channels, out_channels):

[0069] self.conv1 = nn.Conv2d(in_channels, out_channels, kernel_size = 3, padding = 1)

[0070] self.bn1 = nn.BatchNorm2d(out_channels)

[0071] self.conv2 = nn.Conv2d(out_channels, out_channels, kernel_size = 3, padding = 1)

[0072] self.bn2 = nn.BatchNorm2d(out_channels)

[0073] self.prelu = nn.PReLU()

[0074] The frequency-domain feature extraction includes performing a block discrete cosine transform (DCT) on the input image in the frequency domain. First, perform a block DCT transform on the input image to obtain DCT transform coefficients, and then perform block recombination normalization based on the DCT transform coefficients to perform frequency-domain feature extraction, etc.

[0075] The block DCT transform refers to dividing the image into blocks (for example, 4×4 in size), and then performing an orthogonal DCT transform on each block to maintain the integrity of the frequency-domain information. The specific implementation can use def apply_dct(self, img, block_size = 4) to perform the block DCT transform on each channel separately.

[0076] The block recombination uses the Pixel Unshuffle operation to expand the channels as: C -> C*(block_size^2), maintaining spatial correlation and facilitating subsequent convolution processing. The specific implementation can use def pixel_unshuffle(self, img, block_size) to recombine the feature map, reduce the spatial resolution, and increase the number of channels.

[0077] For frequency-domain feature extraction, the PAN image can use a 48-channel output, and the MS image can use a 192-channel output. Use BatchNorm for normalization and use non-linear activation to enhance feature expression.

[0078] S3. Concatenate the spatial features of the panchromatic image and the multispectral image to obtain the image spatial features, and concatenate the frequency-domain features of the panchromatic image and the multispectral image to obtain the image frequency-domain features.

[0079] Concatenate the spatial features and the frequency-domain features through the channel attention mechanism, adopt adaptive weight learning to evaluate the feature importance, use torch.cat for feature concatenation, and adjust the feature weights through the attention mechanism to obtain the concatenated image spatial features and image frequency-domain features.

[0080] S4. Use the Residual Fusion Block (RFB) to perform residual fusion on the spatial features and frequency domain features of the image.

[0081] As Figure 4 shown, the steps for the Residual Fusion Block (RFB) to perform residual fusion on the spatial features and frequency domain features of the image include:

[0082] S41. Perform spatial convolution processing on the spatial features of the image to obtain the spatially convolved features

[0083] S42. At the same time, perform frequency domain convolution and upsampling processing on the frequency domain features of the image to obtain the frequency domain features after convolution;

[0084] S43. Concatenate the spatially convolved features and the frequency domain features to obtain the concatenated features;

[0085] S44. Input the concatenated features into the ESMSA module to output the fused features, and perform residual fusion on the fused features and the spatial features of the image.

[0086] Specifically, the Residual Fusion Block (RFB) mainly includes a spatial branch processing, a frequency domain branch processing, an ESMSA module, and a residual connection module, which can maintain the original feature information and improve the gradient flow through the above structure. The specific implementation can use class ResidualFusionBlock(nn.Module): def__init__(self,Cs,Cf); to implement the residual fusion block for spatial convolution and frequency domain processing.

[0087] As Figure 5 shown, the ESMSA module uses the multi-head self-attention mechanism for local-global feature interaction. The ESMSA module adopts a dual-branch structure, including the Epo (Enhanced Position-wise Operation) branch and the Epr (Enhanced Pairwise Relation) branch. The Epo branch enhances the expression ability of local features through grouped convolution operations, while the Epr branch uses the self-attention mechanism to capture the long-range dependence relationship between features. The outputs of these two branches are fused through adaptive weights, which not only preserves the local details but also establishes global semantic connections. The steps include:

[0088] S44a. Process the splicing features in parallel with multiple attention heads; each attention head generates three groups of feature vectors, namely query, key, and value, through linear transformation; then each attention head calculates the dot product of the query vector and the key vector, and after scaling and softmax operations, obtains the attention weights. After multiplying the attention weights by the value vectors, the attention output is obtained. The specific implementation can adopt class ESMSA(nn.Module), def__init__(self,C,Cs,h=8,b=16); to implement enhanced spatial multi-head self-attention.

[0089] S44b. Convolve the splicing features. After extracting the spatial features, obtain the spatial feature output.

[0090] S44c. Fuse the attention output and the spatial feature output, and after normalization processing, output the fused features.

[0091] S5. Reconstruct the high-resolution multi-spectral image based on the fused features. After quality evaluation, output the qualified image.

[0092] The specific implementation can adopt:

[0093] self.output_conv = nn.Sequential(

[0094] nn.Conv2d(128, 64, kernel_size = 3, padding = 1),

[0095] nn.BatchNorm2d(64),

[0096] nn.PReLU(),

[0097] nn.Conv2d(64, ms_channels, kernel_size = 3, padding = 1),

[0098] nn.Tanh()

[0099] Perform channel dimensionality reduction and feature reconstruction, and finally output using the PReLU and Tanh functions.

[0100] Training the two-stream fusion network (TSFNet) includes: calculating the mean square error loss between the reconstructed image and the target image, calculating the high-frequency information preservation loss, calculating the peak signal-to-noise ratio loss, and optimizing the two-stream fusion network training after comprehensive loss.

[0101] In terms of spectral preservation, a dedicated spectral loss term is adopted. This loss term first downsamples the fusion result to the resolution of the original multi-spectral image, and then calculates the spectral vector angular difference between the two. This ensures the accurate preservation of frequency-domain information during the fusion process. At the same time, by calculating the correlation of spectral features, the fidelity of frequency-domain information is further constrained.

[0102] To preserve spatial details, a spatial detail loss term is added to the loss function. This loss term extracts the edge information of the image through a designed high-pass filter, and then compares the matching degree of the edge features between the fusion result and the panchromatic image. This design effectively ensures that while the fusion result preserves spectral information, it can also accurately reconstruct spatial details.

[0103] The PSNR metric is directly introduced into the loss function, and the target PSNR value is set as the optimization goal. By dynamically adjusting the weights of each loss term, the fusion effect is gradually improved during the training process. Using PSNR calculation for structural similarity evaluation, a PSNR threshold of 47.0 can be set for loss function threshold checking, and the ImprovedLoss class is used for evaluation. Based on comprehensive judgment of multiple metrics, this PSNR-based optimization strategy can more directly guide the network to generate high-quality fusion results.

[0104] As Figure 3 shown, based on the above method, the present invention also discloses an image fusion system, including:

[0105] A feature extraction module, which includes a spatial feature extraction unit and a frequency-domain feature extraction unit. The spatial feature extraction unit contains a multi-layer convolutional neural network and a batch normalization layer; the frequency-domain feature extraction unit contains a discrete cosine transform layer and a convolutional layer; the feature extraction module is used to extract the image spatial features and image frequency-domain features of the input panchromatic image and multi-spectral image;

[0106] A feature fusion module, including a spatial processing unit, a frequency-domain processing unit, an ESMSA unit, a residual connection unit, etc. Among them, the spatial processing unit is used to process spatial features, the frequency-domain processing unit is used to process frequency-domain features, the ESMSA unit is used for feature fusion, and the residual connection unit is used for information transmission; the feature fusion module is used to perform residual fusion on the image spatial features and image frequency-domain features.

[0107] Among them, the ESMSA unit includes a matrix construction unit, a weight calculation unit, a feature weighting unit, etc.; the matrix construction unit is used to construct a query matrix, a key matrix, and a value matrix; the weight calculation unit is used to calculate attention weights; the feature weighting unit is used to perform feature weighted fusion.

[0108] An image reconstruction module, which reconstructs a high-resolution multi-spectral image based on the residual fusion features.

[0109] The improved Two-Stream Fusion Network (TSFNet) is trained by using a loss calculation unit and a network optimization unit. The loss calculation unit is used to calculate the reconstruction loss, the high-frequency information preservation loss, and the peak signal-to-noise ratio loss. The network optimization unit is used to optimize the network parameters based on the loss function.

[0110] The present invention uses a feature extraction module to obtain the panchromatic image data and the multispectral image data of the input layer, performs convolutional processing on the panchromatic image data to obtain the panchromatic image spatial features, and performs convolutional layer processing on the multispectral image data to obtain the multispectral image spatial features. The discrete cosine transform is performed on the panchromatic image data and the multispectral image data to obtain the frequency domain data, and convolutional processing is performed to obtain the panchromatic image frequency domain features and the multispectral image frequency domain features. The panchromatic image spatial features and the multispectral image spatial features are concatenated to obtain the image spatial features and the image frequency domain features. Through the feature fusion module, a query matrix, a key matrix, and a value matrix are constructed, the attention weights are calculated based on the query matrix and the key matrix, and weighted with the value matrix to obtain the attention output. The spatial feature output and the attention output are feature-fused, and the fused features are used for residual fusion with the image spatial features. The image reconstruction module is used to reconstruct the high-resolution multispectral image based on the fused features.

[0111] Specific implementation examples are described as follows:

[0112] The improved TSFNet network structure is adopted. The network structure is as Figure 1 shown. The network structure of this embodiment mainly includes the following core components:

[0113] The network structure of the input layer receives four inputs: Panchromatic image (PAN): 3 channels, with a size of 256×256; Multispectral image (MS): 3 channels, with an original size of 64×64, upsampled to 256×256; The DCT transformation result of the PAN image; The DCT transformation result of the MS image;

[0114] Then, a feature extraction module is used for feature extraction. The feature extraction module includes the following key component codes:

[0115] python

[0116] Copy

[0117] self.pan_input_conv = ConvPReLUConv(pan_channels = 3, out_channels = 32)

[0118] self.ms_input_conv = ConvPReLUConv(ms_channels = 3, out_channels = 96)

[0119] self.pan_dct_conv = ConvPReLUConv(pan_channels * 4, out_channels = 48)

[0120] self.ms_dct_conv = ConvPReLUConv(ms_channels * 4, out_channels = 192)

[0121] Among them, the specific implementation of the ConvPReLUConv module is as follows:

[0122] python

[0123] Copy

[0124] class ConvPReLUConv(nn.Module):

[0125] def __init__(self, in_channels, out_channels):

[0126] super(ConvPReLUConv, self).__init__()

[0127] self.conv1 = nn.Conv2d(in_channels, out_channels, kernel_size = 3, padding = 1)

[0128] self.bn1 = nn.BatchNorm2d(out_channels)

[0129] self.conv2 = nn.Conv2d(out_channels, out_channels, kernel_size = 3, padding = 1)

[0130] self.bn2 = nn.BatchNorm2d(out_channels)

[0131] self.prelu = nn.PReLU()

[0132] Then, the feature fusion module is used for feature fusion. The feature fusion module adopts an improved attention mechanism ESMSA (Enhanced Spatial-Spectral Multi-head Self-Attention), and its core implementation includes:

[0133] Multi-head self-attention calculation:

[0134] python

[0135] Copy

[0136] Q = Q.view(B, -1, self.h, C / / self.h).permute(0, 2, 3, 1)

[0137] K = K.view(B, -1, self.h, C / / self.h).permute(0, 2, 3, 1)

[0138] V = V.view(B, -1, self.h, C / / self.h).permute(0, 2, 1, 3)

[0139] attn = torch.matmul(Q, K.transpose(-2, -1))

[0140] attn = attn * self.sigma.view(1, -1, 1, 1) / ((C / / self.h) ** 0.5 + 1e - 8)

[0141] attn = torch.softmax(attn, dim = -1)

[0142] Residual Fusion Block:

[0143] python

[0144] Copy

[0145] class ResidualFusionBlock(nn.Module):

[0146] def __init__(self, Cs, Cf):

[0147] self.spatial_conv = ConvPReLUConv(Cs, Cs)

[0148] self.freq_conv = ConvPReLUConv(Cf, Cf)

[0149] self.t_conv = nn.Sequential(nn.ConvTranspose2d(Cf, Cf, kernel_size = 4, stride = 2, padding = 1), nn.PReLU())

[0150] self.e_smsa = ESMSA(Cs+Cf, Cs, h=8, b=16)

[0151] The output reconstruction module adopts a progressive reconstruction strategy, including:

[0152] python

[0153] Copy

[0154] self.output_conv = nn.Sequential(nn.Conv2d(128, 64, kernel_size=3, padding=1),

[0155] nn.BatchNorm2d(64), nn.PReLU(), nn.Conv2d(64, ms_channels, kernel_size=3, padding=1), nn.Tanh())

[0156] The training strategy for the network includes: designing a loss function. In this embodiment, an improved multi-task loss function is adopted:

[0157] python

[0158] Copy

[0159] class ImprovedLoss(nn.Module):

[0160] def __init__(self, alpha=5.0, beta=0.0, gamma=0.0, psnr_weight=0.5, target_psnr=47.0):

[0161] super(ImprovedLoss, self).__init__()

[0162] self.mse = nn.MSELoss()

[0163] self.alpha = alpha

[0164] self.beta = beta

[0165] self.gamma = gamma

[0166] self.psnr_weight = psnr_weight

[0167] self.target_psnr = target_psnr

[0168] The loss function consists of four parts:

[0169] Reconstruction loss (L_hrms): Measures the pixel-level difference between the generated image and the target image

[0170] Multi-spectral consistency loss (L_ms): Ensures the preservation of frequency domain information

[0171] Spatial structure loss (L_spatial): Preserves spatial detail information

[0172] PSNR-guided loss: Optimizes the overall reconstruction quality

[0173] An optimization strategy is adopted for the network. The AdamW optimizer is used, and warmup and stepwise learning rate adjustment strategies are designed:

[0174] python

[0175] Copy

[0176] optimizer = AdamW(model.parameters(), lr = 2e-4, weight_decay = 0.05, betas = (0.9, 0.999), eps = 1e-8)

[0177] scheduler = StepLRScheduler(optimizer, init_lr = 2e-4, milestones = [50, 100, 150, 200, 250, 300], gamma = 0.5, warmup_epochs = 5, warmup_start_lr = 1e-6)

[0178] The main parameter settings for the training process are as follows: batch size: 8; number of training epochs: 2000; gradient accumulation steps: 4; validation interval: 5 epochs; early stopping mechanism: stop if there is no improvement in 20 epochs; checkpoint saving interval: 20 epochs.

[0179] The present invention has carefully designed and optimized the network structure and training parameters. In the setting of the number of feature channels, the spatial features gradually increase from the initial 32 channels to 128 channels, while the frequency domain features increase from 48 channels to 240 channels. This progressive channel expansion strategy not only ensures the sufficiency of feature extraction but also avoids excessive consumption of computing resources.

[0180] In terms of the training strategy, the present invention adopts mixed-precision training to improve the training efficiency. The batch size is set to 8, and the gradient accumulation technique is used to update the parameters once every 4 steps. This setting ensures the training stability while also considering the limitations of hardware resources. The learning rate adopts the OneCycleLR strategy, with the initial value set to 4e-4, and it is dynamically adjusted according to the cosine annealing rule during the training process. This learning rate scheduling scheme can effectively improve the convergence speed and performance of the model.

[0181] In terms of the parameter configuration of the loss function, the reconstruction loss weight is set to 1.0 as the main optimization objective; the weights of the spectral loss and the spatial loss are set to 0.1 and 0.01 respectively to balance the reconstruction effects in different aspects; the weight of the PSNR loss is set to 0.1, and the target PSNR value is set to 40.0. The settings of these parameters have been verified through a large number of experiments and can achieve good fusion effects.

[0182] The present invention utilizes an improved two-stream fusion network (TSFNet) with an encoder-decoder structure design. At the input end, the network receives the panchromatic image and the multispectral image respectively; for the panchromatic image, its original spatial resolution is maintained as the information source providing high-spatial-frequency details; for the multispectral image, it is upsampled through bicubic interpolation to have the same spatial resolution as the panchromatic image. In the preprocessing stage, the present invention processes the input images in both the spatial domain and the frequency domain. In the spatial domain, the spatial structure features of the images are directly extracted through a multi-layer convolutional network; in the frequency domain, first, the input images are subjected to block discrete cosine transform (DCT), and then the frequency-domain information is obtained through a specially designed frequency-domain feature extraction network. This dual-path feature extraction mechanism can comprehensively capture different-level features of the images and provide a sufficient information basis for subsequent high-quality fusion.

[0183] The present invention also provides an electronic device Figure 6 which is the structural schematic diagram of the electronic device provided by the embodiment of the present invention, as Figure 6 shown. The electronic device may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communications interface, and the memory complete the communication with each other through the communication bus. The processor can call the logical instructions in the memory, for example, to execute the following method:

[0184] S1. Obtain the panchromatic image and the multispectral image, and perform image preprocessing and normalization operations;

[0185] S2. Process the panchromatic image and the multispectral image in both the spatial domain and the frequency domain simultaneously to capture the spatial features and frequency-domain features of the panchromatic image and the multispectral image respectively;

[0186] S3. Stitch the spatial features of the panchromatic image and the multispectral image to obtain the image spatial features, and stitch the frequency domain features of the panchromatic image and the multispectral image to obtain the image frequency domain features;

[0187] S4. Use the Residual Fusion Block (RFB) to perform residual fusion on the image spatial features and the image frequency domain features;

[0188] S5. Reconstruct the high-resolution multispectral image based on the fused features, and after quality assessment, output the qualified image.

[0189] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0190] The embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments, for example, including:

[0191] S1. Obtain the panchromatic image and the multispectral image, and perform image preprocessing and normalization operations;

[0192] S2. Process the panchromatic image and the multispectral image simultaneously in the spatial domain and the frequency domain, and capture the respective spatial features and frequency domain features of the panchromatic image and the multispectral image;

[0193] S3. Stitch the spatial features of the panchromatic image and the multispectral image to obtain the image spatial features, and stitch the frequency domain features of the panchromatic image and the multispectral image to obtain the image frequency domain features;

[0194] S4. Use the Residual Fusion Block (RFB) to perform residual fusion on the image spatial features and the image frequency domain features;

[0195] S5. Reconstruct the high-resolution multispectral image based on the fused features, and after quality assessment, output the qualified image.

[0196] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0197] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multispectral image and panchromatic image fusion method based on attention mechanism, characterized in that: The trained improved two-stream fusion network (TSFNet) is used to reconstruct high-resolution multispectral images. The steps include: S1, obtaining panchromatic images and multispectral images, and performing image preprocessing and normalization operations; S2, processing the panchromatic image and the multispectral image in the spatial domain and the frequency domain simultaneously, capturing the spatial features and frequency domain features of the panchromatic image and the multispectral image respectively; S3, splicing the spatial features of the panchromatic image and the multispectral image to obtain the image spatial features, and splicing the frequency domain features of the panchromatic image and the multispectral image to obtain the image frequency domain features; S4, using the residual fusion module RFB to perform residual fusion on the image spatial features and image frequency domain features; S5. Reconstruct a high-resolution multispectral image based on the fused features, perform quality assessment, and output a qualified image.

2. The fusion method according to claim 1, characterized in that: The image preprocessing and normalization operations performed in S1 include: For panchromatic images, the original spatial resolution is maintained; for multispectral images, bicubic interpolation is used to upsample them to have the same spatial resolution as the panchromatic image.

3. The fusion method according to claim 1, characterized in that: The S2 processes the panchromatic image and the multispectral image in both the spatial domain and the frequency domain, including: In the spatial domain, the spatial structural features of the image are extracted through the convolutional network; In the frequency domain, the input image is first subjected to block discrete cosine DCT transform, and then the image frequency structure features are extracted through a convolutional network.

4. The fusion method according to claim 3, characterized in that: The input image is subjected to a block discrete cosine DCT transform in the frequency domain, including: performing a block DCT transform on the input image, obtaining DCT transform coefficients, and performing block reorganization and normalization processing according to the DCT transform coefficients.

5. The fusion method according to claim 1, characterized in that: The S4 comprises the steps of: S41, perform spatial convolution processing on the image spatial features to obtain the spatial features after convolution S42, performing frequency domain convolution and upsampling processing on the frequency domain features of the image at the same time to obtain the frequency domain features after convolution; S43, concatenating the spatial features and frequency domain features after the convolution to obtain concatenated features; S44, input the splicing features into the ESMSA module to output the fusion features, and perform residual fusion of the fusion features with the image space features.

6. The fusion method according to claim 5, characterized in that: In S44, the splicing features are input into the ESMSA module to output the fusion features, including: S44a, the concatenated features are divided into multiple attention heads for parallel processing; each attention head generates three sets of feature vectors, namely query, key and value, through linear transformation; then each attention head calculates the dot product of the query vector and the key vector, and obtains the attention weight after scaling and softmax operation, and obtains the attention output after multiplying the attention weight with the value vector; S44b, convolving the concatenated features, extracting the spatial features, and obtaining the spatial feature output; S44c, fuse the attention output and the spatial feature output, normalize them, and then output the fused feature.

7. The fusion method according to claim 1, characterized in that: The improved two-stream fusion network (TSFNet) is trained, including: calculating the mean square error loss between the reconstructed image and the target image, calculating the high-frequency information retention loss, calculating the peak signal-to-noise ratio loss, and optimizing the two-stream fusion network training after comprehensive loss.

8. An image fusion system based on the method according to any one of claims 1 to 7, characterized in that: include: A feature extraction module is used to extract image spatial features and image frequency domain features from the input full-color image and multi-spectral image; Feature fusion module, which performs residual fusion on image spatial features and image frequency domain features; Image reconstruction module reconstructs high-resolution multispectral images based on residual fusion features.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Panchromatic multispectral image fusion method and device

    CN114331930A

  • High-resolution multispectral image reconstruction method based on dynamic edge-guided network

    CN118505509B

  • Hyperspectral and panchromatic image fusion method based on CNN and Laplace pyramid

    CN112669248A

  • Double-flow remote sensing image fusion method based on residual channel attention mechanism

    CN113920043A

  • Remote sensing image space-spectrum fusion method and device, electronic equipment and storage medium

    CN117079105A

Cited By

  • Remote sensing image panchromatic sharpening method and terminal

    CN120355625A

  • Panchromatic sharpening image fusion method and device based on double-domain flexible converter

    CN120807318A

  • Multi-spectral image generation method and system based on W-Transform and medium

    CN120997060A

  • Double-flow space-spectrum fusion method and device, electronic equipment and storage medium

    CN121482545A