Multi-angle panchromatic and multispectral image progressive fusion method
By adopting the progressive fusion method of multi-angle full-color and multi-spectral images in image fusion, and using multi-scale feature extraction and multi-stream feature fusion modules, the problems of insufficient high-frequency information capture and limited multi-scale feature processing capabilities are solved, and high-quality image fusion is achieved.
Patent Information
- Application Number
- CN202510044121.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has problems in the fusion of full-color and multi-spectral images, insufficient high-frequency information capture, limited multi-scale feature processing capabilities, and spatial and spectral distortion, which affects the overall quality of the fusion image.
A multi-angle full-color and multi-spectral image synthesis method is designed, and a multi-scale feature extraction module, an enhanced fusion flow module and a multi-stream feature fusion module are adopted. Through the frequency domain interaction module, the spatial domain interaction module and the dual-domain feature fusion module, the complementary frequency domain and the spatial domain characteristics and the optimization fusion of deep information are achieved.
The spatial details and spectral fidelity of the fusion image are improved, and the shortcomings of existing methods are solved when retaining high-frequency and global information at the same time are solved, thereby achieving higher quality remote sensing image fusion.
Smart Images

Figure CN119942285A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a method for progressive fusion of multi-angle panchromatic and multi-spectral images. Background Art
[0002] The fusion of panchromatic and multispectral images is called panchromatic sharpening, which is a technique that uses high-resolution panchromatic images to improve the spatial resolution of low-resolution multispectral images. It aims to generate images with both high spatial resolution and rich spectral information. Mainstream methods include component replacement, multiresolution analysis, and optimization-based methods, but they usually cause spectral distortion or make it difficult to retain high-frequency and global information at the same time. In recent years, deep learning-based technologies have achieved better fusion through convolutional neural networks, such as PNN, PANNET and other models. Although the performance has been improved, they still rely mainly on spatial domain information and are difficult to effectively capture high-frequency details, resulting in insufficient texture in the generated image. Overall, existing methods still have limitations in fusing high-frequency and global information.
[0003] The Chinese patent publication number is "CN114140359B", and the name is "A remote sensing image fusion sharpening method based on progressive cross-scale neural network". This method first downsamples the high-resolution multispectral image and constructs an input data pair with the panchromatic image; then, the panchromatic image is decomposed into image features of different scales through Gaussian pyramid decomposition, which is combined with the downsampled multispectral image as the input of each pyramid sub-network; then, each sub-network extracts and fuses multi-scale features through feature fusion module and cross-scale attention module, and gradually performs upsampling processing; in the reconstruction module of the sub-network, detail information is retained through jump connection, and the fusion feature map of each layer is output; finally, the top-level pyramid network outputs the fused high-resolution multispectral image. This method has deficiencies in insufficient high-frequency information capture, limited multi-scale feature processing capabilities, spectral fidelity and detail sharpening effect, which affects the overall quality of the fused image. Summary of the invention
[0004] In view of the defects existing in the prior art, the present invention provides a method for progressive fusion of multi-angle panchromatic and multispectral images to solve the problems of insufficient capture of high-frequency information, limited multi-scale feature processing capabilities, spatial and spectral distortion, etc. in the prior art.
[0005] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:
[0006] A multi-angle panchromatic and multispectral image progressive fusion method comprises the following steps:
[0007] S1, prepare down-resolution dataset and full-resolution dataset: use several pre-processed remote sensing images to construct down-resolution experimental dataset and full-resolution experimental dataset respectively, and divide the down-resolution experimental dataset into training set, validation set and test set. The full-resolution experimental dataset is the test set;
[0008] S2, constructing a multi-angle panchromatic and multi-spectral image progressive fusion network model: The multi-angle panchromatic and multi-spectral image progressive fusion network model consists of a multi-scale feature extraction module, an enhanced fusion stream module, and a multi-stream feature fusion module;
[0009] S3, determine the loss function and determine the optimal evaluation index of this method: the loss function includes spatial loss, frequency domain loss and mutual information loss. Set the training loss threshold and iteratively optimize the model until the predetermined number of training times is reached or the loss function converges, and save the model parameters. Select the quantitative indicators of commonly used simulation data and the quantitative indicators of real data to evaluate the fusion performance and accuracy of the model;
[0010] S4, training network model: training a multi-angle panchromatic and multispectral image progressive fusion network model, according to the training set prepared in step S1, inputting the panchromatic image and multispectral image in the data set into the fusion network model constructed in step S2 for training, and obtaining training weights;
[0011] S5, save the model: After each round of training in step S3, the validation set is input into the model for verification. If the loss value is less than the set threshold, the best performing model is saved. In the actual multimodal image fusion task, the multimodal image can be directly input into the trained end-to-end network model to generate the final fused image;
[0012] Furthermore, S1 uses the data from the Gaofen-2 satellite and constructs a down-resolution experimental data set and a full-resolution experimental data set based on multiple pre-processed original remote sensing images. The down-resolution experimental data set is constructed according to the Wald protocol, including a training set, a validation set, and a test set. Each data set contains 3 types of data: multispectral images, reference images, and panchromatic images. The full-resolution experimental data set includes panchromatic images and multispectral images.
[0013] Furthermore, the multi-scale feature extraction module in S2 is composed of six instance residual blocks and two downsampling layers. The downsampling module is used to obtain remote sensing images of different scales; instance residual block 1, instance residual block 2, instance residual block 3, instance residual block 4, instance residual block 5, and instance residual block 6 are used to extract basic information of images of different scales to ensure the accuracy of fusion of features of different scales;
[0014] Furthermore, the enhanced fusion flow module in S2 is composed of multiple 3×3 convolutional layers and R-type activation functions, which are used to preliminarily fuse shallow features of the image to further improve the quality of the fused image in the future;
[0015] Furthermore, the multi-stream feature fusion module in S2 includes a frequency domain interaction module, a spatial domain interaction module and a dual-domain feature fusion module. The frequency domain interaction module consists of three parts: amplitude fusion, phase fusion and information integration. It extracts global frequency domain information through multimodal interactive Fourier transform to supplement the high-frequency information of the panchromatic image for the low-resolution multispectral image. This module uses frequency domain perception attention to fuse frequency domain features, capture the overall structure and contextual information of the image, and thus improve image quality and information representation;
[0016] Furthermore, the spatial domain interaction module in the multi-stream feature fusion module of S2 is composed of three serially connected residual depth separable convolution blocks and channel feature optimization modules. Through detail-enhanced wavelet convolution, point-by-point convolution and 3×3 convolution, image details are effectively captured and fused to ensure spatial clarity. This module combines the depth separable convolution and channel feature optimization modules to enhance subtle feature extraction and spatial feature expression, and realizes efficient fusion of frequency domain and spatial domain, improving the robustness and quality of multimodal image fusion;
[0017] Furthermore, the dual-domain feature fusion module in S2 is composed of three fusion reversible neural blocks, a channel feature optimization module and a 1×1 convolution. The fusion reversible neural block integrates frequency domain and spatial domain features, and the channel feature optimization module dynamically adjusts feature weights to strengthen key features. The two work together to achieve more accurate remote sensing image fusion and improve the overall quality and detail performance of the fused image;
[0018] Furthermore, in step S4, the loss function of the entire fusion network training process is composed of spatial loss, frequency domain loss and mutual information loss, and the integrity of the fusion image output by the network is dynamically adjusted by minimizing the loss function;
[0019] Compared with the prior art, the present invention provides a multi-angle panchromatic and multispectral image progressive fusion method, which has the following beneficial effects:
[0020] The present invention provides a progressive fusion method for multi-angle panchromatic and multispectral images. The method first designs a frequency domain interaction module and a space domain interaction module, which respectively generate rich two-domain fusion features based on the unique characteristics of the frequency domain and the space domain, and effectively realize the complementarity of the frequency domain and space domain characteristics at different scales; then a dual-domain feature fusion module is designed, and the deep information fusion of the panchromatic image and the multispectral image under multiple angles is optimized by further feature interaction and fusion of the dual-domain features. The overall progressive design improves the spatial details and spectral fidelity of the fused image, and solves the shortcomings of the existing methods in retaining both at the same time, thereby achieving higher quality remote sensing image fusion.
[0021] The present invention designs a frequency-domain-aware attention block, which replaces the traditional amplitude splicing method with adaptive weighted fusion, generates frequency-domain features with weight adjustment according to the importance of features, highlights key frequency information and suppresses irrelevant components, thereby improving the overall coordination of the image while retaining details. This design effectively avoids information redundancy and enhances feature expression capabilities.
[0022] The present invention designs a fusion reversible neural block to ensure that the information is fully preserved during the feature conversion process. Compared with ordinary convolutional networks, this network has higher flexibility and lower information loss on the feature layer plane, effectively improving the expression ability of full-color and multi-spectral features. This module focuses on the effective interaction of frequency domain and spatial domain features, enhancing the ability to integrate local and global information, thereby improving the spectral consistency and fusion quality of the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flow chart of a progressive fusion method of multi-angle panchromatic and multispectral images;
[0024] Figure 2 This is a network structure diagram of a multi-angle panchromatic and multispectral image progressive fusion method;
[0025] Figure 3 This is a structural diagram of the multi-scale feature extraction module of the present invention;
[0026] Figure 4 This is a residual block structure diagram for feature extraction of the present invention;
[0027] Figure 5 This is the structural diagram of the enhanced fusion flow module of the present invention;
[0028] Figure 6 This is a structural diagram of the frequency domain interaction module of the present invention;
[0029] Figure 7 This is a frequency domain perception attention structure diagram of the present invention;
[0030] Figure 8This is a structural diagram of the spatial domain interaction module of the present invention;
[0031] Fig. 9 This is a structural diagram of the dual-domain feature fusion module of the present invention;
[0032] Fig.10 This is a structural diagram of the fused reversible neural block of the present invention;
[0033] Fig.11 Schematic diagram of evaluation indicators of the method proposed in the present invention. DETAILED DESCRIPTION
[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] Embodiment 1:
[0036] like Figure 1 As shown, an embodiment of the present invention provides a multi-angle panchromatic and multispectral image progressive fusion method, which specifically includes the following steps:
[0037] Step S1, prepare down-resolution data set and full-resolution data set: This method uses the data of Gaofen-2 satellite, based on multiple pre-processed original remote sensing images, to construct down-resolution experimental data set and full-resolution experimental data set respectively, and divide the down-resolution experimental data set into training set, validation set and test set. The full-resolution experimental data set is the test set.
[0038] Step S2, construct a network model for progressive fusion of multi-angle panchromatic and multi-spectral images: the network consists of a multi-scale feature extraction module, an enhanced fusion stream module and a multi-stream feature fusion module, realizing progressive learning fusion based on the dual domains of space and frequency. The multi-scale feature extraction module is responsible for transmitting semantic information of different scales to the multi-stream feature fusion module to realize cross-domain feature fusion and information interaction. The enhanced fusion stream module consists of a convolutional layer and an activation function, which is used to initially fuse shallow features. The feature extraction module contains six instance residual blocks and two downsampling layers, aiming to capture structural and semantic information at different resolutions and provide support for multi-stream feature fusion. The multi-stream feature fusion module includes a frequency domain interaction module, a spatial domain interaction module and a dual-domain feature fusion module. The frequency domain interaction module extracts and completes global frequency domain information through multimodal interactive Fourier transform, and after interacting with amplitude and phase information, generates a fused frequency domain feature map through inverse Fourier transform. The spatial domain interaction module consists of a residual depth separable convolution block and a channel feature optimization module, which is responsible for extracting details and low-frequency features and generating a spatial domain feature map. Finally, the frequency domain and spatial domain feature maps are concatenated and input into the dual-domain feature fusion module. The features are enhanced by multiple fusion of reversible neural network blocks and fused with the features of the enhanced fusion stream module to generate a high-resolution multispectral image.
[0039] Step S3, determine the loss function and evaluation index to optimize the model performance. The selected loss functions include spatial loss, frequency domain loss and mutual information loss. Set the loss threshold and iteratively optimize the model until the number of training times or loss converges. Select the quantitative index of commonly used simulation data and the quantitative index of real data to evaluate the fusion performance and accuracy of the model;
[0040] Step S4, training the network model: training the constructed multi-angle panchromatic and multispectral image progressive fusion network, inputting the processed data set prepared in step S1 into the network model in step S2, and obtaining the model weights through training;
[0041] Step S5, save the model: After each round of training in step S3, input the validation set into the model for validation. If the loss value is less than the set threshold, save the best performing model. In the actual multimodal image fusion task, the multimodal image can be directly input into the trained end-to-end network model to generate the final fused image;
[0042] Embodiment 2:
[0043] like Figure 1 As shown, an embodiment of the present invention provides a multi-angle panchromatic and multispectral image progressive fusion method, which specifically includes the following steps:
[0044] Step S1, prepare down-resolution dataset and full-resolution dataset: This method uses data from the Gaofen-2 satellite, and constructs down-resolution experimental dataset and full-resolution experimental dataset based on multiple pre-processed original remote sensing images. The down-resolution experimental dataset is constructed according to the Wald protocol, including a training set, a validation set and a test set. Each dataset contains 3 types of data: multispectral images, reference images and panchromatic images, and the data sizes are 32×32, 128×128, and 128×128 respectively. The training set, validation set and test set contain 18,500 pairs, 1,500 pairs and 20 pairs of photos respectively. The full-resolution experimental dataset includes panchromatic images and multispectral images, with data sizes of 512×512 and 128×128 respectively, and the data volume is 20 pairs of photos;
[0045] Step S2, construct a multi-angle panchromatic and multi-spectral image progressive fusion network model: the network model structure is as follows Figure 2 As shown in the figure, the entire network model consists of a multi-scale feature extraction module, an enhanced fusion stream module and a multi-stream feature fusion module. The operation process of the entire network is as follows: First, the input image passes through the multi-scale feature extraction module and the enhanced fusion stream module in sequence to generate multi-scale feature information and enhanced fusion stream feature information. Then, the multi-stream feature fusion module takes the feature information of a certain scale and the enhanced fusion stream feature information as input, and gradually generates the fused enhanced fusion stream feature information. This module sequentially processes the combination of feature information from the first scale to the third scale and then from the third scale back to the first scale and the enhanced fusion stream feature information, and finally obtains the enhanced fusion stream feature information containing rich multi-scale fusion features. Finally, the feature information is added to the mapping information of the multispectral image to obtain a multispectral image with high spatial resolution.
[0046] The multi-scale feature extraction module of the present invention is composed of six instance residual blocks and two downsampling blocks, and its structure is as follows: Figure 3 As shown in the figure. The module is divided into three stages. The first stage consists of 3×3 convolution, the first and second instance residual blocks, which are used to extract high-resolution detail information and form the first-scale features. The second stage consists of the third and fourth instance residual blocks and downsampling convolution layers, which are used to extract medium-scale information and obtain the second-scale features. The third stage consists of the fifth and sixth instance residual blocks and downsampling convolution layers, which are used for low-resolution context information and obtain the third-scale information. Through the integration of multi-scale features, the subsequent processing performance and the feature expression ability of the network are significantly enhanced.
[0047] The formulas for each stage are as follows:
[0048] F MS1 / F p1 =Re s2(Re s1(Con 3x3 (X))),
[0049] F MS2 / F p2 = Down(Re s4(Re s3(F MS1 / F p1 ))),
[0050] F MS3 / F p3 = Down(Re s6(Re s5(F MS2 / F p2 ))),
[0051] Where X represents a panchromatic image or a multispectral image, and F MSi (i=1,2,3) represents the feature information of multispectral images at different scales, F pi (i=1,2,3) represents the feature information of different scales of the full-color image, Re si(i=1,2...,6) represents different instance residual blocks, and Down(·) represents the downsampling layer. 3x3 (·) represents a convolution operation with a kernel size of 3×3.
[0052] The residual block of the present invention is composed of two convolutional layers and an identity map, and its structure is as follows: Figure 4 The overall design aims to improve the performance and stability in image processing tasks through efficient feature transfer and normalization.
[0053] The calculation formula of the instance residual block is as follows:
[0054] out1 = LeakyReLU(conv 3x3 (x)),
[0055] out 1_1 ,out 1_2 =chunk(out1),
[0056] Re s(x)=Leaky(conv 3x3 (Cat(IN(out 1_1 )+out 1_2 ))+indentity(x),
[0057] where conv 3x3 (·) is a 3×3 convolution, and identity(·) is a residual map. Leaky(·) is an L-type activation function, chunk(·) is a channel split, IN(·) is an instance normalization layer, out1, out 1_1 and out 1_2 Output features for the first layer, split feature 1 and split feature 2.
[0058] The L-type activation function formula is as follows:
[0059]
[0060] The enhanced fusion flow module of the present invention is composed of three 3×3 convolutions and a linear activation function, and its structure is shown in the following figure: Figure 5 As shown in the figure, 3×3 convolution and linear activation function are used to transform the full image and multispectral image into the same feature space respectively, and then connected in series along the channel dimension. Then, the enhanced fusion flow features are obtained by roughly fusion through parameterized convolution blocks, so as to further improve the quality of the fused image.
[0061] The linear activation function formula is as follows:
[0062]
[0063] The multi-stream feature fusion module of the present invention is composed of three parts: a frequency domain interaction module, a space domain interaction module and a dual-domain feature fusion module. The input of this module is the multi-scale features after multi-scale feature extraction and the enhanced fusion stream features.
[0064] The frequency domain interaction module of the present invention is composed of three parts: amplitude fusion, phase fusion and information integration. Figure 6 As shown in the figure. The panchromatic image and multispectral image are converted to the frequency domain through fast Fourier transform, and the amplitude and phase information are extracted respectively. The amplitude fusion module is mainly composed of frequency-domain-aware attention and 1×1 convolutional layer. The high-resolution features of the panchromatic image are used to enhance the detail performance of the multispectral image. The weights of different frequency bands are adjusted through convolution and frequency-domain-aware attention mechanism to generate more representative amplitude features. The phase fusion module is composed of 1×1 convolutional layer and channel splicing. While retaining the spectral features of the multispectral image, it combines the geometric information of the panchromatic image to ensure spatial consistency. The information integration module converts the fused amplitude and phase information back to the spatial domain through inverse Fourier transform, and integrates the feature map through channel splicing and 1×1 convolutional layer to ensure the effective interaction of information from different channels. Finally, after multi-channel splicing and addition operations, a fused image containing rich spectral and spatial information is output.
[0065] The frequency domain perception attention described in the present invention is used to fuse the amplitude information of the full-color image and the multi-spectral image. The structure is as follows: Figure 8As shown. This module contains two branches to obtain global and local information of features: the first channel-focused attention branch consists of global average pooling, global maximum pooling, dimensional transformation, and channel feature matrix transformation, which is used to extract and fuse the global and local channel information of the input feature map, thereby generating an attention map that enhances the features of a specific channel. The second spatially focused attention branch consists of multiple dimensional transformations, average pooling, standard pooling, and 1×1 convolution. This branch first performs pooling and convolution on the feature map in different spatial dimensions to extract feature responses at different locations. By focusing on high-response areas in space and generating an attention map for strengthening specific areas of space, the retention of detail features and the expression of spatial information are improved. This enhances the effect of amplitude information fusion.
[0066] The frequency domain perception attention formula is as follows:
[0067] The first channel focuses on the attention branch:
[0068] X Avg =AvgPool(X),
[0069] X Max =MaxPool(X),
[0070] X Max_dig = dig(X Max ), X Avg_dig = dig(X Avg ),
[0071] X Max_band = band(X Max ), X Avg_band = band(X Avg ),
[0072]
[0073] Where AvgPool(·) and MaxPool(·) are global average pooling and global maximum pooling respectively. is the adaptive weight of the channel, is the input feature, S(·) and θ are the linear activation function and dynamic learning parameters of S, band(·) and dig(·) are the operations to generate band matrix and diagonal matrix respectively, X Avg and X Max are the results of global average pooling and global maximum pooling, respectively, X Max_dig and X Max_band are the diagonal matrix and band matrix of the global maximum pooling result respectively. Avg_dig and X Avg_band are the diagonal matrix and band matrix of the global average pooling results, respectively. and They are local channel weight and global channel weight respectively.
[0074] The second spatially focused attention branch:
[0075] X norm 1=Conv 1×1 (Norm(AvgPool(Permute(X,(1,0,2))))),
[0076] X norm 2=Conv 1×1 (Norm(AvgPool(Permute(X,(2,0,1))))),
[0077] X reshape1 =Permute(X norm 1,(1,0,2)),
[0078] X reshape2 =Permute(X norm 2,(2,1,0)),
[0079] in Indicates the dimension transformation of X, where X is the transformation object, (1,0,2) means (height, number of channels, width), and so on. Norm(·) means standard error pooling. Conv 1×1 Represented as a 1×1 convolution. and Represents the results of extracting and fusing information in different dimensions. and is the final result after dimension transformation.
[0080] Output attention map:
[0081]
[0082] Among them, ⊙ represents element-by-element multiplication, represents matrix multiplication, is the final attention map output, S(·) is the S linear activation function, and its formula is as follows:
[0083]
[0084] The overall process of the frequency domain interaction module is summarized as follows:
[0085] Use the Fourier transform to generate the magnitude and phase components:
[0086]
[0087] in and represents the amplitude and phase value after Fourier transform, where is the Fourier transform function.
[0088]
[0089] Among them, Conv 1×1 (·) is a 1x1 convolution, Cat(·) is channel concatenation, and ω is the feature attention map generated by frequency-domain aware attention.
[0090] Finally, the fused frequency domain feature map F is obtained through inverse Fourier transform. ff :
[0091]
[0092] in is the inverse Fourier transform function, F ff It is the final frequency domain fusion output.
[0093] The spatial domain interaction module of the present invention is composed of three serially connected residual depth separable convolution blocks and channel feature optimization modules, and the structure is shown in Figure 8. The residual depth separable convolution block is composed of a detail enhancement wavelet convolution layer, a point-by-point convolution layer and a 3×3 convolution layer. MSi ,F pi ] as input, and extract local features and low-frequency information through residual deep separable convolution blocks. The three residual deep separable convolution blocks use 9×9, 3×3 and 5×5 convolution kernels respectively, and the number of output channels is half of the input. After being processed by the residual deep separable convolution block, the output of the last layer is passed to the channel feature optimization module to flexibly adjust the weights of the feature channels, highlight key features and reduce redundant information, thereby enhancing the performance of the spatial feature map in details and local information.
[0094] The formula of the spatial domain interaction module is as follows:
[0095] F spa =ACEM(WDSCon(WDSCon(WDSCon([F MSi ,F pi ])))),
[0096] Among them, WDSCon(·) is the residual depth separable convolution block function, and ACEM(·) is the channel feature optimization module function. MSi (i=1,2,3) represents the feature information of multispectral images at different scales, F pi (i=1, 2, 3) represents the feature information of different scales of the full-color image.
[0097] The workflow of the residual depthwise separable convolutional block can be described by the following formula:
[0098] WDSCon=conv 1×1 (Wcon([F MSi ,F pi ]))+[F MSi ,F pi ],
[0099] where conv 1×1 (·) is point-wise convolution, Wcon(·) is detail-enhanced wavelet convolution, and F MSi (i=1,2,3) represents the feature information of multispectral images at different scales, F pi (i=1, 2, 3) represents the feature information of different scales of the full-color image.
[0100] The channel feature optimization module consists of global average pooling and multiple 3×3 convolutions, which aims to dynamically adjust the channel weights to strengthen key features and suppress redundant information, thereby improving the overall performance of the model, and finally inputting the spatial domain feature map F spa .
[0101] The dual-domain feature fusion module of the present invention is composed of three fusion reversible neural blocks, a channel feature optimization module and a 1x1 convolution, and its structure is as follows: Fig. 9 As shown. This module will take the output of the two interactive modules in the spatial domain and the frequency domain as input. The three fused reversible neural blocks are densely connected, and each transformation unit is based on the principle of affine transformation. Taking the first layer of fused reversible neural blocks as an example, the input features are split into two independent parts, feature 1 and feature 2, where feature 1 is additively transformed, and feature 2 is enhanced with affine transformation, and then the fused features are generated by channel splicing. After passing through the three fused reversible neural blocks in sequence, the output is spliced in the channel dimension and passed to the channel feature optimization module to generate a fused feature with rich high-frequency details, and finally fused with the previous enhanced fused stream feature to obtain the enhanced fused stream feature of the final output of the module.
[0102] The overall process of the above dual-domain feature fusion module is as follows:
[0103] F1=INM([F ff ,F spa ]), F2=INM(F1), F3=INM(F2),
[0104] F fui+1 =Conv 1×1 (Cat(ACEM(Cat(F1,F2,F3)),F fui )),
[0105] Among them, INM(·) is the fusion reversible neural block, Cat(·) is the channel splicing, ACEM(·) is the channel feature optimization module function, F fui and F fui+1 (i=0,1,2,3,4,5) is the enhanced fusion flow feature of the previous level and this level, F i (i=1,2,3) are the output features of three fused reversible neural blocks. Conv 1×1 Represents a convolution with a convolution kernel of 1×1.
[0106] The fusion reversible neural block in the present invention is composed of channel splitting, mapping function α module, mapping function β (γ) module and channel splicing. The two mapping function modules have the same structure. The specific structure is as follows: Fig.10 As shown in the figure. First, the input features are split into two parts in the channel dimension and passed to the upper and lower feature branches respectively. The mapping function α module consists of 3×3 convolution, instance normalization and identity mapping, which is used to generate the mapping function α. This transformation transforms the long-distance features into a form suitable for fusion with local features to achieve preliminary fusion of information. The mapping function β(γ) module generates β and γ functions. This transformation adjusts the expression ability of the fused features by scaling and displacement, thereby enhancing the flexibility and retention of information. Finally, the two-way features are merged in the channel splicing module to generate fused features containing multi-scale information, providing rich support for subsequent image processing.
[0107] Taking the first layer of fused reversible neural blocks as an example, the formula is as follows:
[0108] Input Features Split into two independent features and
[0109] X=Cat(X1,X2),
[0110] Where Cat(·) represents channel splicing.
[0111] For feature X1, an additive transformation is performed, and the additive change Y1 is applied. For feature 2, an enhanced affine transformation Y2 is applied, and finally Y is generated through a channel concatenation operation. The formula is as follows:
[0112]
[0113] Y1=X1+α(Y2),
[0114] Y=Cat(Y1,Y2),
[0115] Among them, α(·), β(·) and γ(·) represent mapping functions. Cat(·) represents channel concatenation, Represents the mathematical exponential function. Where Y1, Y2 and Y represent the upper branch integration feature, the lower branch integration feature and the final fusion feature respectively.
[0116] Step S3, determine the loss function and determine the optimal evaluation index of this method: the loss function of this network includes spatial loss, frequency domain loss and mutual information loss. In the reduced resolution experiment, quantitative indicators such as spectral angle mapper and global comprehensive error are used to evaluate the fusion effect. In the full resolution experiment, the spatial distortion index D is used. s and spectral distortion index D λ To further verify the fusion performance and accuracy of the model.
[0117] The total loss function of the present invention is formulated as a weighted sum of spatial loss, frequency loss, and mutual information loss:
[0118]
[0119] in For space loss, is the frequency domain loss, is the mutual information loss, α and λ are hyperparameters, and their values are both 0.1.
[0120] The spatial loss described in the present invention optimizes the network model by minimizing the loss function between the fused image and the corresponding reference image. It aims to maintain the spatial details and pixel-level similarity of the image, ensuring that the image is clearer, more natural and more realistic. Using L1 loss, The loss is expressed as follows:
[0121]
[0122] Among them, I H Represents the reconstructed HRMS image, I GT Represents the corresponding reference image.
[0123] The frequency domain loss described in the present invention aims to enhance the global characteristics and detail performance of the image, especially focusing on the improvement of high-frequency details and the consistency of low-frequency information, thereby improving the perception of high-frequency information. The L1 loss is used in the frequency domain to calculate the amplitude difference and phase difference between H and GT. The loss is expressed as follows:
[0124]
[0125] The amplitude and phase components of the Fourier transform are respectively and express.
[0126] The mutual information loss in the present invention aims to minimize the loss function by encouraging the learning of complementary information between panchromatic images and multispectral images, reducing information redundancy, and thus improving the quality of image reconstruction. The loss calculation formula is as follows:
[0127]
[0128] in is the joint entropy of the two features, and is the entropy of the two features, and It is the low-dimensional feature vector obtained by processing the output of the multi-scale feature extraction module through a 3×3 convolutional layer and two linear layers.
[0129] Step S4, training the network model: Based on the training set prepared in step S1, the panchromatic image and multispectral image in the data set are input into the fusion network model constructed in step S2 for training. The optimizer used in the iterative optimization process of the model is the Adam optimizer, and the total number of training rounds is 1000. The batch size is set to 16, the initial learning rate is set to 5×10-3, and the learning rate is reduced by 0.5 times after every 200 rounds.
[0130] Step S5, save the model: After each round of training in step S3, input the data in the validation set into the model for verification. If the loss value is less than the set threshold, save the current best performing fusion network model.
[0131] The fusion network of the present invention performs well in multiple spatial and spectral evaluation indicators, surpassing the existing technical level. In summary, the present invention uses end-to-end training, adopts a structurally efficient fusion network, and uses a deep learning algorithm to fully explore the interaction between frequency domain and spatial domain features to achieve accurate fusion of high-resolution multispectral images. This method provides a practical and efficient solution for future multimodal fusion based on remote sensing images.
[0132] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for progressive fusion of multi-angle panchromatic and multispectral images, characterized by: The steps include: S1, prepare down-resolution dataset and full-resolution dataset: use several pre-processed remote sensing images to construct down-resolution experimental dataset and full-resolution experimental dataset respectively, and divide the down-resolution experimental dataset into training set, validation set and test set, and use the full-resolution experimental dataset as the test set; S2, constructing a multi-angle panchromatic and multi-spectral image progressive fusion network model: The multi-angle panchromatic and multi-spectral image progressive fusion network model consists of a multi-scale feature extraction module, an enhanced fusion stream module, and a multi-stream feature fusion module; The multi-scale feature extraction module consists of six instance residual blocks and two downsampling layers, which are used to extract feature information of images of different scales, thereby achieving the capture and expression of multi-level features; The enhanced fusion flow module is composed of multiple convolutional layers, which is used to initially fuse the shallow features of the image to further improve the quality of the fused image later; The multi-stream feature fusion module includes a frequency domain interaction module, a space domain interaction module and a dual-domain feature fusion module, which are used to effectively interact and fuse the dual-modal information; The frequency domain interaction module is used to extract and process the frequency domain feature information of the image, capture the high-frequency details and change areas in the image through feature analysis in the frequency domain dimension, and thus enhance the expression of global information; The spatial domain interaction module is used to extract and process the spatial domain feature information of the image, thereby improving the expressiveness of the local spatial information of the image; The dual-domain feature fusion module is used to deeply fuse the frequency domain and space domain features to achieve the integration of multi-dimensional features, and finally output a more expressive comprehensive feature map to improve the fusion effect of the model; S3, determine the loss function and determine the optimal evaluation index of this method: the loss function includes spatial loss, frequency domain loss and mutual information loss; set the training loss threshold, and iteratively optimize the model until the predetermined number of training times is reached or the loss function converges, and save the model parameters; select the quantitative index of commonly used simulation data and the quantitative index of real data to evaluate the fusion performance and accuracy of the model; S4, training network model: training a multi-angle panchromatic and multispectral image progressive fusion network model, according to the training set prepared in step S1, inputting the panchromatic image and multispectral image in the data set into the fusion network model constructed in step S2 for training, and obtaining training weights; S5, save the model: After each round of training in step S3, the validation set is input into the model for verification. If the loss value is less than the set threshold, the best performing model is saved. In the actual multimodal image fusion task, the multimodal image can be directly input into the trained end-to-end network model to generate the final fused image.
2. The multi-angle panchromatic and multispectral image progressive fusion method according to claim 1, characterized in that: In step S1, the required data set is constructed using the GF-2 satellite data, and a reduced-resolution experimental data set and a full-resolution experimental data set are constructed based on multiple preprocessed original remote sensing images; the reduced-resolution experimental data set is constructed according to the Wald protocol, including a training set, a validation set and a test set.
3. The multi-angle panchromatic and multispectral image progressive fusion method according to claim 2, characterized in that: Each dataset contains three types of data: multispectral images, reference images, and panchromatic images; the full-resolution experimental dataset includes panchromatic images and multispectral images.
4. The method for progressive fusion of multi-angle panchromatic and multi-spectral images according to claim 1, characterized in that: The frequency domain interaction module in the multi-stream feature fusion module in step S2 consists of three parts: amplitude fusion, phase fusion and information integration. It uses the frequency domain perception attention mechanism to assign differentiated weights to the high-frequency and low-frequency information of the two modalities.
5. The multi-angle panchromatic and multi-spectral image progressive fusion method according to claim 4, characterized in that: The frequency-domain-aware attention module consists of two branches. The first channel-focused attention branch consists of global average pooling, global maximum pooling, dimensionality transformation, and channel feature matrix transformation. The second spatial-focused attention branch consists of multiple dimensionality transformations, average pooling, standard pooling, and 1×1 convolution.
6. The multi-angle panchromatic and multispectral image progressive fusion method according to claim 1, characterized in that: The dual-domain feature fusion module in the multi-stream feature fusion module in step S2 consists of three fused reversible neural modules, a channel feature optimization module and a 1×1 convolution.
7. The multi-angle panchromatic and multispectral image progressive fusion method according to claim 1, characterized in that: The fused reversible neural module consists of channel splitting, mapping function α module, mapping function β module and channel splicing, which further extracts and fuses the input features.
8. The multi-angle panchromatic and multispectral image progressive fusion method according to claim 7, characterized in that: The channel feature optimization module consists of global average pooling and multiple 3×3 convolutions, which are used to dynamically adjust channel weights, strengthen key features and suppress redundant information. Finally, the fusion features are combined with the enhanced fusion features through 1×1 convolution to form new enhanced fusion flow features.
9. The method for progressive fusion of multi-angle panchromatic and multi-spectral images according to any one of claims 1 to 8, characterized in that: The loss function in the entire network training process in step S3 is composed of spatial loss, frequency domain loss and mutual information loss. The spatial loss is used to measure the difference between the spatial details and pixels of the reference image and the predicted image. The frequency domain loss is used to measure the amplitude and phase difference between the reference image and the predicted image. The mutual information loss is used to quantify the similarity and complementarity between the panchromatic image features and the multispectral image features.
Citation Information
Cited By
Remote sensing image super-resolution fusion method based on frequency domain feature prior
CN120279368A
Feature integration method for interactive convolution and dynamic focusing of infrared image
CN120339779A
Feature integration method of interactive convolution and dynamic focusing for infrared images
CN120339779B
Water body extraction method and device based on multi-source multi-scale optical remote sensing data
CN120411810A
Underwater sound source localization method based on frequency domain feature enhancement and multi-scale wavelet convolution
CN120507718A