Multispectral image fusion method and system based on multi-task joint semi-supervised network

By introducing dynamic perception modules, attention aggregation modules and feature reconstruction modules into the multi-task joint semi-supervised network, combined with the degenerate network, the problem of difficulty in reducing spectral features and spatial structures and insufficient generalization capabilities of multi-spectral image data processing in the prior art is solved, and a more efficient image fusion effect is achieved.

CN120013775APending Publication Date: 2025-05-16DALIAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510029379.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When processing large-scale multispectral image data, it is difficult to effectively restore the spectral characteristics and spatial structure of the image, and the generalization ability of unknown data is weak.

Method used

Using a method based on a multi-task joint semi-supervised network, by constructing a dynamic perception module, attention aggregation module and feature reconstruction module, the multi-spectral image is dynamically assigned weights to different bands, information exchange and feature reconstruction are realized, and spatial and spectral degradation are simulated through the degradation network.

Benefits of technology

It improves the network's ability to generalize unknown data, restores the spectral characteristics and spatial structure of the image more accurately, and improves the accuracy and completeness of the fusion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013775A_ABST
    Figure CN120013775A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral image fusion method and system based on a multi-task joint semi-supervised network, and relates to the technical field of image fusion. The dynamic sensing module can dynamically and locally decompose spectrum information of the multispectral image into different sub-bands; different weighted values are given to different wavebands according to the influence of the wavebands on the fusion result, so that the importance of the different wavebands can be better captured. The attention aggregation module promotes information interaction and fusion among different source images; by learning the correlation and importance between source images, the module can dynamically adjust the weights of different input features to realize effective communication and integration of information, so that the performance of the model is improved. The feature reconstruction module aims at recovering spectrum and space information at the same time, the spectrum features and the space structure of the image can be better restored through the synergistic effect of convolution operations in different directions and a residual network, and the accuracy and the integrity of a fusion result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and in particular to a multispectral image fusion method and system based on a multi-task joint semi-supervised network. Background Art

[0002] High-resolution multispectral images play a vital role in the field of remote sensing. They are indispensable research data because they are rich in ground object information. However, it is not realistic to directly obtain such images due to the constraints of production costs and industrial technology. At present, researchers mainly use pan-sharpening technology, that is, fusing panchromatic images with low-resolution multispectral images, to indirectly obtain high-resolution multispectral images. Among them, panchromatic images can reveal small-scale ground object features, such as buildings, road networks, etc., while low-resolution multispectral images provide spectral characteristics of different bands, which are helpful for ground object classification, vegetation monitoring, environmental change analysis, etc.

[0003] In the early research of remote sensing image fusion, traditional methods were widely popular because of their simplicity. However, with the surge in data volume, traditional methods are unable to cope with large-scale data and complex image scenes. In contrast, deep learning-based models not only have the ability to learn complex nonlinear correlations, but also can quickly analyze massive image data through distributed and parallel computing technologies. Therefore, remote sensing image fusion methods based on deep learning have become a research hotspot in this field.

[0004] Deep learning-based architectures are divided into supervised and unsupervised learning, depending on whether a reference image is needed as guidance. Supervised learning relies on reduced-resolution data obtained by downsampling full-resolution data as input, and requires full-resolution images as references. Although this method is effective, it is highly dependent on reference image data, has weak generalization capabilities, and may not work well when processing unknown real data. Unsupervised learning can autonomously learn the distribution and characteristics of data without reference images, but it is sensitive to parameter settings. In addition, current research methods often directly upsample multispectral images and then cascade them with panchromatic images to achieve fusion, but ignore the modal differences brought by different sensors. For the band information in multispectral images, existing unsupervised frameworks usually adopt the same extraction strategy, failing to consider that not all spectral bands contribute equally to the fusion task. Summary of the invention

[0005] The purpose of the present invention is to propose a multispectral image fusion method and system based on a multi-task joint semi-supervised network, which can improve the generalization of the network in processing unknown data, better restore the spectral characteristics and spatial structure of the image, and improve the accuracy and completeness of the fusion results.

[0006] According to a first aspect of an embodiment of the present disclosure, a multispectral image fusion method based on a multi-task joint semi-supervised network is provided, comprising the following steps:

[0007] Construct the first dynamic perception module and the second dynamic perception module, and dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion results, so as to obtain the correlation and difference between different bands;

[0008] Constructing a first attention aggregation module, a second attention aggregation module and a third attention aggregation module for processing multispectral images or panchromatic images, and realizing information exchange between multispectral images and panchromatic images;

[0009] Construct the first feature reconstruction module and the second feature reconstruction module, and use the synergistic effect of convolution operations in different directions and the residual network to restore image information;

[0010] Construct degradation networks to simulate how spatial and spectral degradation occurs.

[0011] In one embodiment, the multispectral image passes through a 3×3 convolution block to obtain a feature map f, and the first dynamic perception module performs independent filtering operations on each channel of the feature map f, as shown below:

[0012] f 1 =Conv3(Dconv(δ(f)))

[0013] f 2 =ε(BN(PConv(Avg(FS))))

[0014] Where δ and ε represent ReLU and softmax activation functions respectively; DConv(·) and PConv(·) represent the corresponding depthwise convolution and pointwise convolution respectively; BN(·) represents batch normalization operation; Avg(·) is the global average pooling operation; Conv3(·) represents a 3×3 convolution layer; f 2 Equivalent to convolution kernels with different weights, and feature map f 1 Multiply, then perform upsampling to get the feature map F1∈R 2h×2w×C ;

[0015] The feature map F1 passes through a 3×3 convolution block to obtain the feature map f'. The second dynamic perception module performs independent filtering operations on each channel of the feature map f' as shown below:

[0016] f 1 =Conv3(Dconv(δ(f')))

[0017] f 2 =ε(BN(PConv(Avg(FS))))

[0018] where f 2 Equivalent to convolution kernels with different weights, and feature map f 1 Multiply, then perform upsampling to obtain the feature map F2∈R 4h×4w×C .

[0019] In one embodiment, the first attention aggregation module uses a residual architecture with an activation function and a convolution process to process full-color image features. It is the feature of the full-color image after 3×3 convolution; full-color image feature The size matches F2, so by placing Add F2 to get the mutual information F, and use the improved attention mechanism to operate F as follows:

[0020]

[0021] Where Mean(·) and Max(·) represent the mean pooling layer and the maximum pooling layer respectively; FC(·) represents the fully connected layer; σ(·) represents the sigmoid function, and ⊙ represents the element-by-element point multiplication; CRC(·) consists of two 3×3 convolutional layers and a ReLU activation function; the generated F2 and Add them together as the output of the first dynamic perception module;

[0022] The second attention aggregation module uses a residual architecture with activation functions and convolutional processes to process full-color image features. It is the feature of the first attention aggregation module output after 1×1 convolution; full-color image feature The size matches F1, so by Add the mutual information F' to F1, and use the improved attention mechanism to operate F' as follows:

[0023]

[0024] Generated F1 and Add them together as the output of the second attention aggregation module;

[0025] The third attention aggregation module uses a residual architecture with activation functions and convolutional processes to process full-color image features. It is the feature of the second attention aggregation module output after 3×3 convolution; full-color image feature The size matches the multispectral image MS, so by The mutual information F is obtained by adding it to MS, and the improved attention mechanism is used to operate F as follows:

[0026]

[0027] Generated MS and Added together as the output of the third attention aggregation module.

[0028] In one embodiment, the first feature reconstruction module includes three feature reconstruction blocks FRB; the output of the third attention aggregation module is subjected to 1×1 convolution and upsampling, and finally connected with the output of the second attention aggregation module, and the result is recorded as the feature feature As the input of FRB, it is as follows:

[0029]

[0030] The difference between f1, f2 and f3 is the direction of the last convolutional layer. Then f1, f2 and f3 are processed differently to get the output of FRB:

[0031]

[0032] Where DConv{·} indicates that both f1 and f3 are subjected to deep convolution; Conv1(·) represents a 1×1 convolutional layer. After convolution and activation function, the feature map and and Multiply them together to get the final output feature f b :

[0033]

[0034] f b Representing the feature information output by the feature reconstruction block, connecting the output results of the feature reconstruction blocks as the output of the entire first feature reconstruction module;

[0035] The second feature reconstruction module includes three feature reconstruction blocks FRB. The output of the first feature reconstruction module is subjected to 3×3 convolution and upsampling, and finally added to the output of the first attention aggregation module. The result is recorded as the feature feature As the input of FRB, it is as follows:

[0036]

[0037] The difference between f1, f2 and f3 is the direction of the last convolutional layer. Then f1, f2 and f3 are processed differently to get the output of FRB:

[0038]

[0039]

[0040] Where DConv·} indicates that both f1 and f3 are subjected to deep convolution; after convolution and activation function, the feature map and and Multiply them together to get the final output feature f b :

[0041]

[0042] f b It represents the feature information output by the feature reconstruction block, and connects the output results of the feature reconstruction blocks as the output of the second feature reconstruction module.

[0043] In one embodiment, the degradation network includes a spectrum degradation network, which uses a spectral response function to degrade the fusion result HM, and is implemented by convolution plus ReLU function plus convolution in a convolutional neural network:

[0044]

[0045] δ(·) represents the ReLU function, ⊙ represents the element-by-element point multiplication, and Conv1(·) represents a 1×1 convolutional layer.

[0046] In one embodiment, the degradation network includes a detail degradation network for maintaining spectral information loss while maintaining texture details consistent with the source full-color image; the process is as follows:

[0047]

[0048] Where Avg(·) is the global average pooling operation, Conv3(·) represents the 3×3 convolutional layer, and ε(·) represents the softmax function; and They are all intermediate feature maps; the pseudo full-color image is obtained by deformation, as follows:

[0049]

[0050] In one embodiment, after obtaining the degraded pseudo panchromatic image and pseudo multispectral image, the quality of the fusion result is improved by optimizing the loss between the source image and the degraded image; the expression is as follows:

[0051]

[0052] Where N and M represent the number of inputs with and without reference images, respectively, ||·||2 represents the L2 loss function, α and β are balance coefficients, and GT represents the reference image.

[0053] According to a second aspect of an embodiment of the present disclosure, a multispectral image fusion system based on a multi-task joint semi-supervised network is provided, comprising:

[0054] The first dynamic perception module and the second dynamic perception module dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion result, so as to obtain the correlation and difference between different bands;

[0055] The first attention aggregation module, the second attention aggregation module and the third attention aggregation module are used to process the multispectral image or the panchromatic image, and can also realize the information exchange between the multispectral image and the panchromatic image;

[0056] The first feature reconstruction module and the second feature reconstruction module utilize the synergistic effect of convolution operations in different directions and the residual network to restore image information;

[0057] Degraded networks that model how spatial and spectral degradation occurs.

[0058] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored and running on the memory, wherein when the processor executes the program, the multispectral image fusion method based on a multi-task joint semi-supervised network is implemented.

[0059] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the multispectral image fusion method based on a multi-task joint semi-supervised network is implemented.

[0060] Compared with the prior art, the above technical solution adopted by the present invention has the following advantages: In the present invention, a multi-task joint semi-supervised network is proposed, which is specifically used for remote sensing image fusion tasks. The network significantly enhances its generalization ability by combining unlabeled data with supervised learning methods. Specifically, the parallel training of two collaborative tasks, degradation and fusion, is realized, so that they can share knowledge, thereby improving the overall processing capability of the network. The dynamic perception module can assign different weight values ​​to different bands of the multispectral image; the attention aggregation module can fully extract the feature information of the source image and promote the interaction between information; the feature reconstruction module uses the synergy of convolution operations in different directions and residual connections to effectively restore the image information. At the same time, the degradation network is used to learn and simulate the spatial and spectral degradation process, which not only optimizes the degradation model, but also further improves the visual performance of the fused image. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The drawings in the specification, which constitute a part of the present application, are used to provide further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.

[0062] Figure 1 This is the overall framework diagram of the multi-task joint semi-supervised network;

[0063] Figure 2 This is a schematic diagram of the structure of the first dynamic perception module;

[0064] Figure 3 This is a schematic diagram of the structure of the second dynamic perception module;

[0065] Figure 4 This is the structural principle diagram of the first attention aggregation module;

[0066] Figure 5 This is the structural schematic diagram of the second attention aggregation module;

[0067] Figure 6 This is the structural principle diagram of the third attention aggregation module;

[0068] Figure 7 Reconstructing the module structure schematic diagram for the first feature;

[0069] Figure 8 Reconstruct the module structure schematic diagram for the second feature;

[0070] Fig. 9 It is the schematic diagram of the degenerate network structure;

[0071] Fig.10 The figure is a qualitative comparison between the present invention and other advanced fusion methods on the QuickBird, IKONOS and WorldView-II datasets. DETAILED DESCRIPTION

[0072] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.

[0073] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0074] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0075] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and systems according to various embodiments of the present disclosure. It should be noted that each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram can be implemented using a dedicated hardware-based system that performs a specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0076] Embodiment 1:

[0077] like Figure 1 As shown, this embodiment provides a multispectral image fusion method based on a multi-task joint semi-supervised network, comprising the following steps:

[0078] S1. Construct the first dynamic perception module and the second dynamic perception module, and dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion results, so as to obtain the correlation and difference between different bands;

[0079] Specifically, the low-resolution multispectral image MS∈R h×w×C and the full-color image P∈R H×W×c , where h, w and H, W represent the size of the image, C and c represent the channel information, and H / h=W / w=4.

[0080] The first dynamic perception module is Figure 2As shown in the figure, the multispectral image mainly provides spectral information for the fusion result, but different bands contribute differently to different scenes. Therefore, a dynamic perception module is designed to assign different weight values ​​to different bands according to their influence on the fusion result. This module enables the network to better learn the correlation and difference between different bands to ensure that more attention is paid to the bands that have a greater impact on the final result during the fusion process, thereby improving the quality of the fusion result. At the same time, due to the size difference between the panchromatic image and the multispectral image, the size of the multispectral image is changed during the upsampling operation. First, the multispectral image passes through a simple 3×3 convolution block to obtain the feature map f. In order to obtain different local receptive fields, the first dynamic perception module performs independent filtering operations on each channel, as shown below:

[0081] f 1 =Conv3(Dconv(δ(f)))

[0082] f 2 =ε(BN(PConv(Avg(FS))))

[0083] Where δ and ε represent ReLU and softmax activation functions respectively. DConv(·) and PConv(·) represent the corresponding depthwise convolution and pointwise convolution respectively. BN(·) represents batch normalization operation. Avg(·) is the global average pooling operation. Conv3(·) represents a 3×3 convolution layer. f 2 Equivalent to convolution kernels with different weights, and feature map f 1 Multiply them to focus the image on important channel information, and then perform upsampling to obtain F1∈R 2h×2w×C .

[0084] The second dynamic perception module is Figure 3 As shown, first, F1 passes through a simple 3×3 convolution block to obtain the feature map f'. In order to obtain different local receptive fields, the second dynamic perception module performs independent filtering operations on each channel as shown below:

[0085] f 1 =Conv3(Dconv(δ(f')))

[0086] f 2 =ε(BN(PConv(Avg(FS))))

[0087] Where δ and ε represent ReLU and softmax activation functions respectively. DConv(·) and PConv(·) represent the corresponding depthwise convolution and pointwise convolution respectively. BN(·) represents batch normalization operation. Avg(·) is the global average pooling operation. Conv3(·) represents a 3×3 convolution layer. f 2Equivalent to convolution kernels with different weights, and feature map f 1 Multiply to focus the image on important channel information, and then perform upsampling to obtain F2∈R 4h×4w×C .

[0088] S2. construct a first attention aggregation module, a second attention aggregation module and a third attention aggregation module for processing multispectral images or panchromatic images, and also realize information exchange between multispectral images and panchromatic images;

[0089] The first attention aggregation module is Figure 4 As shown in the figure, an attention aggregation module is designed when processing multispectral images and panchromatic images. This module can not only process multispectral images or panchromatic images separately, but also realize information exchange between multispectral images and panchromatic images. The attention mechanism can help the network focus on important features, thereby improving the representation ability of features. Through the attention aggregation module, it is possible to automatically learn which features in different source images are more critical to the task and thus extract important information more effectively. The first attention aggregation module mainly uses a residual architecture with activation functions and convolution processes to process panchromatic image features. It is the feature of the full-color image after 3×3 convolution. Full-color image feature The size matches F2, so by placing Add F2 to get the mutual information F, and use the improved attention mechanism to operate F as follows:

[0090]

[0091] Where Mean(·) and Max(·) represent mean pooling layer and maximum pooling layer respectively. FC(·) represents fully connected layer. σ(·) represents sigmoid function, and ⊙ represents element-wise point multiplication. CRC(·) consists of two 3×3 convolutional layers and a ReLU activation function. Generated F2 and Added together as the output of attention aggregation module 1.

[0092] The second attention aggregation module is Figure 5 As shown, the residual architecture with activation function and convolution process is mainly used to process full-color image features. It is the feature of the output of the first attention aggregation module after 1×1 convolution. Full-color image feature The size matches F1, so by Add the mutual information F' to F1, and use the improved attention mechanism to operate F as follows:

[0093]

[0094]

[0095] Where Mean(·) and Max(·) represent mean pooling layer and maximum pooling layer respectively. FC(·) represents fully connected layer. σ(·) represents sigmoid function, and ⊙ represents element-wise point multiplication. CRC(·) consists of two 3×3 convolutional layers and a ReLU activation function. Generated F1 and Added together as the output of the second attention aggregation module.

[0096] The third attention aggregation module is Figure 6 As shown, the residual architecture with activation function and convolution process is mainly used to process full-color image features. It is the feature of the output of the second attention aggregation module after 3×3 convolution. Full-color image feature The size matches the multispectral image MS, so by Add the mutual information F' and MS to obtain the mutual information F', and use the improved attention mechanism to operate F as follows:

[0097]

[0098] Where Mean(·) and Max(·) represent mean pooling layer and maximum pooling layer respectively. FC(·) represents fully connected layer. σ(·) represents sigmoid function, and ⊙ represents element-wise point multiplication. CRC(·) consists of two 3×3 convolutional layers and a ReLU activation function. Generated MS and Added together as the output of the third attention aggregation module.

[0099] S3. construct a first feature reconstruction module and a second feature reconstruction module, and utilize the synergistic effect of convolution operations in different directions and the residual network to restore image information;

[0100] In order to reconstruct the image using the feature information of the attention aggregation module, the influence of spectral information and texture information on the fusion result is balanced. A feature reconstruction module is designed to restore image information using the synergistic effect of convolution operations in different directions and residual networks, such as Figure 7As shown in the figure. By performing convolution operations from different directions, the feature reconstruction module can capture feature information from different directions in the image. The residual connection can reduce information loss, help the model better retain image details and structures, and improve the quality of the fusion result. The first feature reconstruction module is mainly composed of three feature reconstruction blocks FRB. The output of the third attention aggregation module is 1×1 convolved, upsampled, and finally connected with the output of the second attention aggregation module. The result is recorded as As the input of FRB, it is as follows:

[0101]

[0102] The difference between f1, f2 and f3 is mainly due to the different directions of the last convolutional layer. Then f1, f2 and f3 are processed differently to get the output of FRB:

[0103]

[0104] DConv{·} indicates that both f1 and f3 are subjected to deep convolution. Conv1(·) represents a 1×1 convolution layer. After convolution and activation function, the feature map and and Multiply them together to get the final output feature f b :

[0105]

[0106] where f b It represents the feature information output by the feature reconstruction block, and the output results of the feature reconstruction blocks are connected as the output of the entire first feature reconstruction module.

[0107] The second feature reconstruction module is Figure 8 As shown in Figure 1, it is mainly composed of three feature reconstruction blocks FRB; the output of the first feature reconstruction module undergoes 3×3 convolution and upsampling, and is finally added to the output of the first attention aggregation module. The result is recorded as As the input of FRB, it is as follows:

[0108]

[0109] The difference between f1, f2 and f3 is mainly due to the different directions of the last convolutional layer. Then f1, f2 and f3 are processed differently to get the output of FRB, whose mathematical expression is as follows:

[0110]

[0111] Where DConv{·} indicates that both f1 and f3 are subjected to deep convolution. After convolution and activation function, the feature map and and Multiply them together to get the final output feature f b . It is expressed as:

[0112]

[0113] where f b It represents the feature information output by the feature reconstruction block, and the output results of the feature reconstruction blocks are connected as the output of the entire second feature reconstruction module.

[0114] S4. Construct a degradation network to simulate how spatial and spectral degradation occurs.

[0115] Since pan-sharpening is to integrate the spectral information in the multispectral image and the spatial information in the panchromatic image into one image, the ideal pan-sharpening result is a pseudo multispectral image generated after degradation in the spatial and spectral domains. and pseudo-panchromatic images should be consistent with the input image. Using this assumption, Fig. 9 As shown in Figure 1, a degradation network is developed to try to simulate the way spatial and spectral degradation occurs. The first is the spectrum degradation network, which uses the spectral response function to degrade the fusion result HM, and uses convolution plus ReLU function plus convolution in the convolutional neural network:

[0116]

[0117] δ(·) represents the ReLU function, ⊙ represents the element-by-element point multiplication, and Conv1(·) represents a 1×1 convolutional layer. For the detail degradation network, the focus is to keep the spectral information missing while the texture details are consistent with the source full-color image; the process is as follows:

[0118]

[0119]

[0120] Where Avg(·) is the global average pooling operation, Conv3(·) represents the 3×3 convolutional layer, and ε(·) represents the softmax function. and They are all intermediate feature maps. The pseudo full-color image is obtained by simple deformation, as follows:

[0121]

[0122] After obtaining the degraded pseudo-panchromatic and pseudo-multispectral images, the quality of the fusion result is improved by optimizing the loss between the source image and the degraded image. The expression is as follows:

[0123]

[0124] Where N and M represent the number of inputs with and without reference images, respectively, ||·||2 represents the L2 loss function, α and β are balance coefficients, and GT represents the reference image. When N, i.e., the input image, is a reference-free image, α is set to 0, otherwise it is 1. Empirical experiments have shown that the ratio γ = M / N between reference-free input and reference input is set to 0.3, and β is set to 0.05.

[0125] The present invention selects test image sequences on the QuickBird, IKONOS and WorldView-III datasets to compare with eleven state-of-the-art remote sensing image fusion methods. Fig.10 It shows the overall effect and local feature details. Fig.10 The experimental results of the images with reduced resolution are shown on the left of each group, and the experimental results of the images with full resolution are shown on the right. The first group is the QuickBird dataset. Since the low-resolution dataset is synthesized, there will be some deviations. The reason for choosing to display the images with errors is that the data used in all experiments are consistent, so the fusion results are fair. On the other hand, what we want to show is that even in the case of destruction, the fusion results of the present invention are more in line with the actual effect than other aspects. This may be related to the constraint network designed during training to better fit the input image in terms of spectrum and detail texture. The images on the left of the first group show different degrees of white fog, while the images of the present invention are least affected by white fog. It can be clearly seen from the detail magnification window of the image on the right that the fused image produced using traditional technology has color distortion. Models based on deep learning usually have better fusion results than traditional ones. Although the fusion results based on deep learning do not show color deviation, there are still slight artifacts. Compared with these methods, the experimental results of the present invention are closer to the visual effect in terms of color and spatial fidelity. The second and third groups are the comparison results between the IKONOS and WorldView-II datasets, respectively. It can be seen from the results that when conducting these two types of comparative experiments, the present invention can retain more significant features in detail processing while avoiding serious spectral distortion problems, which plays a vital role in the practical application of remote sensing images.

[0126] In addition to subjective qualitative analysis, the present invention also uses objective quantitative indicators to evaluate the performance of the fusion results. Six currently popular reference metric methods are used to evaluate the down-resolution fusion results, including CC, FSIM, ERGAS, RASE, SAM and SSIM. The full-resolution experiment is implemented on the original image without available reference images. Therefore, the most commonly used non-reference image evaluation indicators are used to evaluate the performance of the fused image, including D λ , D S and QNR. The present invention uses 150 pairs of untrained full-resolution and reduced-resolution images as test sets to complete different remote sensing image fusion tasks, and the quantitative results are shown in the table below. In addition to the SSIM index, we achieved the highest scores on the other five reference indicators of the reduced-resolution dataset. The results show that even when dealing with biased data, the method of the present invention can still obtain more realistic fusion results. Based on the QNR index results, the fusion network of the present invention outperforms the previous fusion model in the full-resolution test, although not all indicators get the highest score. It may be due to the influence of the semi-supervised network. Under the action of the reference image, satisfactory fusion results can be maintained even in the face of biased data. The influence of different networks on the model refers to the ablation experiment content. It should be noted that when evaluating SCC, FSIM, and SSIM, the larger the calculated results are, the better the quality of the fused image. For ERGAS, SAM, and RASE, the smaller the value, the better the quality of the fused image. D λ It is used to measure the spectral correlation between the original multispectral image and the fusion result. S It is used to evaluate the spatial correlation between the original panchromatic image and the fusion result. QNR is dependent on D λ and D S A relatively comprehensive measure of the two. QNR can roughly reflect the panchromatic sharpening performance at full resolution. A large QNR indicates better quality of the fused image, while a small QNR indicates better quality of the fused image. λ and D S This means less distortion.

[0127] Table 1 Quantitative comparison between the proposed method and other advanced fusion methods on the QuickBird dataset

[0128]

[0129] Embodiment 2:

[0130] This embodiment provides a multispectral image fusion system based on a multi-task joint semi-supervised network, including:

[0131] The first dynamic perception module and the second dynamic perception module dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion result, so as to obtain the correlation and difference between different bands;

[0132] The first attention aggregation module, the second attention aggregation module and the third attention aggregation module are used to process the multispectral image or the panchromatic image, and can also realize the information exchange between the multispectral image and the panchromatic image;

[0133] The first feature reconstruction module and the second feature reconstruction module utilize the synergistic effect of convolution operations in different directions and the residual network to restore image information;

[0134] Degraded networks that model how spatial and spectral degradation occurs.

[0135] The dynamic perception module in the present invention can dynamically and locally decompose the spectrum information of the multispectral image into different sub-bands; by assigning different weight values ​​to different bands according to the influence of the bands on the fusion results, the importance of different bands can be better captured. The attention aggregation module promotes information interaction and fusion between different source images; by learning the correlation and importance between source images, the module can dynamically adjust the weights of different input features to achieve effective communication and integration of information, thereby improving the performance of the model. The feature reconstruction module aims to restore spectral and spatial information at the same time. Through the synergy of convolution operations in different directions and the residual network, the spectral characteristics and spatial structure of the image can be better restored, and the accuracy and completeness of the fusion results can be improved. During training, unlabeled data and labeled data are input into the network at the same time. The results of the fused image after mapping in different domains can be learned through the degradation network, and then constrained with the input image; this overcomes the problem of no reference data, and at the same time uses the uncertainty of unlabeled data to improve the generalization of the network in processing unknown data.

[0136] Embodiment three:

[0137] An electronic device includes a memory, a processor, and a computer program stored and running on the memory, wherein the processor implements the multispectral image fusion method based on a multi-task joint semi-supervised network when executing the program, including:

[0138] Construct the first dynamic perception module and the second dynamic perception module, and dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion results, so as to obtain the correlation and difference between different bands;

[0139] Constructing a first attention aggregation module, a second attention aggregation module and a third attention aggregation module for processing multispectral images or panchromatic images, and realizing information exchange between multispectral images and panchromatic images;

[0140] Construct the first feature reconstruction module and the second feature reconstruction module, and use the synergistic effect of convolution operations in different directions and the residual network to restore image information;

[0141] Construct degradation networks to simulate how spatial and spectral degradation occurs.

[0142] Embodiment 4:

[0143] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multispectral image fusion method based on a multi-task joint semi-supervised network, including:

[0144] Construct the first dynamic perception module and the second dynamic perception module, and dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion results, so as to obtain the correlation and difference between different bands;

[0145] Constructing a first attention aggregation module, a second attention aggregation module and a third attention aggregation module for processing multispectral images or panchromatic images, and realizing information exchange between multispectral images and panchromatic images;

[0146] Construct the first feature reconstruction module and the second feature reconstruction module, and use the synergistic effect of convolution operations in different directions and the residual network to restore image information;

[0147] Construct degradation networks to simulate how spatial and spectral degradation occurs.

[0148] Those skilled in the art should understand that the modules or steps of the present disclosure can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present disclosure is not limited to any specific combination of hardware and software.

[0149] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0150] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A multispectral image fusion method based on a multi-task joint semi-supervised network, characterized in that: The following steps are involved: Construct the first dynamic perception module and the second dynamic perception module, and dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion results, so as to obtain the correlation and difference between different bands; Constructing a first attention aggregation module, a second attention aggregation module and a third attention aggregation module for processing multispectral images or panchromatic images, and realizing information exchange between multispectral images and panchromatic images; Construct the first feature reconstruction module and the second feature reconstruction module, and use the synergistic effect of convolution operations in different directions and the residual network to restore image information; Construct degradation networks to simulate how spatial and spectral degradation occurs.

2. The multispectral image fusion method based on a multi-task joint semi-supervised network according to claim 1, characterized in that: The multispectral image passes through a 3×3 convolution block to obtain the feature map f. The first dynamic perception module performs independent filtering operations on each channel of the feature map f as shown below: f 1 =Conv3(Dconv(δ(f))) f 2 =ε(BN(PConv(Avg(FS)))) Where δ and ε represent ReLU and softmax activation functions respectively; DConv(·) and PConv(·) represent the corresponding depthwise convolution and pointwise convolution respectively; BN(·) represents batch normalization operation; Avg(·) is the global average pooling operation; Conv3(·) represents a 3×3 convolution layer; f 2 Equivalent to convolution kernels with different weights, and feature map f 1 Multiply, then perform upsampling to get the feature map F1∈R 2h×2w×C ; The feature map F1 passes through a 3×3 convolution block to obtain the feature map f'. The second dynamic perception module performs independent filtering operations on each channel of the feature map f' as shown below: f 1 =Conv3(Dconv(δ(f'))) f 2 =ε(BN(PConv(Avg(FS)))) where f 2 Equivalent to convolution kernels with different weights, and feature map f 1 Multiply, then perform upsampling to obtain the feature map F2∈R 4h×4w×C .

3. The multispectral image fusion method based on a multi-task joint semi-supervised network according to claim 2 is characterized in that: The first attention aggregation module uses a residual architecture with activation functions and convolutional processes to process full-color image features. It is the feature of the full-color image after 3×3 convolution; full-color image feature The size matches F2, so by placing Add F2 to get the mutual information F, and use the improved attention mechanism to operate F as follows: Where Mean(·) and Max(·) represent the mean pooling layer and the maximum pooling layer respectively; FC(·) represents the fully connected layer; σ(·) represents the sigmoid function, and ⊙ represents the element-by-element point multiplication; CRC(·) consists of two 3×3 convolutional layers and a ReLU activation function; the generated F2 and Add them together as the output of the first dynamic perception module; The second attention aggregation module uses a residual architecture with activation functions and convolutional processes to process full-color image features. It is the feature of the first attention aggregation module output after 1×1 convolution; full-color image feature The size matches F1, so by adding Add the mutual information F' to F1, and use the improved attention mechanism to operate F' as follows: Generated F1 and Add them together as the output of the second attention aggregation module; The third attention aggregation module uses a residual architecture with activation functions and convolutional processes to process full-color image features. It is the feature of the second attention aggregation module output after 3×3 convolution; full-color image feature The size matches the multispectral image MS, so by The mutual information F is obtained by adding it to MS, and the improved attention mechanism is used to operate F as follows: Generated MS and Added together as the output of the third attention aggregation module.

4. The multispectral image fusion method based on a multi-task joint semi-supervised network according to claim 1, characterized in that: The first feature reconstruction module includes three feature reconstruction blocks FRB; the output of the third attention aggregation module is subjected to 1×1 convolution and upsampling, and finally connected with the output of the second attention aggregation module, and the result is recorded as the feature feature As the input of FRB, it is as follows: The difference between f1, f2 and f3 is the direction of the last convolutional layer. Then f1, f2 and f3 are processed differently to get the output of FRB: Where DConv} means that both f1 and f3 are subjected to deep convolution; Conv1(·) represents a 1×1 convolution layer. After convolution and activation function, the feature map and and Multiply them together to get the final output feature f b : f b Representing the feature information output by the feature reconstruction block, connecting the output results of the feature reconstruction blocks as the output of the entire first feature reconstruction module; The second feature reconstruction module includes three feature reconstruction blocks FRB. The output of the first feature reconstruction module is subjected to 3×3 convolution and upsampling, and finally added to the output of the first attention aggregation module. The result is recorded as the feature feature As the input of FRB, it is as follows: The difference between f1, f2 and f3 is the direction of the last convolutional layer. Then f1, f2 and f3 are processed differently to get the output of FRB: Where DConv{·} indicates that both f1 and f3 are subjected to deep convolution; after convolution and activation function, the feature map and and Multiply them together to get the final output feature f b : f b It represents the feature information output by the feature reconstruction block, and connects the output results of the feature reconstruction blocks as the output of the second feature reconstruction module.

5. The multispectral image fusion method based on multi-task joint semi-supervised network according to claim 1, characterized in that: The degradation network includes a spectrum degradation network, which uses a spectrum response function to degrade the fusion result HM, and is implemented by using convolution plus ReLU function plus convolution in a convolutional neural network: δ(·) represents the ReLU function, ⊙ represents the element-by-element point multiplication, and Conv1(·) represents a 1×1 convolutional layer.

6. The multispectral image fusion method based on multi-task joint semi-supervised network according to claim 1, characterized in that: The degradation network includes a detail degradation network for maintaining the missing spectral information while the texture details are consistent with the source full-color image; the process is as follows: Where Avg(·) is the global average pooling operation, Conv3(·) represents the 3×3 convolutional layer, and ε(·) represents the softmax function; and They are all intermediate feature maps; the pseudo full-color image is obtained by deformation, as follows:

7. The multispectral image fusion method based on multi-task joint semi-supervised network according to claim 1, characterized in that: After obtaining the degraded pseudo-panchromatic image and pseudo-multispectral image, the quality of the fusion result is improved by optimizing the loss between the source image and the degraded image; the expression is as follows: Where N and M represent the number of inputs with and without reference images, respectively, ||·||2 represents the L2 loss function, α and β are balance coefficients, and GT represents the reference image.

8. A multi-spectral image fusion system based on a multi-task joint semi-supervised network, characterized in that: include: The first dynamic perception module and the second dynamic perception module dynamically assign corresponding weight values ​​to different bands of the multispectral image according to the influence of the bands on the fusion result, so as to obtain the correlation and difference between different bands; The first attention aggregation module, the second attention aggregation module and the third attention aggregation module are used to process the multispectral image or the panchromatic image, and can also realize the information exchange between the multispectral image and the panchromatic image; The first feature reconstruction module and the second feature reconstruction module utilize the synergistic effect of convolution operations in different directions and the residual network to restore image information; Degraded networks that model how spatial and spectral degradation occurs.

9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the multispectral image fusion method based on the multi-task joint semi-supervised network is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multispectral image fusion method based on a multi-task joint semi-supervised network is implemented.

Citation Information

Cited By

  • Optical image synthesis method based on spectrum feature perception diffusion model

    CN120526275A

  • Crop disease and insect pest intelligent monitoring system based on multispectral imaging and use method thereof

    CN120823435A

  • Crop disease and pest intelligent monitoring system based on multispectral imaging and method of using the same

    CN120823435B