A variational inference based blind multispectral and panchromatic image fusion method and system

CN118397413BActive Publication Date: 2026-09-11WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410533063.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2026-09-11
Estimated Expiration
2044-04-30

AI Technical Summary

Technical Problem

然而,全色图像与多光谱图像融合技术面临以下困难:(1)现有的深度学习类图像融合方法尽管可以取得较好的效果,但是模型的性能受训练数据的影响很大,往往会在训练数据上过拟合,导致模型的性能优劣难以控制,且当训练数据与测试数据之间存在较大差异时模型的性能会下降

Benefits of technology

[0120] Compared with existing technologies, the advantages and beneficial effects of this invention are as follows: This invention proposes a novel method for blind multispectral and panchromatic image fusion based on variational inference. It models the image fusion problem using variational inference, considers more complex degradation modes, and estimates the degradation parameters of multispectral and panchromatic images using variational inference to improve the generalization ability of the method. Simultaneously, it integrates the advantages of deep learning in learning from massive amounts of data with the theoretical guidance of variational inference, and designs a multispectral and panchromatic image fusion network based on the results of variational inference. Extensive simulation and practical experimental results demonstrate the effectiveness and practicality of this invention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118397413B_ABST
    Figure CN118397413B_ABST
Patent Text Reader

Abstract

The application provides a blind multispectral and panchromatic image fusion method and system based on variational inference. The multispectral and panchromatic image fusion problem is modeled by using variational inference to obtain an inverse problem to be solved. The inverse problem is taken as an optimization object of a network, a multispectral image and panchromatic image fusion network VBPN based on a convolutional neural network and degradation parameter estimation is designed, the image fusion network is composed of four sub-networks, and respectively includes a multispectral image blur kernel estimation sub-network, a multispectral image noise parameter estimation sub-network, a panchromatic image noise parameter estimation sub-network and a feature integration and image fusion sub-network. Finally, a loss function suitable for the fusion network is designed to guide network training, and the GaoFen-2 simulation dataset based on the Wald protocol is trained, so that a multispectral and panchromatic image fusion model is finally obtained to perform image fusion. A large number of simulation and actual experiment results prove the effectiveness and practicability of the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and deep learning technology, and relates to a blind multispectral and panchromatic image fusion method and system based on variational inference, which is suitable for multispectral and panchromatic image fusion scenarios with high precision and multiple scene requirements. Background Technology

[0002] High-quality multispectral images are an important category of images, often used in remote sensing fields such as agriculture, disaster detection, and urban planning due to their rich spatial and spectral details. However, acquiring high-quality multispectral images faces two major challenges. Firstly, current satellite imaging sensors include panchromatic and multispectral cameras. Panchromatic cameras have only a single band but high spatial resolution, capturing rich information about the spatial structure of ground features. Multispectral cameras can capture information about ground features across multiple frequency bands, possessing rich spectral characteristics, but often have lower spatial resolution and less detailed spatial features, making it difficult to directly obtain high-quality multispectral images. Secondly, while increasing the pixel count of a multispectral camera sensor can improve performance, it significantly increases manufacturing costs and also increases the sensor's size and weight, which is detrimental to satellite operation.

[0003] The fusion of panchromatic and multispectral images is an important research direction in the fields of remote sensing imaging and computer vision. It is widely used in visual fields such as semantic segmentation and object detection because it can make full use of the rich spatial details of panchromatic images and improve the spatial resolution of multispectral images while ensuring the spectral fidelity of multispectral images. However, the fusion technology of panchromatic and multispectral images faces the following difficulties: (1) Although existing deep learning-based image fusion methods can achieve good results, the performance of the model is greatly affected by the training data. It often overfits to the training data, making it difficult to control the performance of the model. Moreover, the performance of the model will decrease when there is a large difference between the training data and the test data. (2) The datasets used for training of current supervised deep learning image fusion methods are all generated based on the Wald protocol simulation. The image degradation is relatively simple and it is difficult to adapt to satellite data with different sensor performance.

[0004] Existing panchromatic and multispectral image fusion methods can be divided into two main categories: traditional methods and learning-based methods. Traditional methods can be further divided into CS-based methods, MRA-based methods, and VO-based methods. CS-based methods project low-resolution multispectral images into the transform domain and separate their spectral and spatial information, replacing the spatial information with a panchromatic image. Common CS methods include principal component analysis (PCA) (Kwarteng P, Chavez A. Extracting spectral contrast in Landsat Thematic Mapperimage data using selective principal component analysis[J]. Photogramm.Eng.Remote Sens,1989,55(1):339-348.). The fusion results obtained by this type of method have clear spatial details but are prone to spectral distortion. MRA-based methods regard the generalized sharpening process as a detail extraction and injection process under multi-scale decomposition. By decomposing the low-resolution multispectral image and the panchromatic image at multiple scales and injecting the spatial information extracted from the panchromatic image into the low-resolution multispectral image, the fused image is finally obtained through inverse transformation. Some researchers used wavelet transform to improve the spatial resolution of low-resolution multispectral images by adding high-order wavelet transform coefficients from panchromatic images with high spatial resolution to the spatial components (Nunez J, Otazu X, Fors O, et al. Multiresolution-based image fusion with additive wavelet decomposition[J].IEEE Transactions on Geoscience and Remote Sensing, 1999, 37(3):1204-1211.). This type of method has good spectral detail fidelity, but it is prone to spatial structure distortion. VO-based methods mainly model the relationship between panchromatic images, low-resolution multispectral images, and the final fused image, aiming to minimize the loss function, treating it as an ill-posed problem. Ballester et al. proposed the first variational image fusion method and achieved good results (Ballester C, Caseles V, Igual L, et al. A variational model for P+XS image fusion[J]. International Journal of Computer Vision, 2006, 69:43-58.). This type of method requires a large amount of computation and has high complexity. This type of method has a strong physical mechanism, but it is difficult to achieve a balance between spatial and spectral aspects.

[0005] In recent years, thanks to the rapid development of deep learning, many data-driven network-based image fusion methods have been proposed. Inspired by the application of convolutional neural networks (CNNs) in super-resolution of remote sensing images, Masi et al. proposed the first CNN-based image fusion network, PNN (Masi G, Cozzolino D, Verdoliva L, et al. Pansharpening by convolutional neural networks[J]. Remote Sensing, 2016, 8(7): 594.); Yang et al. improved spectral accuracy by designing a high-pass filter to preserve spatial information and adding upsampled multispectral images to the network output, thus constructing a deep neural network, PanNet, to achieve image fusion (Yang J, Fu X, Hu Y, et al. PanNet: A deep network architecture for pan-sharpening[C] / / Proceedings of the IEEE international conference on computer vision. 2017: 5449-5457.); Wei et al. proposed DRPNN, which improves the nonlinear fitting ability of the network by introducing residual learning to design a deep neural network with more layers (Wei Y, Yuan Q, Shen H, et al. Boosting the accuracy of multispectral image pansharpening by learning a deep residual network[J].IEEE Geoscience and Remote Sensing Letters,2017,14(10):1795-1799.); FusionNet proposed by Deng et al. combines traditional methods with machine learning methods to estimate nonlinear injection models and finally designs neural networks to achieve image fusion (Deng LJ, Vivone G, Jin C, et al. Detail injection-based deep convolutional neural networks for pansharpening[J].IEEE Transactions on Geoscience and Remote Sensing,2020,59(8):6995-7010.).The aforementioned deep learning-based multispectral and panchromatic image fusion methods can implicitly learn image features through a large amount of training data. However, they often lack clear theoretical guidance and are prone to overfitting and poor generalization ability. Therefore, many methods combine model-based methods with deep learning methods, utilizing the theoretical guidance of model-based methods with the data-driven approach and nonlinear fitting capabilities of deep learning. For example, VP-Net proposed by Tian et al. combines the physical modeling of the VO method with the construction of a spatial structure prior for nonlinear fitting through deep learning, achieving good results (Tian X, Li K, Wang Z, et al. VP-Net: An interpretable deep network for variational pansharpening[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 1-16.). Therefore, current deep learning-based methods often lack the guidance of physical models, resulting in reduced model generalization ability and poor image fusion performance in real-world scenarios. Summary of the Invention

[0006] This invention addresses the shortcomings of existing technologies by proposing a deep learning-based blind image fusion method based on variational inference. The technical solution employed in this invention is as follows: a blind multispectral and panchromatic image fusion method based on variational inference. First, variational inference modeling is performed on the image fusion problem, considering both blur and noise degradation in low-resolution multispectral images and noise degradation in panchromatic images. Second, in terms of network design, a blind image fusion network (VBPN) with four sub-networks is constructed. VBPN includes a blur kernel estimation sub-network for low-resolution multispectral images, a noise estimation sub-network for panchromatic images, and a feature integration and fusion sub-network. The proposed method exhibits good image fusion performance and better generalization ability. Finally, a loss function suitable for the network is designed to guide network training, which is then performed on the GaoFen-2 simulation dataset to obtain a multispectral and panchromatic image fusion model for image fusion. The method includes the following steps:

[0007] Step 1: Model the multispectral and panchromatic image fusion problem using variational inference. Decompose the inverse problem of multispectral and panchromatic image fusion into two parts: the part with low-resolution multispectral image M as the observation variable and the part with high-resolution panchromatic image P as the observation variable. Combine the blurring and noise degradation of low-resolution multispectral image and the noise degradation of high-resolution panchromatic image to solve the set of latent variables in the degradation process. Then, use the problem to be solved in the set of latent variables as the loss function of the fusion network to guide the training of the fusion network.

[0008] The set of latent variables Θ is:

[0009]

[0010] Where F represents the high-resolution multispectral image obtained by fusion, σ 2 Λ represents the noise variance in the multispectral image degradation process, and Λ represents the blur kernel parameter in the multispectral image degradation process. The image gradient representing the fusion result; δ 2 Represents the noise variance during the panchromatic image degradation process;

[0011] Step 2: Design a multispectral and panchromatic image fusion network based on convolutional neural network and degradation parameter estimation. The fusion network VBPN consists of four sub-networks, including a multispectral image blur kernel estimation sub-network, a multispectral image noise parameter estimation sub-network, a panchromatic image noise parameter estimation sub-network, and a feature integration and image fusion sub-network. The fusion network takes low-resolution multispectral and panchromatic images with unknown degradation processes as input.

[0012] The processing procedure of the multispectral image and panchromatic image fusion network is as follows:

[0013] First, the low-resolution multispectral image is input into the multispectral image noise parameter estimation subnetwork and the multispectral image blur kernel estimation subnetwork to estimate the degradation parameters of the low-resolution multispectral image. Then, the panchromatic image is input into the panchromatic image noise parameter estimation subnetwork to estimate the noise parameters of the panchromatic image. Finally, the low-resolution multispectral image is upsampled and input together with the panchromatic image into the feature integration and image fusion subnetwork to obtain the fusion result.

[0014] Step 3: Train a network for fusing multispectral and panchromatic images using a loss function, and then use the trained network to reconstruct high-resolution multispectral images.

[0015] Furthermore, the portion of the low-resolution multispectral image M used as the observation variable includes:

[0016] The transition from low-resolution multispectral to high-resolution multispectral can be understood as a multi-channel image super-resolution problem:

[0017] M=(F*K)↓ s +N

[0018] Where F represents the high-resolution multispectral image obtained by fusion, K represents the blur kernel, * represents the convolution operation, ↓ s Let M represent the downsampling of the spatial scale, and N represent additive white Gaussian noise; therefore, the probability distribution of M is modeled as:

[0019]

[0020] To more accurately estimate the blur and noise parameters of multispectral images, they are modeled at the pixel level, where M... j Represents the j-th pixel of M. Representing a Gaussian distribution, d is the total number of pixels in the low-resolution multispectral image. Represents the noise variance during the multispectral image degradation process;

[0021] Specifically for F and K introduce priors for Introducing the inverse Gamma distribution as its prior:

[0022]

[0023] Where IG(·) represents the inverse Gamma distribution. Let ξ be the noise variance of the j-th pixel in the multispectral image, α0 be a hyperparameter controlling the shape of the distribution, and ξ be the noise variance of the j-th pixel in the multispectral image. j The noise parameters are estimated based on real high-resolution multispectral images and low-resolution multispectral images in the dataset;

[0024] For the fused high-resolution multispectral image, a Gaussian distribution is introduced as its prior distribution:

[0025]

[0026] Where F i Let x be the i-th pixel of the fused high-resolution multispectral image, m be the total number of pixels in the high-resolution multispectral image, and x be the ith pixel. i For the real high-resolution multispectral images in the dataset, The hyperparameters are artificially defined to measure minute differences between the two; the fuzzy kernel K is considered to be an anisotropic Gaussian distribution.

[0027]

[0028]

[0029] Where G(·) represents an anisotropic Gaussian distribution, k i′j′ Let ρ be the value in the i′ row and j′ column of the multispectral image blur kernel, where ρ represents the Pearson correlation coefficient, λ1 and λ2 represent the noise variance in the two directions, and S represents the coordinates of the spatial point where the blur kernel acts. When the kernel size is (2r+1)*(2r+1), its range is [-r, r]. When the blur kernel size is determined, the covariance matrix can uniquely determine the blur kernel. Therefore, the three variables in the covariance matrix are represented as Λ:

[0030]

[0031] but:

[0032] K = G(Λ)

[0033]

[0034] Where G(Λ) represents an anisotropic Gaussian distribution, and Λ represents the blur kernel parameter in the multispectral image degradation process. It controls ρ and The hyperparameters representing the difference between them are ρ, which represents the Pearson correlation coefficient; κ0 is the hyperparameter controlling the inverse Gamma distribution. and These are given during the dataset simulation process.

[0035] Furthermore, the portion using the high-resolution panchromatic image P as the observed variable includes:

[0036] To model the relationship between panchromatic and multispectral images, we can consider the gradient perspective:

[0037]

[0038] in For gradient operators, The gradient of the panchromatic image is given by N, where N represents the number of spectral bands, Ω represents additive white Gaussian noise, and f is the gradient of the panchromatic image. n t represents a single band of the fused high-resolution multispectral image. n The weighting parameter represents the corresponding waveband;

[0039] set up:

[0040]

[0041] in, The image gradient representing the fusion result, t = {t1, t2, ... t N} represents a given set of parameters;

[0042] Similarly, modeling is performed at the pixel level, assuming that the gradient between the two follows a Gaussian distribution:

[0043]

[0044] Where m represents the number of pixels in the panchromatic image, which is numerically the same as the number of pixels in the high-resolution multispectral image, and therefore both are represented by m. This represents the noise variance during the panchromatic image degradation process; therefore, the following prior forms are introduced for the distribution of the Gaussian distribution's mean and variance:

[0045]

[0046]

[0047] Where y i Represents the gradient of the real high-resolution multispectral image in the dataset. Let α1 be a hyperparameter used to measure the difference between the two, and γ be a hyperparameter controlling the inverse Gamma distribution. i The obtained noise variance parameter is used for estimation.

[0048] Furthermore, according to Bayes' theorem, the posterior distribution of the problem to be solved for the set of latent variables can be written as:

[0049]

[0050] Let P(y|x) be the posterior distribution of variable y given the observed variable x, where P(·) represents the probability distribution and q(·) represents the approximate probability distribution that needs to be estimated through a network. Since this posterior distribution is difficult to solve, based on the principle of variational inference, we hope to obtain an approximate posterior distribution q(Θ|M,P) that approximates the true posterior distribution. The numerator of the posterior formula is rewritten as:

[0051]

[0052]

[0053] Furthermore, based on the idea of ​​variational inference, the approximate posterior distribution q(Θ|M,P) can be written as:

[0054]

[0055] According to the mean-field theorem, we assume that the latent variables to be solved are independent of each other. Therefore, we expand the two posterior distributions as follows:

[0056] q(F,Λ,σ 2 |M)=q(F|M)q(Λ|M)q(σ 2 |M)

[0057]

[0058] Based on the modeling of the degradation process, and according to the prior distribution of each latent variable to be solved, the approximate posterior distribution form is given:

[0059]

[0060]

[0061]

[0062]

[0063]

[0064] Where, {μ j ,β j ,m,η l ,u i ,a i ,b i} represents the distribution parameters of the latent variables that need to be estimated via the network. These are the hyperparameters given in the prior form of each latent variable;

[0065] According to Bayes' theorem, the joint probability distribution of two observed variables M and P can be modeled in the following form:

[0066]

[0067] in D represents the lower bound of evidence. KL (·) represents the KL divergence, which measures the distance between two distributions. Its specific form is:

[0068]

[0069]

[0070] Since the joint probability distribution logp(M,P) of the two observed variables M and P is a constant that is difficult to solve, the two approximate posterior probability distributions q(F,Λ,σ) are minimized according to the modeling form of the joint probability distribution. 2 |M), The distance between the evidence and the corresponding prior is transformed into maximizing the lower bound of the evidence;

[0071] The analytical solutions for each term in the lower bound of the evidence are given below:

[0072] By using reparameterization and resampling techniques, we can obtain:

[0073]

[0074] in, From q(F∣M), q(Λ∣M) and q(σ 2 Resampling was performed in |M).

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081] Where Γ(·) and ψ(·) represent the Gamma and Digamma functions, respectively.

[0082] Furthermore, the multispectral image blur kernel estimation subnetwork is composed of a convolutional neural network with a residual channel attention module CAB, which extracts feature information of each spectral channel from the low-resolution multispectral image and integrates it to obtain the feature information of the multispectral image blur kernel.

[0083] The multispectral image noise parameter estimation subnetwork and the panchromatic image noise parameter estimation subnetwork are composed of the existing image denoising network DnCNN with the same input and output scale, and a pooling function is added at the network output to obtain the noise variance parameter;

[0084] The feature integration and image fusion subnetwork consists of an encoder-decoder network with a feature integration layer. Both the encoder and decoder are composed of FE layers. There are image downsampling and upsampling operations between adjacent encoder and decoder layers to extract and integrate features of the input image from multiple scales, and finally obtain the fused high-resolution multispectral image.

[0085] Furthermore, the processing procedure of the multispectral image blur kernel estimation subnetwork is as follows:

[0086] First, shallow features are extracted through a 3×3 convolutional layer. Then, the shallow features are used to extract features from each spectral component through an 8-layer cascaded CAB module. Finally, the obtained features are passed through a 3×3 convolutional layer to obtain the estimated probability distribution parameters m, η1, η2 of the blur kernel of the low-resolution multispectral image.

[0087] The CAB module consists of convolutional layers, Leaky ReLU activation functions, and channel attention layers (CA layers). It takes shallow features from a multispectral image as input, which are then passed through a 3×3 convolutional layer and a Leaky ReLU activation function, followed by another 3×3 convolutional layer for further feature extraction before being input into the channel attention layer to extract features from each spectral channel. To facilitate network training, residual connections are added at the input and output of this module. The mathematical expression of the Leaky ReLU activation function is as follows:

[0088]

[0089] Where: αLeaky The parameter represents the slope of the negative half-axis of the activation function, where x is the independent variable;

[0090] Furthermore, assume that the features of the input CA Layer are F in The output feature is F out The process of transferring input features in CALayer is as follows:

[0091] F1 = AveragePool(F in )

[0092] F2 = LeakyRelu(Conv(F1))

[0093] F out =Sigmod(Conv(F2))

[0094] Where F1 and F2 are intermediate variables, AveragePool(·) is the average pooling operation, and Sigmod(·) is the activation function, whose mathematical expression is as follows:

[0095]

[0096] Furthermore, the processing procedures of the multispectral image noise parameter estimation subnetwork and the panchromatic image noise parameter estimation subnetwork are as follows;

[0097] The input image first passes through a 3×3 convolutional layer and a ReLU activation function to extract shallow features non-linearly. The features are then input into a 5-layer cascaded 3×3 convolutional layer and a Leaky ReLU activation function layer. Finally, the estimated multispectral image noise distribution parameter β and panchromatic image noise distribution parameters a and b are obtained through a 3×3 convolutional layer.

[0098] Furthermore, the processing procedure of the feature integration and image fusion subnetwork is as follows;

[0099] The input image is fed into the first encoder layer after passing through a 3×3 convolutional layer, and then encoded by several encoder layers. Skip connections are added between the encoder layer output and the decoder layer input at the same scale to integrate the encoder's feature information into the decoder, thereby improving the quality of image fusion. Finally, the probability distribution parameters μ of the high-resolution multispectral image F to be solved are obtained through the decoder, thus obtaining the high-resolution multispectral image F.

[0100] The FE Layer, which incorporates existing features, consists of two cascaded 3×3 convolutional layers and a Leaky ReLU activation function, and is used to extract feature information from multispectral and panchromatic images.

[0101] Furthermore, the loss function used comprises five parts: namely, the loss function of the multispectral image blur kernel estimation subnetwork is... The loss function of the multispectral image noise parameter estimation subnetwork is: The loss function of the panchromatic image noise parameter estimation subnetwork is: The loss function of the feature integration and image fusion subnetwork is: The loss function that serves to jointly constrain the training of the four sub-networks is:

[0102] in and Each term is composed of the KL divergence term from the variational inference result in step one, and is used to guide the training of each sub-network. Their mathematical expressions are as follows:

[0103]

[0104]

[0105]

[0106]

[0107] The loss function that jointly constrains the four subnetworks to achieve parameter transfer and joint optimization is... Composed of the expectation term derived from the variational inference in step one, its mathematical expression is:

[0108]

[0109] Ultimately, the overall loss function is:

[0110]

[0111] The present invention also provides a blind multispectral and panchromatic image fusion system based on variational inference, comprising the following steps:

[0112] The fusion problem modeling module is used to model the fusion problem of multispectral and panchromatic images using variational inference. It decomposes the inverse problem of multispectral and panchromatic image fusion into two parts: the part with low-resolution multispectral image M as the observation variable and the part with high-resolution panchromatic image P as the observation variable. It combines the blurring and noise degradation of low-resolution multispectral image and the noise degradation of high-resolution panchromatic image to solve the set of latent variables in the degradation process, and uses the problem to be solved in the set of latent variables as the loss function of the fusion network to guide the training of the fusion network.

[0113] The set of latent variables Θ is:

[0114]

[0115] Where F represents the high-resolution multispectral image obtained by fusion, σ 2 Λ represents the noise variance in the multispectral image degradation process, and Λ represents the blur kernel parameter in the multispectral image degradation process. The image gradient representing the fusion result; δ 2 Represents the noise variance during the panchromatic image degradation process;

[0116] The fusion network construction module is used to design a multispectral and panchromatic image fusion network based on convolutional neural networks and degradation parameter estimation. The fusion network VBPN consists of four sub-networks, including a multispectral image blur kernel estimation sub-network, a multispectral image noise parameter estimation sub-network, a panchromatic image noise parameter estimation sub-network, and a feature integration and image fusion sub-network. The fusion network takes low-resolution multispectral and panchromatic images with unknown degradation processes as input.

[0117] The processing procedure of the multispectral image and panchromatic image fusion network is as follows:

[0118] First, the low-resolution multispectral image is input into the multispectral image noise parameter estimation subnetwork and the multispectral image blur kernel estimation subnetwork to estimate the degradation parameters of the low-resolution multispectral image. Then, the panchromatic image is input into the panchromatic image noise parameter estimation subnetwork to estimate the noise parameters of the panchromatic image. Finally, the low-resolution multispectral image is upsampled and input together with the panchromatic image into the feature integration and image fusion subnetwork to obtain the fusion result.

[0119] The final fusion module is used to train a fusion network for multispectral and panchromatic images by combining a loss function, and to reconstruct high-resolution multispectral images using the trained fusion network.

[0120] Compared with existing technologies, the advantages and beneficial effects of this invention are as follows: This invention proposes a novel method for blind multispectral and panchromatic image fusion based on variational inference. It models the image fusion problem using variational inference, considers more complex degradation modes, and estimates the degradation parameters of multispectral and panchromatic images using variational inference to improve the generalization ability of the method. Simultaneously, it integrates the advantages of deep learning in learning from massive amounts of data with the theoretical guidance of variational inference, and designs a multispectral and panchromatic image fusion network based on the results of variational inference. Extensive simulation and practical experimental results demonstrate the effectiveness and practicality of this invention. Attached Figure Description

[0121] Figure 1 This is a diagram of the CAB module and FE Layer structure.

[0122] Figure 2 This is a diagram of the subnetwork structure for the multispectral image blur kernel estimator.

[0123] Figure 3 This is a diagram of the subnetwork structure for estimating noise parameters in multispectral images.

[0124] Figure 4 This is a diagram of the subnetwork structure for estimating noise parameters in a panchromatic image.

[0125] Figure 5 This is a diagram of the sub-network structure for feature integration and image fusion.

[0126] Figure 6 It is a graph showing the interaction between the network and the loss function.

[0127] Figure 7 This is the GaoFen-2 simulation and QuickBird real dataset test input diagram of the embodiment.

[0128] Figure 8 This is an image showing the fusion result of multispectral and panchromatic images from the GaoFen-2 simulation and the real QuickBird dataset in the example.

[0129] Figure 9 This is a comparison chart of the simulation and real dataset processing results of the embodiment with the results of different fusion methods. Detailed Implementation

[0130] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.

[0131] This invention primarily addresses the application requirements of multispectral and panchromatic image fusion, proposing a blind multispectral and panchromatic image fusion method based on variational inference. Variational inference is used to model the image fusion problem, simultaneously considering the blurring and noise degradation of low-resolution multispectral images and the noise degradation of panchromatic images. The variational inference results are used to construct an image fusion network (VBPN) with four subnetworks. VBPN includes a blur kernel estimation subnetwork for low-resolution multispectral images, a noise estimation subnetwork for panchromatic images, and a feature integration and fusion subnetwork.

[0132] Step 1: Perform variational inference modeling for the multispectral and panchromatic image fusion problem, decomposing the inverse problem of multispectral and panchromatic image fusion into two parts: the part with the low spatial resolution multispectral image M as the observation variable and the part with the high spatial resolution panchromatic image P as the observation variable.

[0133] The transition from low-resolution multispectral to high-resolution multispectral can be understood as a multi-channel image super-resolution problem:

[0134] M=(F*K)↓ s +N

[0135] Where F represents the high-resolution multispectral image obtained by fusion, K represents the blur kernel, * represents the convolution operation, ↓ s Let M represent the downsampling of the spatial scale, and N represent additive white Gaussian noise. Therefore, the probability distribution of M can be modeled as:

[0136]

[0137] To more accurately estimate the blur and noise parameters of multispectral images, they are modeled at the pixel level, where M... j Represents the j-th pixel of M. Representing a Gaussian distribution, d is the total number of pixels in the low-resolution multispectral image, and σ 2 This represents the noise variance during the multispectral image degradation process.

[0138] Specifically for F and K introduce priors for Introducing the inverse Gamma distribution as its prior:

[0139]

[0140] Where IG(·) represents the inverse Gamma distribution. Let ξ be the noise variance of the j-th pixel in the multispectral image, and α0 be a hyperparameter controlling the shape of the distribution. In this embodiment, α0 = 40.5, which is set manually. j The noise parameters are estimated based on real high-resolution multispectral images and low-resolution multispectral images in the dataset.

[0141] For the fused high-resolution multispectral image, a Gaussian distribution is introduced as its prior distribution:

[0142]

[0143] Where F i Let x be the i-th pixel of the fused high-resolution multispectral image, m be the total number of pixels in the high-resolution multispectral image, and x be the ith pixel. i For the real high-resolution multispectral images in the dataset, Hyperparameters are manually set parameters used to measure minute differences between the two; in this embodiment... For the fuzzy kernel K, it is assumed to be an anisotropic Gaussian distribution:

[0144]

[0145]

[0146] Where G(·) represents an anisotropic Gaussian distribution, k i′j′ Let ρ be the value in the i′ row and j′ column of the multispectral image blur kernel, where ρ represents the Pearson correlation coefficient, λ1 and λ2 represent the noise variances in the two directions, and S represents the coordinates of the spatial point where the blur kernel acts. When the kernel size is (2r+1)*(2r+1), its range is [-r, r]. When the blur kernel size is determined, the covariance matrix can uniquely determine the blur kernel. Therefore, the three variables in the covariance matrix are represented as Λ:

[0147]

[0148] but:

[0149] K = G(Λ)

[0150]

[0151] Where G(·) represents an anisotropic Gaussian distribution, and Λ represents the blur kernel parameter in the multispectral image degradation process. It controls ρ and The hyperparameter of the difference between them, ρ represents the Pearson correlation coefficient. In this embodiment... κ0 is a hyperparameter controlling the inverse Gamma distribution; in this embodiment, κ0 = 50. and These are given during the dataset simulation process.

[0152] To model the relationship between panchromatic and multispectral images, we can consider the gradient perspective:

[0153]

[0154] in For gradient operators, The gradient of the panchromatic image is given by N, which represents the number of spectral bands (N = 4 in this embodiment), and Ω represents additive white Gaussian noise. n t represents a single band of the fused high-resolution multispectral image. n These represent the weight parameters for the corresponding waveband. For ease of modeling, we assume:

[0155]

[0156] in, The image gradient representing the fusion result, t = {t1, t2, ... t N} represents a given set of parameters; in this embodiment, t = {1,1,1,1}.

[0157] Similarly, modeling is performed at the pixel level, assuming that the gradient between the two follows a Gaussian distribution:

[0158]

[0159] Where m represents the number of pixels in the panchromatic image, which is numerically the same as the number of pixels in the high-resolution multispectral image, and therefore both are represented by m. This represents the noise variance during the panchromatic image degradation process. Furthermore, the following prior forms are introduced for the distributions of the Gaussian distribution's mean and variance:

[0160]

[0161]

[0162] Where p() represents the probability distribution function, y i Represents the gradient of the real high-resolution multispectral image in the dataset. This is a hyperparameter used to measure the difference between the two; in this embodiment... α1 is a hyperparameter controlling the inverse Gamma distribution; in this embodiment, α1 = 40.5, γ i The noise variance parameter is estimated.

[0163] Furthermore, based on the above modeling process, the set of latent variables to be solved can be obtained as follows:

[0164]

[0165] For ease of representation, let P(y|x) be the posterior distribution of variable y given the observed variable x, P(·) denote the probability distribution, and q(·) denote the approximate probability distribution that needs to be estimated through a network. According to Bayes' theorem, the posterior distribution of this problem can be written as:

[0166]

[0167] Since the posterior distribution is difficult to solve, based on the principle of variational inference, we hope to obtain an approximate posterior distribution q(Θ|M,P) that approximates the true posterior distribution. The numerator of the posterior formula can be rewritten as:

[0168]

[0169]

[0170] Furthermore, based on the idea of ​​variational inference, the approximate posterior distribution q(Θ|M,P) can be written as:

[0171]

[0172] According to the mean-field theorem, the latent variables to be solved are considered to be independent of each other. Therefore, the two posterior distributions can be expanded as follows:

[0173] q(F,Λ,σ 2 |M)=q(F|M)q(Λ|M)q(σ 2 |M)

[0174]

[0175] Based on the modeling of the degradation process, the approximate posterior distribution can be derived from the prior distribution of each latent variable to be solved:

[0176]

[0177]

[0178]

[0179]

[0180]

[0181] Where, {μ i ,β i ,m,η l ,u i ,a i ,b i} represents the distribution parameters of the latent variables that need to be estimated via the network. These are the hyperparameters given in the prior form of each latent variable.

[0182] According to Bayes' theorem, the joint probability distribution of two observed variables M and P can be modeled in the following form:

[0183]

[0184] in D represents the lower bound of evidence. KL (·) represents the KL divergence, which measures the distance between two distributions. Its specific form is:

[0185]

[0186]

[0187] Since the joint probability distribution logp(M,P) of the two observed variables M and P is a constant that is difficult to solve, the two approximate posterior probability distributions q(F,Λ,σ) are minimized according to the modeling form of the joint probability distribution. 2 |M), The distance between the given prior and the evidence can be transformed into maximizing the lower bound of evidence. The analytical solutions for each term in the lower bound of evidence are given below:

[0188] By using reparameterization and resampling techniques, we can obtain:

[0189]

[0190] in, From q(F∣M), q(Λ∣M) and q(σ 2 Resampling was performed in |M).

[0191]

[0192]

[0193]

[0194]

[0195]

[0196]

[0197] Where Γ(·) and ψ(·) represent the Gamma and Digamma functions, respectively.

[0198] Step 2: Design a multispectral image blur kernel estimation subnetwork: as shown in the attached diagram. Figure 2 As shown, the multispectral image blur kernel estimation subnetwork consists of convolutional layers of the Residual Channel Attention Module (CAB). This subnetwork takes a low-resolution multispectral image as input. First, it extracts shallow features through a 3×3 convolutional layer. Then, these shallow features are processed through an 8-layer cascaded CAB module to extract features from each spectral component. Finally, the resulting features are processed through another 3×3 convolutional layer to obtain the probability distribution parameters m, η1, and η2 of the estimated low-resolution multispectral image blur kernel. This subnetwork can extract feature information from each spectral channel of the low-resolution multispectral image and integrate it to obtain the feature information of the multispectral image blur kernel.

[0199] Further details are attached. Figure 1As shown, the CAB module consists of convolutional layers, a Leaky ReLU activation function, and a channel attention layer (CA layer). It takes shallow features from a multispectral image as input, which are then passed through a 3×3 convolutional layer and a Leaky ReLU activation function, followed by another 3×3 convolutional layer for further feature extraction before being input into the channel attention layer to extract features from each spectral channel. To facilitate network training, residual connections are added at the input and output of this module. The mathematical expression of the Leaky ReLU activation function is as follows:

[0200]

[0201] Where: α Leaky The parameter represents the slope of the negative half-axis of the activation function.

[0202] Furthermore, assume that the features of the input CA Layer are F in The output feature is F put The process of transferring input features in CALayer is as follows:

[0203] F1 = AveragePool(F in )

[0204] F2 = LeakyRelu(Conv(F1))

[0205] F out =Sigmod(Conv(F2))

[0206] Where AveragePool(·) is the average pooling operation, and Sigmod(·) is the activation function, whose mathematical expression is as follows:

[0207]

[0208] Step 3: Design the multispectral image noise parameter estimation subnetwork and the panchromatic image noise parameter estimation subnetwork. (See attached diagram) Figure 3 With appendix Figure 4 As shown, the multispectral image noise parameter estimation subnetwork and the panchromatic image noise parameter estimation subnetwork are constructed from existing image denoising networks (DnCNN) with the same input and output scales. A pooling function is added at the network output to obtain the noise variance parameter. The two noise parameter estimation subnetworks take low-resolution multispectral images and panchromatic images as inputs, respectively. The input images first pass through a 3×3 convolutional layer and a ReLU activation function to extract shallow features non-linearly. These features are then input to a 5-layer cascaded 3×3 convolutional layer and a Leaky ReLU activation function layer. Finally, the 3×3 convolutional layer yields the multispectral image noise distribution parameter β and the panchromatic image noise distribution parameters a and b. The mathematical expression of the ReLU activation function is:

[0209] ReLU(x) = max(0,x)

[0210] Step 4: Design the feature integration and image fusion subnetwork. (See attached image) Figure 5 As shown, the feature integration and image fusion subnetwork consists of an encoder-decoder network with an existing feature integration layer (FE Layer). Each encoder and decoder layer is composed of an FE Layer. Image downsampling and upsampling operations are performed between adjacent encoder and decoder layers to extract features from multispectral and panchromatic images at different spatial scales. The input image is fed into the first encoder layer after passing through a 3×3 convolutional layer, and then encoded by four encoder layers. Skip connections are added between the encoder layer output and decoder layer input at the same scale to integrate the encoder's feature information into the decoder, improving the quality of image fusion. Finally, the probability distribution parameters μ of the high-resolution multispectral image F to be solved are obtained through the decoder, thus obtaining the high-resolution multispectral image F. In this example, there are a total of 4 encoder-decoder layers.

[0211] Furthermore, the aforementioned FE Layer consists of two cascaded 3×3 convolutional layers and a Leaky ReLU activation function, used to extract feature information from multispectral and panchromatic images.

[0212] Step 5: Design a loss function suitable for this network. The loss function consists of five parts: the loss function of the multispectral image blur kernel estimation subnetwork is as follows. The loss function of the multispectral image noise parameter estimation subnetwork is: The loss function of the panchromatic image noise parameter estimation subnetwork is: The loss function of the feature integration and image fusion subnetwork is: The loss function that serves to jointly constrain the training of the four sub-networks is:

[0213] in and Each term is composed of the KL divergence term from the variational inference result in step 1, and is used to guide the training of each sub-network. Their mathematical expressions are as follows:

[0214]

[0215]

[0216]

[0217]

[0218] The loss function that jointly constrains the four subnetworks to achieve parameter transfer and joint optimization is... Composed of the expectation term derived from the variational inference in step 1, its mathematical expression is:

[0219]

[0220] Ultimately, the overall loss function is:

[0221]

[0222] As attached Figure 6 As shown, the loss function constrains the training of each network and enables parameter transfer between networks.

[0223] During network training, the Adam optimizer was used, with its parameters fixed at β1 = 0.9 and β2 = 0.999. The initial learning rate was 3 × 10⁻⁶. -4 The network was trained on the GaoFen-2 simulation dataset generated via the Wald protocol. The trained multispectral and panchromatic image fusion model was then tested on the corresponding test sets. Sample input images for both test sets are attached. Figure 7 As shown. The final processing results on the two datasets are attached. Figure 8 As shown.

[0224] Based on the multispectral and panchromatic image fusion results obtained from the above steps, in order to compare with other methods, PNN (Masi G, Cozzolino D, Verdoliva L, et al. Pansharpening by convolutional neural networks[J]. Remote Sensing, 2016, 8(7): 594.), MSDCNN (Yuan Q, Wei Y, Meng X, et al. A multiscale and multidepth convolutional neural network for remote sensing imagery pan-sharpening[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2018, 11(3): 978-989.), and FusionNet (Deng LJ, Vivone G, Jin C, et al. Detail injection-based deep convolutional neural networks for pansharpening[J]. IEEE Transactions on Geoscience and RemoteSensing, 2020, 59(8):6995-7010.), Hyper-DSNet (Zhuo YW, Zhang TJ, Hu JF, etal. Adeep-shallow fusion network with multidetail extractor and spectralattention for hyperspectral pansharpening[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing,2022,15:7539-7555.), ADKNet (Peng S, Deng LJ, Hu JF, et al. Source-Adaptive Discriminative Kernels basedNetwork for Remote Sensing Pansharpening[C] / / IJCAI.2022:1283-1289.The method of this invention was compared with the method of the present invention as a comparative method, and the results are attached. Figure 9 As shown.

[0225] To quantitatively evaluate the results of multispectral and panchromatic image fusion, peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), universal image quality index (UIQI,Q), comprehensive global error (ERGAS), and spectral angle mapping function (SAM) are introduced as supervised evaluation indicators. Spectral distortion index (D) is also included. λ Spatial distortion index (D) S The Comprehensive Distortion Evaluation Index (QNR) and other metrics are used as unsupervised evaluation metrics to measure the effect of multispectral and panchromatic image fusion. In supervised evaluation metrics, higher values ​​for the first three metrics indicate better image fusion, while lower values ​​for the last two metrics indicate better fusion. In unsupervised evaluation metrics, lower values ​​for the first two metrics indicate better fusion, while higher values ​​for the last metric indicate better fusion. GaoFen-2 satellite data was selected as the data source for the simulation dataset. The simulation dataset was generated according to the Wald protocol and includes unknown levels of degradation, i.e., blur downsampling and downsampling were performed on the multispectral and panchromatic images respectively, and these were used as training data pairs with the unsampled multispectral images. The real dataset consists of QuickBird satellite multispectral and panchromatic images. During actual training and testing, both multispectral and panchromatic images were cropped into smaller image patches for training, ensuring that the ground cover coverage of the multispectral and panchromatic images was consistent. Quantitative comparison results on the GaoFen-2 simulation dataset are shown in Table 1.

[0226] Table 1. Quantitative comparison results on the GaoFen-2 simulation dataset.

[0227]

[0228] The quantitative comparison results on the real QuickBird dataset are shown in Table 2:

[0229] Table 2 Quantitative comparison results on the real QuickBird dataset

[0230]

[0231] Quantitative results show that the image fusion results obtained by the method proposed in this invention are superior to existing methods on both simulated and real datasets, achieving better image fusion results and improving the spatial resolution of multispectral images.

[0232] On the other hand, the present invention also provides a blind multispectral and panchromatic image fusion system based on variational inference, comprising the following modules:

[0233] The fusion problem modeling module is used to model the fusion problem of multispectral and panchromatic images using variational inference. It decomposes the inverse problem of multispectral and panchromatic image fusion into two parts: the part with low-resolution multispectral image M as the observation variable and the part with high-resolution panchromatic image P as the observation variable. It combines the blurring and noise degradation of low-resolution multispectral image and the noise degradation of high-resolution panchromatic image to solve the set of latent variables in the degradation process, and uses the problem to be solved in the set of latent variables as the loss function of the fusion network to guide the training of the fusion network.

[0234] The set of latent variables Θ is:

[0235]

[0236] Where F represents the high-resolution multispectral image obtained by fusion, σ 2 Λ represents the noise variance in the multispectral image degradation process, and Λ represents the blur kernel parameter in the multispectral image degradation process. The image gradient representing the fusion result; δ 2 Represents the noise variance during the panchromatic image degradation process;

[0237] The fusion network construction module is used to design a multispectral and panchromatic image fusion network based on convolutional neural networks and degradation parameter estimation. The fusion network VBPN consists of four sub-networks, including a multispectral image blur kernel estimation sub-network, a multispectral image noise parameter estimation sub-network, a panchromatic image noise parameter estimation sub-network, and a feature integration and image fusion sub-network. The fusion network takes low-resolution multispectral and panchromatic images with unknown degradation processes as input.

[0238] The processing procedure of the multispectral image and panchromatic image fusion network is as follows:

[0239] First, the low-resolution multispectral image is input into the multispectral image noise parameter estimation subnetwork and the multispectral image blur kernel estimation subnetwork to estimate the degradation parameters of the low-resolution multispectral image. Then, the panchromatic image is input into the panchromatic image noise parameter estimation subnetwork to estimate the noise parameters of the panchromatic image. Finally, the low-resolution multispectral image is upsampled and input together with the panchromatic image into the feature integration and image fusion subnetwork to obtain the fusion result.

[0240] The final fusion module is used to train a fusion network for multispectral and panchromatic images by combining a loss function, and to reconstruct high-resolution multispectral images using the trained fusion network.

[0241] The specific implementation methods of each module are the same as those of each step, and will not be described in this invention.

[0242] It should be understood that any parts not described in detail in this specification belong to the prior art.

[0243] It should be understood that the above description of the embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art can make substitutions or modifications under the guidance of this invention without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A method for fusion of blind multispectral and panchromatic images based on variational inference, characterized in that, Includes the following steps: Step 1: Model the multispectral and panchromatic image fusion problem using variational inference. Decompose the inverse problem of multispectral and panchromatic image fusion into two parts: one is the low-resolution multispectral image... M The part as observed variable and the high-resolution panchromatic image P As part of the observed variables, the set of latent variables in the degradation process is solved by combining the blurring and noise degradation of low-resolution multispectral images and the noise degradation of high-resolution panchromatic images. The problem to be solved in the set of latent variables is used as the loss function of the fusion network to guide the training of the fusion network. Among them, the set of latent variables for: in, The high-resolution multispectral image obtained by fusion is represented. Represents the noise variance in the multispectral image degradation process. Represents the blur kernel parameters in the multispectral image degradation process; The image gradient represents the fusion result; Represents the noise variance during the panchromatic image degradation process; Step 2: Design a multispectral and panchromatic image fusion network based on convolutional neural network and degradation parameter estimation. The fusion network VBPN consists of four sub-networks, including a multispectral image blur kernel estimation sub-network, a multispectral image noise parameter estimation sub-network, a panchromatic image noise parameter estimation sub-network, and a feature integration and image fusion sub-network. The fusion network takes low-resolution multispectral and panchromatic images with unknown degradation processes as input. The processing procedure of the multispectral image and panchromatic image fusion network is as follows: First, the low-resolution multispectral image is input into the multispectral image noise parameter estimation subnetwork and the multispectral image blur kernel estimation subnetwork to estimate the degradation parameters of the low-resolution multispectral image. Then, the panchromatic image is input into the panchromatic image noise parameter estimation subnetwork to estimate the noise parameters of the panchromatic image. Finally, the low-resolution multispectral image is upsampled and input together with the panchromatic image into the feature integration and image fusion subnetwork to obtain the fusion result. Step 3: Train a network for fusing multispectral and panchromatic images using a loss function, and then use the trained network to reconstruct high-resolution multispectral images.

2. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 1, characterized in that: From low-resolution multispectral images M The components that are observed variables include: The transition from low-resolution multispectral to high-resolution multispectral image can be understood as a multi-channel image super-resolution problem: in The high-resolution multispectral image obtained by fusion is represented. Represents the fuzzy kernel, Represents the convolution operation. Downsampling, representing spatial scale This represents additive white Gaussian noise; therefore, M The probability distribution is modeled as follows: To more accurately estimate the blur and noise parameters of multispectral images, they are modeled at the pixel level, where... represent M The j 1 pixel, Represents a Gaussian distribution. d This represents the total number of pixels in a low-resolution multispectral image. Represents the noise variance during the multispectral image degradation process; Specifically for , as well as Introducing priors, for We introduce the inverse Gamma distribution as its prior: in Represents the inverse Gamma distribution. For the multispectral image The noise variance of each pixel. To control the hyperparameters of the distribution shape, The noise parameters are estimated based on real high-resolution multispectral images and low-resolution multispectral images in the simulation dataset; For the fused high-resolution multispectral image, a Gaussian distribution is introduced as its prior distribution: in The first high-resolution multispectral image obtained by fusion i 1 pixel, for m This represents the total number of pixels in a high-resolution multispectral image. For the real high-resolution multispectral images in the simulation dataset, The hyperparameters are artificially defined to measure minute differences between the two; the fuzzy kernel K is considered to be an anisotropic Gaussian distribution. in Represents an anisotropic Gaussian distribution. The first in the multispectral image blur kernel Line 1 The value of the column, Represents the Pearson correlation coefficient. and These represent the noise variance in two directions, respectively. S The coordinates of the spatial points affected by the fuzzy kernel are given when the kernel size is 1. At that time, its range is When the fuzzy kernel size is determined, the covariance matrix can uniquely determine the fuzzy kernel. Therefore, the three variables in the covariance matrix are represented as follows: : but: in, Represents the blur kernel parameters in the multispectral image degradation process, where It is control and The hyperparameters of the difference between them; To control the hyperparameters of the inverse Gamma distribution, the parameters , and These are given during the dataset simulation process.

3. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 2, characterized in that: High-resolution panchromatic image P The components that are observed variables include: To model the relationship between panchromatic and multispectral images, we can consider the gradient perspective: in For gradient operators, The gradient of the panchromatic image. Represents the number of spectral bands. Represents additive white Gaussian noise. Represents a single band of the fused high-resolution multispectral image. The weighting parameter represents the corresponding waveband; set up: in, The image gradient representing the fusion result, For a given set of parameters; Similarly, modeling is performed at the pixel level, assuming that the gradient between the two follows a Gaussian distribution: in The number of pixels representing a panchromatic image is numerically the same as the number of pixels in a high-resolution multispectral image, therefore both are represented by [missing information - likely a typo]. m express, This represents the noise variance during the panchromatic image degradation process; therefore, the following prior forms are introduced for the distribution of the Gaussian distribution's mean and variance: in Represents the gradient of the real high-resolution multispectral image in the simulation dataset. This is a hyperparameter used to measure the difference between the two. To control the hyperparameters of the inverse Gamma distribution The noise variance parameter is estimated.

4. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 3, characterized in that: According to Bayes' theorem, the posterior distribution of the problem to be solved for the set of latent variables can be written as: Regulation In order to know the observed variables x Time variable The posterior distribution, Represents a probability distribution. q This represents the approximate probability distribution that needs to be estimated via a network. Since this posterior distribution is difficult to solve, based on the principle of variational inference, we hope to obtain an approximate posterior distribution. Approximating the true posterior distribution; the numerator of the posterior formula is rewritten as: Furthermore, based on the idea of ​​variational inference, the approximate posterior distribution... Written as: According to the mean-field theorem, the latent variables to be solved are considered to be independent of each other. Therefore, the two posterior distributions are expanded as follows: Based on the modeling of the degradation process, and according to the prior distribution of each latent variable to be solved, the approximate posterior distribution is given: in, For the distribution parameters of the latent variables that need to be estimated via a network, These are the hyperparameters given in the prior form of each latent variable; According to Bayes' theorem, the two observed variables... and The joint probability distribution is modeled in the following form: in Represents the lower bound of evidence. represent KL Divergence, used to measure the distance between two distributions, is expressed as follows: Due to two observed variables and joint probability distribution Since the given value is a constant that is difficult to solve, we minimize the two approximate posterior probability distributions based on the modeling form of the joint probability distribution. , The distance between the evidence and the corresponding prior is transformed into maximizing the lower bound of the evidence; The analytical solutions for each term in the lower bound of the evidence are given below: Using reparameterization and resampling techniques, we obtain: in, 、 、 from , and obtained by medium resampling, , ; in and These represent the Gamma and Digamma functions, respectively.

5. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 1, characterized in that: The multispectral image blur kernel estimation subnetwork is composed of a convolutional neural network with a residual channel attention module (CAB). It extracts feature information of each spectral channel from the low-resolution multispectral image and integrates it to obtain the feature information of the multispectral image blur kernel. The multispectral image noise parameter estimation subnetwork and the panchromatic image noise parameter estimation subnetwork are composed of the existing image denoising network DnCNN with the same input and output scale, and a pooling function is added at the network output to obtain the noise variance parameter; The feature integration and image fusion subnetwork consists of an encoder-decoder network with a feature integration layer. Both the encoder and decoder are composed of FE layers. There are image downsampling and upsampling operations between adjacent encoder and decoder layers to extract and integrate features of the input image from multiple scales, and finally obtain the fused high-resolution multispectral image.

6. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 1, characterized in that: The processing procedure of the multispectral image blur kernel estimation subnetwork is as follows: Firstly, through a Convolutional layers extract shallow features, which are then passed through an 8-layer cascaded CAB module to extract features from each spectral component. The final features are then processed by a... The convolutional layer yields the estimated probability distribution parameters of the blur kernel in a low-resolution multispectral image. , , ; The CAB module consists of convolutional layers, Leaky ReLU activation functions, and channel attention layers (CA layers). It takes shallow features of a multispectral image as input and processes them through a... The convolutional layer and the Leaky ReLU activation function are then passed through another... After further feature extraction by the convolutional layer, the features are input into the channel attention layer to extract features from each spectral channel. To facilitate network training, residual connections are added at the input and output of this module. The mathematical expression of the activation function is as follows: in: The parameter represents the slope of the negative half-axis of the activation function. x As the independent variable; Furthermore, assume that the features of the input CA Layer are The output features are The process of input feature transfer in CA Layer is as follows: Among them, F1 and F2 are intermediate variables. For average pooling operation, The activation function is expressed mathematically as follows:

7. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 1, characterized in that: The processing procedures of the multispectral image noise parameter estimation subnetwork and the panchromatic image noise parameter estimation subnetwork are as follows; The input image first passes through Convolutional layers with ReLU activation function are used to extract shallow features non-linearly, and these features are then input into a 5-layer cascade. Convolutional layers and Leaky ReLU activation function layers are ultimately passed through The convolutional layer obtains the estimated multispectral image noise distribution parameters. With panchromatic image noise distribution parameters , .

8. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 1, characterized in that: The processing procedure of the feature integration and image fusion subnetwork is as follows; The input image is processed by a The image is fed into the first encoder layer after the convolutional layer, and then encoded through several encoder layers. Skip connections are added between the encoder layer output and the decoder layer input at the same scale to integrate the encoder's feature information into the decoder, thereby improving the quality of image fusion. Finally, the decoder obtains the high-resolution multispectral image to be solved. probability distribution parameters This leads to the acquisition of high-resolution multispectral images. ; The existing feature integration layer FE Layer consists of two cascaded layers. The convolutional layer, combined with the Leaky ReLU activation function, is used to extract feature information from multispectral and panchromatic images.

9. The blind multispectral and panchromatic image fusion method based on variational inference as described in claim 4, characterized in that: The loss function used consists of five parts: the loss function of the multispectral image blur kernel estimation subnetwork is... The loss function of the multispectral image noise parameter estimation subnetwork is: The loss function of the panchromatic image noise parameter estimation subnetwork is: The loss function of the feature integration and image fusion subnetwork is: The loss function that serves to jointly constrain the training of the four sub-networks is... ; in , , and Each term is composed of the KL divergence term from the variational inference result in step one, and is used to guide the training of each sub-network. Their mathematical expressions are as follows: The loss function that jointly constrains the four subnetworks to achieve parameter transfer and joint optimization is... Composed of the expectation term derived from the variational inference in step one, its mathematical expression is: Ultimately, the overall loss function is: 。 10. A blind multispectral and panchromatic image fusion system based on variational inference, characterized in that, Includes the following steps: The fusion problem modeling module is used to model the multispectral and panchromatic image fusion problem using variational inference. It decomposes the inverse problem of multispectral and panchromatic image fusion into two parts: one is a low-resolution multispectral image... M The part as observed variable and the high-resolution panchromatic image P As part of the observed variables, the set of latent variables in the degradation process is solved by combining the blurring and noise degradation of low-resolution multispectral images and the noise degradation of high-resolution panchromatic images. The problem to be solved in the set of latent variables is used as the loss function of the fusion network to guide the training of the fusion network. Among them, the set of latent variables for: in, The high-resolution multispectral image obtained by fusion is represented. Represents the noise variance in the multispectral image degradation process. Represents the blur kernel parameters in the multispectral image degradation process; The image gradient represents the fusion result; Represents the noise variance during the panchromatic image degradation process; The fusion network construction module is used to design a multispectral and panchromatic image fusion network based on convolutional neural networks and degradation parameter estimation. The fusion network VBPN consists of four sub-networks, including a multispectral image blur kernel estimation sub-network, a multispectral image noise parameter estimation sub-network, a panchromatic image noise parameter estimation sub-network, and a feature integration and image fusion sub-network. The fusion network takes low-resolution multispectral and panchromatic images with unknown degradation processes as input. The processing procedure of the multispectral image and panchromatic image fusion network is as follows: First, the low-resolution multispectral image is input into the multispectral image noise parameter estimation subnetwork and the multispectral image blur kernel estimation subnetwork to estimate the degradation parameters of the low-resolution multispectral image. Then, the panchromatic image is input into the panchromatic image noise parameter estimation subnetwork to estimate the noise parameters of the panchromatic image. Finally, the low-resolution multispectral image is upsampled and input together with the panchromatic image into the feature integration and image fusion subnetwork to obtain the fusion result. The final fusion module is used to train a fusion network for multispectral and panchromatic images by combining a loss function, and to use the trained fusion network to reconstruct high-resolution multispectral images.

Citation Information

Patent Citations

  • Multispectral image fusion method based on adaptive spectrum-spatial gradient sparse regularization

    CN109859153A

  • Hyper-spectral and panchromatic image fusion method for extracting multi-scale spatial-spectral characteristics based on AE

    CN112634137A