Pan-sharpening Method Based on Multimodal Texture Correction and Adaptive Edge Detail Fusion
By introducing multimodal texture correction and adaptive edge detail fusion technology into the full color sharpening method, the problem of insufficient correlation and similarity between MS and PAN images is solved, and more accurate spatial detail extraction and better spectral and spatial information fusion are achieved, improving the full color sharpening effect.
Patent Information
- Application Number
- CN202411587725.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The existing full-color sharpening method has low correlation and similarity between MS and PAN images, resulting in inaccurate spatial details of the extracted, difficult to balance spectral and spatial information, resulting in spatial and spectral distortions in the fused image.
A full-color sharpening method based on multimodal texture correction and adaptive edge detail fusion is proposed. The multimodal texture correction model is used to establish intensity, gradient and depth correction constraints between the texture correction image and the source image, and the degradation filter is determined through the adaptive degradation filter algorithm, and the multimodal texture correction model is optimized to obtain the texture correction image. Then, the details information is extracted and fused through the adaptive edge detail fusion model and injected into the upsampled multispectral image to obtain the final high-resolution multispectral image.
The correlation and similarity between MS and PAN images are improved, the correlation and similarity between texture correction images and source images are enhanced, the spectral and spatial distortion of HRMS images are improved, and more accurate spatial information and better fusion effects are obtained.
Smart Images

Figure CN119515728B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image fusion, and particularly relates to a panchromatic sharpening method based on multimodal texture correction and adaptive edge detail fusion. Background Art
[0002] Due to the limitations of satellite imaging sensor hardware conditions, it is impossible to obtain multispectral images with both high spatial resolution and high spectral resolution simultaneously. However, spectral information-rich but low-spatial-resolution multispectral (MS) images can be obtained using spectral sensors, and panchromatic (PAN) images with high spatial resolution but poor spectral information can be obtained using spatial sensors. Therefore, panchromatic sharpening technology is adopted to improve the spatial resolution of low-resolution multispectral (LRMS) images. By fusing LRMS and PAN images and utilizing their respective advantages, high-resolution multispectral (HRMS) images can be finally obtained.
[0003] Panchromatic sharpening refers to the process of fusing multispectral (MS) images and panchromatic (PAN) images to obtain high-resolution multispectral (HRMS) images. However, due to the low correlation and similarity between MS and PAN images, as well as problems such as inaccurate spatial information injection, serious spectral and spatial distortions occur in HRMS images.
[0004] With the rapid development of panchromatic sharpening technology, it can be divided into four categories: methods based on component substitution (CS), methods based on multiresolution analysis (MRA), methods based on variational optimization (VO), and methods based on deep learning (DL). Methods based on CS can generally better retain spatial details, obtain high spatial quality, are easy to implement, and have high computational efficiency, but are prone to serious spectral distortion. Methods based on MRA can better retain spectral information, but the decomposition of the spatial structure is likely to cause spatial distortion. Methods based on VO can address the issues of spectral and spatial distortion in images, impose spectral prior constraints and spatial prior constraints between MS, PAN, and ideal HRMS images, and perform regularization prior constraint correction to construct a reasonable degradation model, and solve this model through an optimization algorithm. Methods based on VO generally better retain spatial and spectral information and obtain better fusion results. However, once unreasonable model assumptions are made, unpredictable deviations usually occur. Therefore, this type of method requires a more accurate mathematical model and further improvement in efficiency. Methods based on DL generally can achieve good fusion results, but they require a large number of images to train the network, consume a large amount of computing resources, and the test images are highly correlated with the training data. The parameters of the trained network are also fixed, which usually cannot adapt to other new datasets of different sensors, and the accuracy of methods based on DL cannot be further improved.
[0005] Currently, in all the above-mentioned pan-sharpening methods, there are problems of low correlation and similarity between the MS and PAN images, resulting in inaccurate extraction of spatial details and other information, and even only extracting spatial details from the PAN image. It is difficult to balance spectral and spatial information during the fusion process, resulting in spatial and spectral distortions in the fused image, and the final fused effect of the high-resolution multispectral image is not good enough. Even if methods based on deep learning can be used to balance spectral and spatial information, for example, supervised training networks will only be applicable to this dataset during testing, and frequent training on different datasets will lead to a sharp increase in costs such as training time. Summary of the Invention
[0006] To solve the above technical problems, the present invention proposes a pan-sharpening method based on multi-modal texture correction and adaptive edge detail fusion to solve the problems existing in the above-mentioned prior art.
[0007] To achieve the above object, the present invention provides a pan-sharpening method based on multi-modal texture correction and adaptive edge detail fusion, including:
[0008] Obtain a low-resolution multispectral image and a panchromatic image, fuse the upsampled low-resolution multispectral image and the panchromatic image to obtain a fused image, respectively extract the intensity components of the low-resolution multispectral image and the fused image, input the intensity components of the low-resolution multispectral image and the fused image and the panchromatic image into a multi-modal texture correction model, and optimize and solve the multi-modal texture correction model through an optimization method to obtain a texture-corrected image, wherein the multi-modal texture correction model is constructed based on a variational optimization model;
[0009] Extract details and apply edge protection to the texture-corrected image to obtain the first image details; extract details and apply edge protection to the upsampled low-resolution multispectral image to obtain the second image details; adaptively fuse the first image details and the second image details to obtain detail information, and add the detail information to the upsampled low-resolution multispectral image to obtain the final high-resolution multispectral image.
[0010] Optionally, perform linear weighted summation on each band image of the low-resolution multispectral image and perform linear weighted summation on each band image of the fused image to extract the intensity components of the low-resolution multispectral image and the fused image.
[0011] Optionally, fuse the upsampled low-resolution multispectral image and the panchromatic image through a pan-sharpening (A-PNN) model based on an object-adaptive convolutional neural network to obtain a fused image.
[0012] Optionally, the multi-modal texture correction model is:
[0013]
[0014] where T C is the texture correction image, D represents the downsampling matrix, H represents the degradation filter, I 0 represents the intensity component of the low-resolution multi-spectral image, α, β, γ, δ, θ represent the penalty parameters corresponding to different terms, is the Laplacian operator, P represents the panchromatic image, I net represents the intensity component of the fused image, ||·|| F represents the Frobenius norm, ||·|| 1 represents the 1-norm.
[0015] Optionally, the degradation filter H is obtained by an adaptive degradation filter algorithm, where the degradation filter H uses the Gaussian filter H A , and the adaptive degradation filter algorithm is:
[0016]
[0017] where DHT A T C = DF -1 (H A (u,v)F(T C ));
[0018] The frequency domain expression H A (u,v) of the Gaussian filter H A is:
[0019]
[0020] where D C (u,v) represents the distance from the point (u,v) to the center of the frequency domain, σ represents the standard deviation, and the optimal value of σ is obtained according to the correlation and similarity indexes. The optimal value of σ is σ best :
[0021]
[0022] where ρ(DHT A T C ,I 0 ) is the CC index between DHT A T C and I 0 , and S(DHT A T C ,I 0 ) is the SSIM index between DHT A TC The SSIM index between I 0 and I
[0023] Optionally, the multi-modal texture correction model is optimized and solved by the ADMM model.
[0024] Optionally, the process of extracting details from the texture correction image includes:
[0025]
[0026] where D TC is the image detail of the texture correction image, and T CL is the low-resolution version of the texture correction image.
[0027]
[0028] where χ 1 represents the weight coefficient, I UP represents the intensity component of the upsampled low-resolution multi-spectral image, and T CD represents the image obtained by processing the texture correction image with a Gaussian filter.
[0029]
[0030] where x 3 represents the normalized weight, x 1 represents the influence coefficient of I UP and x 2 represents the influence coefficient of T CD The value of x 1 is the mean of the correlation and similarity between T C and I UP The value of x 2 is the mean of the correlation and similarity between T C and T CD and I
[0031] Optionally, the process of adaptively fusing the first image detail and the second image detail includes:
[0032] Enhance the second image detail to the same level as the first image factor according to the scaling factor ξ:
[0033]
[0034] where F 2 represents the second image detail, F 3 represents the enhanced second image detail, and the superscript or subscript i represents the band label corresponding to the image.
[0035] Fuse the enhanced second image details with the first image details to obtain detail information F:
[0036]
[0037] where χ 2 is the weight coefficient, where x 1 represents the influence coefficient of I UP and the value of x 1 is the mean of the correlation and similarity between T C and I UP .
[0038] Optionally, the process of adding the detail information to the upsampled low-resolution multispectral image includes:
[0039]
[0040] where g represents the scaling factor for injecting details, M UP is the upsampled low-resolution multispectral image, B represents the total number of bands, i represents the band label, the superscript or subscript i represents the band label corresponding to the image, F represents the detail information, and M HR is the high-resolution multispectral image.
[0041] Optionally, the scaling factor g for injecting details is:
[0042]
[0043] where cov(.) is the covariance function, σ 2 is the variance function, T C represents the texture-corrected image, M UP is the upsampled low-resolution multispectral image, and the superscript or subscript i represents the band label corresponding to the image.
[0044] Compared with the prior art, the present invention has the following advantages and technical effects:
[0045] (1) To enhance the correlation and similarity between source images, a multimodal texture correction model is proposed. The intensity component of the LRMS image, the PAN image, and the intensity component of the image after A-PNN fusion are used as the input ends of the model, and the output end is the texture-corrected image. The model imposes an intensity correction constraint between the images, a gradient correction constraint between the texture-corrected image, the intensity component of the LRMS image, and the PAN image, and an A-PNN-based deep plug-and-play correction prior between the texture-corrected image and the intensity component of the image after A-PNN fusion.
[0046] (2) Since it is difficult to determine the degradation filter in the intensity correction constraint, in order to ensure the accuracy of establishing each constraint prior, an adaptive degradation filter algorithm is proposed. This algorithm can adaptively determine the degradation filter in the model, thereby enhancing the correlation and similarity between the texture-corrected image and the source image in the multi-modal texture correction model.
[0047] (3) In order to achieve the accuracy of spatial information injection, an adaptive edge detail fusion model is proposed. It adaptively extracts the detail information of the texture-corrected image and applies edge protection. Similarly, it extracts the detail information of the upsampled multi-spectral image and applies edge protection, and raises the spatial information of the upsampled multi-spectral image to the same level as the texture-corrected image. Finally, the spatial information of the texture-corrected image and the upsampled multi-spectral image is adaptively fused to obtain more accurate spatial information. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0049] Figure 1 is a flowchart of the method according to an embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of the iterative convergence result of the WorldView-3 dataset according to an embodiment of the present invention;
[0051] Figure 3 is a subjective evaluation fusion result diagram of the downscaled image in the QuickBird dataset according to an embodiment of the present invention;
[0052] Figure 4 is a subjective evaluation fusion result diagram of the downscaled image in the WorldView-2 dataset according to an embodiment of the present invention;
[0053] Figure 5 is a subjective evaluation fusion result diagram of the downscaled image in the WorldView-3 dataset according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0055] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0056] To solve these problems pointed out in the above technical background, the present invention proposes a panchromatic sharpening method based on multimodal texture correction and adaptive edge detail fusion. In order to obtain a texture-corrected image T that is highly correlated and similar to the multispectral (MS) image C , a panchromatic sharpening (A-PNN) fusion method based on a target adaptive convolutional neural network is introduced. By constructing a multimodal texture correction model, intensity, gradient, and A-PNN-based depth plug-and-play correction constraints are established between the texture-corrected image and the source image, and an algorithm for an adaptive degradation filter is proposed to ensure the accuracy of the establishment of these constraints. Since the obtained texture-corrected image can replace the panchromatic (PAN) image, and the MS image also contains partial spatial information, an adaptive edge detail fusion algorithm is proposed to adaptively extract the detail information of the texture-corrected image and the MS image respectively and apply edge protection. Since the MS image has less spatial information, its spatial information is enhanced proportionally and then adaptively fused. The fused spatial information is injected into the upsampled multispectral (UPMS) image to obtain the final HRMS image. A large number of experimental results show that compared with other methods, the algorithm proposed by the present invention has achieved better results in subjective visual effects and objective evaluation indicators and maintains a high operating efficiency.
[0057] Related work and related technical basis involved in the present invention:
[0058] Injection model:
[0059] The injection model is commonly used in panchromatic sharpening methods. It injects the spatial detail information with high spatial resolution in the panchromatic (PAN) image into the original upsampled multispectral (UPMS) image with high spectral resolution to generate the HRMS image, so as to solve the problem that the LRMS image lacks a large amount of spatial information. Among them, it is assumed that the size of the LRMS image is L×W×B (i.e., length×width×number of bands), and the size of the PAN image is L′×W′, where L′ = L / r, W′ = W / r, and r represents the compression ratio. Then the sizes of the upsampled multispectral (UPMS) and high-resolution multispectral (HRMS) images are L′×W′×B. The specific formula of the injection model can be uniformly expressed as
[0060] M HR = M UP + GS D (1)
[0061] where, M HR is the HRMS image, M UP is the UPMS image, G is the injection gain, and S D is the injected spatial detail information. Extract SD The methods can be uniformly divided into CS-based methods and MRA-based methods. For CS-based methods, S D can be extracted using the following formula.
[0062] S D = P I - I UP (2)
[0063] where P I represents the image obtained by performing histogram matching on the intensity component (I UP ) of the PAN and UPMS images. Histogram matching ensures that the intensity and contrast of the PAN and LRMS images are within the same grayscale value range, guaranteeing the accuracy of spatial information extraction. The formula for P I is as follows.
[0064]
[0065] where P represents the original PAN image, μ P and μ I represent the average values of the P and I UP images respectively, and σ P and σ I represent the variances of the P and I UP images respectively. I UP is obtained by linearly weighting each band of M UP , and its formula is as follows.
[0066]
[0067] where ω represents the linear weighting coefficient, the superscript or subscript i represents the i-th band of the image, and B represents the total number of bands. For MRA-based methods, S D can be extracted using the following formula.
[0068]
[0069] where P D represents the degraded PAN image, which can be obtained by applying a low-pass filter H LP to the PAN image P. H LP has a blurring effect on P.
[0070] However, problems such as inaccurate injected spatial detail information still exist. Since the missing spatial detail information in the LRMS image is generally inferred from the PAN image, problems such as inaccurate inference and possible misalignment of spectral information during the fusion process make it impossible to simultaneously maintain relatively accurate spectral fidelity and spatial fidelity, thereby resulting in spectral and spatial distortion in the fused image.
[0071] Variational optimization model:
[0072] The method of variational optimization has gradually become popular in recent years. It can ensure the accuracy of image spectral and spatial information by establishing a mathematical model. The established mathematical model can be regarded as a degradation model. The ideal HRMS image obtained by fusing this model is restored from the LRMS and PAN images, that is, the ideal HRMS image degenerates into an inverse process of the source image. Therefore, the method based on variational optimization can retain the spatial and spectral information of the LRMS and PAN images through various optimization algorithms and finally restore them into the required ideal HRMS image. To sum up, the method of variational optimization generally establishes an energy function E(M HR ). The energy function can be divided into three terms. The first term is the spectral fidelity term f spectral (M 0 , M HR ), the second term is the spatial fidelity term + f spatial (P, M HR ), and the third term is the regularization prior term + f prior (M HR ). The specific formula is as follows.
[0073] E(M HR ) = f spectral (M 0 , M HR ) + f spatial (P, M HR ) + f prior (M HR ) (6)
[0074] Among them, M 0 is the LRMS image. By performing blurring and downsampling operations on the ideal HRMS image H MR , M 0 can be obtained. M HR can also obtain the PAN image P through linear weighted combination. Therefore, the energy function in formula (6) can be simplified to the following common form.
[0075] E(M HR ) = λ 1 ||(DH LP M HR - M 0 )|| + ||P - CM HR || + λ 2 f prior (M HR ) (7)
[0076] Among them, λ 1 and λ 2 are penalty parameters, D represents the downsampling matrix, and C represents the linear weighted combination matrix. By optimizing and solving the above formula, M HR can be finally obtained.
[0077] Although the variational optimization method can retain relatively accurate spectral and spatial information simultaneously, it depends on the accuracy of the mathematical model establishment. An unreasonable variational optimization model will ignore the correlation and similarity between MS and PAN images, and the spectral and spatial information obtained from the two may be mismatched, which will further lead to spectral and spatial distortion in the final HRMS image. In addition, the efficiency of most variational optimization models is also relatively low.
[0078] The specific related method flow involved in the present invention will be described:
[0079] In order to solve the problems of poor correlation and similarity among LRMS, PAN, and HRMS images, and inaccurate spatial information of the injected UPMS image, a panchromatic sharpening method based on multimodal texture correction and adaptive edge detail fusion is proposed. This method can improve the spectral and spatial distortion of HRMS images.
[0080] The input end of the multimodal texture correction model is the intensity component I 0 of the LRMS image, the PAN image, and the intensity component of the image after A-PNN fusion, and the output end is the texture correction image T C . By establishing an intensity correction prior, the intensity constraint between I 0 and T C images is corrected. By establishing a gradient correction prior, the gradient constraint between I 0 , PAN and T C images is corrected. By establishing a depth plug-and-play correction prior based on A-PNN, the intensity gradient constraint between I net and T C images is corrected. These three correction priors form the basis of the multimodal texture correction model. In addition, an adaptive degradation filter algorithm is proposed, which can be used to obtain an accurate adaptive degradation filter H A in the intensity correction prior to degrade T C , so that the correlation and similarity between the degraded T C and I 0 images reach the highest. Finally, the multimodal texture correction model is optimized by ADMM to obtain the texture correction image T C . Since the texture correction image T CWith a high correlation and similarity to the source image, after keeping the spectral information of the LRMS image unchanged, it inherits the gradient information from the PAN image, and the intensity component I of the image after A-PNN fusion net has more image features and can further maintain the stability of texture information. Therefore, the texture-corrected image T C can be used to replace the PAN image for subsequent fusion operations.
[0081] After obtaining the texture-corrected image T C the texture-corrected image T and the multispectral MS image are fused through an adaptive edge detail fusion model C to generate the final HRMS image;
[0082] In the adaptive edge detail fusion model, since spatial detail information exists not only in the texture-corrected image T that replaces the PAN image C but also in the multispectral MS image. Therefore, the detail information of the texture-corrected image T is adaptively extracted and edge protection is applied, and the detail information in the upsampled multispectral (UPMS) image is extracted by using a Gaussian filter matching the modulation transfer function (MTF) and edge protection is applied. The detail information of the upsampled multispectral (UPMS) image with edge protection is enhanced to the same level as the texture-corrected image T C and is adaptively fused with the detail information of the texture-corrected image T with edge protection to obtain spatial information with a high correlation and similarity to the source image. The spatial information is injected into the UPMS image at an appropriate ratio to obtain the final HRMS image. The flowchart of the method of the present invention is as C shown. The specific process is shown in the following content. C Specifically, the multimodal texture correction model mainly includes an intensity correction prior term, a gradient correction prior term, and a depth plug-and-play correction prior term based on A-PNN; among them, the relevant filters in the intensity correction prior term and the gradient correction prior term are determined by the adaptive degradation filter algorithm, and the multimodal texture correction model is optimized and solved by the optimization model algorithm to obtain the final texture-corrected image. The specific content is as follows. Figure 1 shown. The specific process is shown in the following content.
[0083] Specifically, the multimodal texture correction model mainly includes an intensity correction prior term, a gradient correction prior term, and a depth plug-and-play correction prior term based on A-PNN; among them, the relevant filters in the intensity correction prior term and the gradient correction prior term are determined by the adaptive degradation filter algorithm, and the multimodal texture correction model is optimized and solved by the optimization model algorithm to obtain the final texture-corrected image. The specific content is as follows.
[0084] Intensity correction prior term:
[0085] Based on the spectral fidelity term in the variational optimization model in the above technical basis, the LRMS image can be obtained by blurring and downsampling the HRMS image. The specific formula is as follows:
[0086]
[0087] Among them, H is generally a Gaussian smoothing filter, and ‖·‖ F represents the Frobenius norm. In order to keep the inherent correlation and similarity between bands unchanged, the LRMS and ideal HRMS images of each band are linearly weighted and summed using Equation (4) to obtain I 0 and the intensity component I HR of the ideal HRMS image, and the specific formula is as follows.
[0088]
[0089] Since I HR is unknown, it is assumed that T C is close to I HR and is highly correlated. Therefore, the intensity correction prior term E intensity is as follows.
[0090]
[0091] Gradient correction prior term:
[0092] In the intensity correction prior model, T C maintains the invariance of spectral information. However, the spatial information should also be preserved. Based on the spatial fidelity term of the variational optimization model in the above technical basis, by establishing the spatial fidelity term, the gradient information of the PAN image is retained, and the specific formula is as follows.
[0093]
[0094] Among them, α is the penalty parameter. is the Laplacian operator. Since the process of correcting the gradient information of the PAN image by T C may cause a deviation in the intensity correction of T C and I 0 images, it is necessary to establish another spatial fidelity term to ensure that there is no deviation in the intensity correction during the gradient correction process, and further enhance the correlation and similarity between T C and I 0 images. The specific formula is as follows.
[0095]
[0096] Among them, β is the penalty parameter. To sum up, the gradient correction prior term can be expressed as follows.
[0097]
[0098] Depth Plug-and-Play Correction Prior Term Based on A-PNN:
[0099] To generate more texture features and further improve the C correlation and similarity between T 0 and I net and retain more spectral and spatial information, after fusing PAN and UPMS images through A-PNN, an HRMS image is obtained, denoted as MS net After linearly weighting MS net using Equation (4), the intensity component I net of MS C is obtained. A-PNN is a technique well-known to those skilled in the art and is a panchromatic sharpening method based on an object-adaptive convolutional neural network. Based on the spectral fidelity term of the variational optimization model in the above technical foundation, by establishing a spectral fidelity term between T net and I C to correct the intensity information of T
[0100]
[0101] where γ is a penalty parameter. By establishing a spatial fidelity term between T C and I net to correct the gradient information of T C the specific formula is as follows.
[0102]
[0103] where δ is a penalty parameter. In summary, the depth plug-and-play correction prior based on A-PNN can be expressed as follows.
[0104]
[0105] Multimodal texture correction model:
[0106] To ensure the sparsity of the output texture map and reduce artifacts, in addition to combining the above intensity correction prior term, gradient correction prior term, and depth plug-and-play correction prior term based on A-PNN, a total variation regularization term (TV) is also adopted. Among them, Therefore, the present invention proposes a multimodal texture correction model, and the specific formula is as follows.
[0107]
[0108] where θ is a penalty parameter.
[0109] Adaptive degradation filter algorithm:
[0110] In the model shown in Equation (17), in addition to the Gaussian filter H and the texture correction image T CEverything else has been determined, the texture-corrected image T C can be determined using Algorithm 2, while H is difficult to determine. Therefore, the present invention proposes an adaptive degradation filter algorithm, using a Gaussian filter as the degradation filter, defined as H A , which can be determined by the following formula.
[0111]
[0112] As can be seen from the above formula, when the texture-corrected image DH A T C and the intensity component I 0 of the LRMS image have the smallest difference, that is, when their correlation and similarity reach the highest, the H A at this time is the optimal degradation filter. Therefore, the adaptive degradation filter algorithm comprehensively considers their correlation and similarity, and uses the correlation coefficient (CC) and structural similarity (SSIM) indicators to measure, and finally adaptively determines the optimal degradation filter. When the filter processes the image in the spatial domain, the convolution operation will greatly increase the computational complexity. When processing the image in the frequency domain, the convolution operation is converted into an inner product operation, which will greatly reduce the computational complexity. Therefore, H A is selected to perform the operation in the frequency domain, and the frequency domain expression of H A is as follows.
[0113]
[0114] where D C (u, v) represents the distance from the point (u, v) to the center of the frequency domain, and σ represents the standard deviation. After H A is converted to the frequency domain, T C also needs to be calculated in the frequency domain. Therefore, the fast Fourier transform (FFT) is used to convert T C into the frequency domain, and the inverse fast Fourier transform (IFFT) is used to convert H A T C back to the spatial domain to facilitate the subsequent calculation of the correlation and similarity between DH A T C and I 0 . The specific formula is as follows.
[0115] DH A T C =DF -1 (H A (u, v)F(T C )) (20)
[0116] where F(.) represents the FFT operation, F -1(.) represents the IFFT operation. In summary, to determine H A , the key is to determine the unknown parameter σ. Therefore, by calibrating the correlation and similarity between DH A T C and I 0 , an optimal σ can be found. The correlation is measured by the CC index, denoted as ρ(DH A T C , I 0 ). The similarity is measured by the SSIM index, denoted as S(DH A T C , I 0 ). Using the average value rule to comprehensively consider these two indices, iterative processing is performed with different σ values, and the maximum value is taken for the final result. At this time, σ is the optimal value, denoted as σ best .
[0117] The specific formula is as follows.
[0118]
[0119] In summary, the overall process of the adaptive degradation filter algorithm is shown in Algorithm 1.
[0120] Algorithm 1:
[0121] Algorithm 1: Adaptive Degradation Filter Algorithm
[0122] Input: Texture-calibrated image T C , the intensity component I of the LRMS image 0
[0123] Initialization: Set σ (0) = 1, step size s = 0.5, iteration step k = 0,
[0124] Convert H A to the frequency domain through Equation (19),
[0125] Calculate (DH A T C ) (0) ,
[0126] Calculate ρ (0) (DH A T C , I 0 ) and S (0) (DH A T C , I 0 ).
[0127] Loop: σ (k+1) = σ (k) + s, k = k + 1,
[0128] Optimize (DH through formula (20) A T C ) (k+1) ,
[0129] Optimize ρ (k+1) (DH A T C , I 0 ), and S (k+1) (DH A T C , I 0 ),
[0130] Calculate through formula (21)
[0131] Until σ is satisfied (k+1) <σ (k) , the iteration ends and the loop is exited
[0132] Output: σ best =σ (k) , the adaptive degradation filter H A
[0133] Optimization model algorithm:
[0134] The optimization model algorithm uses ADMM to optimize, which is an optimization method that decomposes the original problem into multiple easy-to-handle sub-problems. To make the optimization process easier, different auxiliary variables are introduced That is, B = DA, then the model of formula (17) can be expressed as
[0135]
[0136] The augmented Lagrangian function of the above formula can be expressed as
[0137]
[0138] Among them, Λ 1 , Λ 2 , Λ 3 are different Lagrange multipliers, μ 1 , μ 2 , μ 3 are different penalty parameters. To minimize the energy function of the above formula, A is optimized by iteration (k+1) , B (k+1) , T C (k+1) , H A (k+1) , C (k+1) , Λ 1 (k+1) , Λ 2(k+1) and Λ 3 (k+1) Finally, we obtain T C , where k is the number of iterations, and the different parameters with superscript k + 1 represent the corresponding parameters in the (k + 1)-th iteration. The specific optimization process is as follows.
[0139] (1) Optimize A (k+1)
[0140] Fix other variables, and the sub-problem of A (k+1) is as follows.
[0141]
[0142] Set the deviation of A (k+1) to 0, that is Then A (k+1) is obtained by the following formula.
[0143]
[0144] where U represents the identity matrix, and the superscript T represents the transpose operator.
[0145] (2) Optimize B (k+1)
[0146] Fix other variables, and the sub-problem of B (k+1) is as follows.
[0147]
[0148] Set the deviation of B (k+1) to 0, that is However, due to the existence of the Laplace operator, the computational complexity increases during the solution process. To improve the computational efficiency, FFT and IFFT are adopted. After performing fast calculations in the frequency domain, it is then transformed back to the spatial domain. Therefore, when A (k+1) is optimized, B (k+1) can be obtained by the following formula.
[0149]
[0150] (3) Optimize T C (k+1)
[0151] Fix other variables, and the sub-problem of T C (k+1) is as follows.
[0152]
[0153] Set T C (k+1)The deviation is set to 0, that is Due to the existence of the Laplace operator, FFT and IFFT are also used for solving. Therefore, when A (k+1) After optimization, T C (k+1) Can be obtained from the following formula.
[0154]
[0155] (4) Optimize C (k+1)
[0156] Fix other variables, the sub-problem of C (k+1) Is as follows.
[0157]
[0158] Further simplify using the soft threshold (SoftThresholding) formula to obtain the following formula.
[0159]
[0160] Among them, sgn(.) is the signal function and max(.) is the maximum value function.
[0161] (5) Optimize the Lagrange multiplier Λ 1 (k+1) , Λ 2 (k+1) And Λ 3 (k+1)
[0162] Fix other variables and obtain Λ through the gradient ascent method 1 (k+1) , Λ 2 (k+1) And Λ 3 (k+1) Sub-problems.
[0163]
[0164] Among them, Is the step size required for gradient ascent, and the formula is as follows.
[0165]
[0166] Among them, τ is the penalty parameter. Generally, let τ>1, which can accelerate the convergence speed. To sum up, the overall process of optimizing the multi-modal texture correction model is shown in Algorithm 2. Among them, in the iteration process, H A (t+1) Is optimized by Algorithm 1. When T CWhen the relative change (RelCha) between two consecutive iterations is less than the tolerance deviation ε, the iterative process is terminated, and T is finally obtained. C Image. The relative change discrimination formula is as follows.
[0167]
[0168] As the iteration progresses, the relative change value RelCha gradually decreases. Therefore, it is necessary to determine the parameter ε, which is slightly larger than RelCha, to balance the efficiency and accuracy of the model. For example, Figure 2 shows the iterative convergence result of the test image in the WorldView-3 dataset. When the number of iterations reaches about 15 times, RelCha at this time tends to converge and is close to 1×10 -4 , that is, ε can be assigned a value of 1×10 -4 .
[0169] Algorithm 2:
[0170] Algorithm 2: Optimization Algorithm for Multimodal Texture Correction Model
[0171] Input: Panchromatic image P, intensity component I of LRMS image 0
[0172] Initialization: k = 0, Obtained through the initialization of Algorithm 1 τ = 1.01
[0173] While RelCha > ε do
[0174] Loop:
[0175] Optimize A through Equation (25) (k+1) ,
[0176] Optimize B through Equation (27) (k+1) ,
[0177] Optimize through Equation (29)
[0178] Optimize through Algorithm 1
[0179] Optimize C through Equation (31) (k+1) ,
[0180] Optimize through Equation (32)
[0181] k = k + 1.
[0182] Until RelCha <= ε is satisfied, the iteration ends and the loop is exited
[0183] Output: Texture-corrected image T C .
[0184] Adaptive edge detail fusion model:
[0185] Adaptive extraction of T C Image details and apply edge protection:
[0186] Obtain T through Algorithm 2 C After that, select the following formula to extract image details D TC .
[0187]
[0188] Among them, T CL is the low-resolution version of T C image. In order to extract details from T C more accurately, from equations (2) and (5), it can be seen that T CL can be obtained by two methods. The first method is that T C obtains the degraded image T by applying Algorithm 1 CD , similar to the method based on MRA to extract details, better retaining spectral information. The second method is to use equation (4) to obtain I UP , similar to the method based on CS to extract details, better retaining spatial information. Therefore, considering the advantages of these two methods comprehensively, adaptively extract D TC , and its algorithm design is as follows.
[0189] T CL = χ 1 I UP +(1 - χ 1 )T CD s.t. 0 < χ 1 < 1 (36)
[0190] Among them, χ 1 represents the weight coefficient to be determined. Since the accuracy of extracting details by the two methods is affected by the correlation and similarity between the source images, the influence coefficient of I UP can be set to x 1 , and the influence coefficient of T CD can be set to x 2 , and its formula is as follows.
[0191]
[0192] Since x1 and x2 do not satisfy the normalization constraint condition of χ 1 , χ 1 should be related to x 1 and x2 Is positively correlated and within a reasonable range, so χ can be obtained using the following formula 1 .
[0193]
[0194] Substitute χ in the above formula 1 into Equation (36) to obtain T CL . Then substitute it into Equation (35) to finally obtain D TC , completing the operation of adaptively extracting the details of the T C image. In order to retain edge information while extracting details, use the following edge detection matrix formula E TC to extract edges.
[0195]
[0196] Among them, is the gradient operator, η and are modulation coefficients. Generally, let η = 1×10 -9 , Therefore, the detailed information of the T C image with edge protection, that is, the first image detail F 1 is as follows.
[0197]
[0198] Extract the details of the UPMS image and apply edge protection:
[0199] Select the following formula to extract the detailed information D of the UPMS image M .
[0200]
[0201] Among them, M UPL represents the low-resolution version of the UPMS image. Since M UPL is unknown, introduce the MTF obtained from the MS sensor as an important indicator for extracting the details of the UPMS image. Therefore, use the Gaussian filter H MG with MTF matching to degrade the UPMS image to obtain the low-resolution version of the UPMS image. The specific process is shown in the following formula.
[0202]
[0203] Substitute the above formula into Equation (41) to obtain the detailed information of the UPMS image. At this time, it is necessary to use the edge detection matrix formula E M to perform edge protection on D M .
[0204]
[0205] Therefore, the detailed information of the UPMS image with edge protection, i.e., the second image detail F 2 is as follows.
[0206]
[0207] Adaptive edge detail fusion process:
[0208] After separately extracting T C and the detailed information of the UPMS image with edge protection, F can be fused 1 and F 2 . However, since the spatial resolution of the UPMS image is lower than that of T C , the detailed information contained in F 2 is less than that in F 1 . Directly fusing them may result in loss of detailed information. To avoid this, the information of F 2 is enhanced to the same level as F 1 before fusion. The specific formula is as follows.
[0209]
[0210] where ξ is the scale factor. The value of ξ is solved using the method of linear regression model. Therefore, the spatially enhanced information F 3 by ξ is expressed as.
[0211]
[0212] At this time, F 1 and F 3 can be adaptively fused to obtain the detailed information F. The specific algorithm is as follows.
[0213]
[0214] where χ 2 is the weight coefficient. The weight assignment of the detailed information is affected by the correlation and similarity between the source image, i.e., T C and the UPMS. Therefore, the relationship x C between T UP and I 1 can be set using Equation (37), while ensuring that χ 2 is within a reasonable range and x 1 is positively correlated with χ 2 . The specific formula is as follows.
[0215]
[0216] Substitute χ in the above formula 2 into Equation (47) to obtain the final detailed information F.
[0217] Final injection of detailed information at the spatial edge:
[0218] Substitute the detailed information F of Equation (47) into the following injection model to obtain the final HRMS image.
[0219]
[0220] where g represents the scaling factor for injecting details, which can be adaptively determined by the following formula.
[0221]
[0222] where cov(.) is the covariance function and σ 2 is the variance function.
[0223] Experimental design and conclusions for the above technical solutions:
[0224] To illustrate the performance advantages and effectiveness of the method proposed in the present invention, the proposed method (Proposed) is compared with 8 methods such as GSA, NIHS, BDSD-PC, SFIM, ATWT-W3, DMPIF, CDIF, and A-PNN, and a large number of experiments are carried out using 3 datasets such as QuickBird, WorldView-2, and WorldView-3. Each pair of images in each dataset contains one MS image and one PAN image. Among them, in the QuickBird dataset, the number of bands of the MS image is 4; while in the WorldView-2 and WorldView-3 datasets, the number of bands of the MS image is 8. The PAN image in all datasets only contains 1 band.
[0225] According to the Wald protocol, in this experiment, the original MS image is used as the reference image, that is, the ground truth (GT) image. At this time, it is necessary to perform 4-fold downsampling degradation processing on the original MS and PAN images respectively. The degraded images can be used as the source images for downscaling. The source images are fused using the algorithm proposed in the present invention, and the fused image is compared with the GT image. The smaller the gap, the better the effect. Therefore, in this experiment, the size of each band of the GT image is cropped to 256*256, the size of each band of the MS image is cropped to 64*64, and the size of the PAN image is cropped to 256*256.
[0226] The specific information of these 3 datasets in this experiment is summarized in Table 1, and Table 1 shows the detailed information of the datasets used in this experiment.
[0227] Table 1
[0228]
[0229]
[0230] To evaluate and compare the image quality of different methods, a combined subjective and objective evaluation criterion is adopted. Six commonly used objective evaluation indicators are used for objective evaluation. Among them, the Q2n index (Q4 represents the 4-band dataset, and Q8 represents the 8-band dataset) is selected to evaluate the spatial and spectral quality of the image, the peak signal-to-noise ratio (PSNR) is used to measure the error degree between the reconstructed image and the reference image, the universal image quality index (UIQI) is used to more comprehensively evaluate the quality difference and similarity between the fused image and the reference image, the relative average spectral error (RASE) is used to evaluate the average spectral difference before and after image fusion, the comprehensive dimensionless relative global error (ERGAS) is used to represent the distortion degree of the spatial and spectral information of the image, and the spectral correlation coefficient (SCC) is used to measure the retention ability of the spectral information of the image. Subjective evaluation is to visualize the fused MS image, extract the red (R), green (G), and blue (B) bands to display the true-color fused image, which can more intuitively reflect the quality difference of the image. Among the above evaluation indicators, the ideal values of Q2n, UIQI, and SCC are 1, while those of RASE and ERGAS are, and the ideal value of PSNR is +∞. All experiments in this section are run on a PC with an Inter Core i7-12700 CPU, a base speed of 2.10 GHz, and 32 GB of memory, and the experimental platform is MATLAB R2021b.
[0231] Comparative experiment:
[0232] QuickBird dataset:
[0233] In the subjective evaluation of the QuickBird dataset, as Figure 3As shown, the subjective evaluation fusion results of the method proposed in the present invention and each comparative method are presented, with the GT image used as the reference image. To more clearly display the spatial and spectral information of the image, the local fusion results are magnified. It can be seen from the magnified local area that the GSA method shows excessive details of the roof of the house. The images of the BDSD-PC and ATWT-M3 methods are relatively blurred and have a darker brightness. Although the SFIM method retains relatively good edge information in some areas, there is a problem of darker brightness. The CDIF method maintains good edge information but generates artifacts. Although the DMPIF and A-PNN methods retain relatively good spatial information, their edges generate excessive color information and the spectral distortion is relatively serious. The result of the method proposed in the present invention is the closest to the GT image, and it better retains the spatial and spectral information. The objective evaluation fusion results are shown in Table 2. Table 2 shows the objective evaluation fusion results of the downscaled images in the QuickBird dataset, and the ideal values are marked in parentheses. The black bold indicates the optimal result. It can be seen that compared with the other eight methods, the method proposed in the present invention has achieved the optimal results in all evaluation indicators, and the time spent by this method is shorter.
[0234] Table 2
[0235]
[0236] WorldView-2 dataset:
[0237] As Figure 4 shown, in the subjective evaluation of the fusion results in each comparative method in the WorldView-2 dataset. It can be seen from the magnified local area that the images in the GSA, BDSD-PC, and ATWT-M3 methods have poor clarity, resulting in relatively serious spatial distortion and darker colors. In the NIHS method, there is a problem of excessive injection of spatial information in some areas. Compared with the GT image, the SFIM method still has a certain gap in spatial information. In the CDIF method, the problems of image spatial distortion and spectral distortion are relatively serious. The clarity of the image of the DMPIF method has a gap compared with the GT image, and there are relatively serious artifact problems in the image. In the A-PNN method, the spectrum in some areas is distorted, and the spatial information retention is poor. The method proposed in the present invention is the closest to the GT image and has a better visual effect than other comparative methods. The objective evaluation fusion results are shown in Table 3. Table 3 shows the objective evaluation fusion results of the downscaled images in the WorldView-2 dataset. Obviously, compared with the other eight methods, the method proposed in the present invention has achieved the best results in all evaluation indicators, and the running time required is also shorter.
[0238] Table 3
[0239]
[0240] WorldView-3 dataset
[0241] As Figure 5 shown in the subjective evaluation of the fusion results of the WorldView-3 dataset for each method. It can be seen from the enlarged local area that in comparison with the GT image, the color of the roof of the house in the GSA method is darker. The image of the NIHS method produces certain artifacts, which affect the quality of the spatial information of the image. The images of the BDSD-PC, SFIM, and ATWT-M3 methods are relatively blurred. Although the CDIF method retains better spectral information, its retained detail information is poor, resulting in relatively serious spatial distortion. The DMPIF and A-PNN methods have a large color change, resulting in relatively serious spectral distortion. The method proposed by the present invention is closest to the GT image, and the subjective visual effect achieves the best result. The objective evaluation of the fusion results is shown in Table 4. Table 4 is the objective evaluation of the fusion results of the downscaled images in the WorldView-3 dataset. It can be seen that the results of the method of the present invention are the best among all evaluation indicators, and the running time is also shorter.
[0242] Table 4
[0243]
[0244] Related conclusions:
[0245] Since the MS and PAN images are obtained by different sensors, this pair of source images usually has low correlation and similarity. Direct fusion may lead to serious spectral distortion and spatial distortion. Secondly, to obtain an ideal HRMS image, the spatial information of the PAN image needs to be injected into the UPMS image. However, inaccurate injection of spatial information will result in a low spatial resolution of the HRMS image. To address these main problems, the present invention proposes a multi-modal texture correction and adaptive edge detail fusion model. To obtain T that is highly correlated and similar to MS and accurately inherits the PAN spatial information, by establishing an intensity constraint between T and I, a gradient constraint between T and PAN, I, and a depth plug-and-play constraint based on A-PNN between T and I, and an algorithm for an adaptive degradation filter is proposed to accurately maintain the constraints of the model. Finally, a multi-modal texture correction model is constructed. This model uses the ADMM algorithm to solve for T, which is used to replace the function of the PAN image. Since spatial detail information does not only exist in T C by establishing the intensity constraint between T C and I 0 the gradient constraint between T C and PAN, I 0 and the depth plug-and-play constraint based on A-PNN between T C and I net and an algorithm for an adaptive degradation filter is proposed to accurately maintain the constraints of the model. Finally, a multi-modal texture correction model is constructed. This model uses the ADMM algorithm to solve for T C to replace the function of the PAN image. Since spatial detail information does not only exist in TC In it, there is also some spatial detail information in the MS image. Therefore, an adaptive edge detail fusion model is proposed. This model extracts the detail information of the T C and UPMS images respectively and applies edge protection. In order to extract detail information more accurately, an adaptive extraction algorithm for T C is used to extract details, and a Gaussian filter matched with MTF is used to extract the UPMS image. The detail information with edge protection applied to T C is adaptively fused with the enhanced detail information with edge protection applied to the UPMS image. Finally, the fused spatial information is injected into the UPMS image using an injection model to obtain the final HRMS image. In the comparative experiment, the performance advantages of the algorithm of the present invention are illustrated, and the parameter analysis and ablation study prove the effectiveness of the algorithm of the present invention. The final result shows that the algorithm proposed by the present invention can obtain better fusion results.
[0246] In the multi-modal texture correction model, since the iterative optimization is carried out in a two-dimensional image, the solution efficiency is greatly improved, and the three correction prior terms set can better retain spatial and spectral information. However, this model still has deficiencies. There are unknown parameters in the correction prior terms that need to be determined through experiments, which may consume a large amount of computing resources and time. In the adaptive edge detail fusion model, in order to obtain accurate spatial information, the edge detail information of T C and UPMS is comprehensively considered. However, problems such as the quantity of the injected spatial information and the ratio of the spectral information of UPMS to the injected spatial information still exist. Therefore, the focus of our future work is to adaptively determine other unknown parameters in the pan-sharpening model and explore more suitable injection model methods to improve the overall performance and efficiency.
[0247] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A full color sharpening method based on multimodal texture correction and adaptive edge detail fusion, characterized in that: include: Acquire a low-resolution multispectral image and a panchromatic image, fuse the upsampled low-resolution multispectral image with the panchromatic image to obtain a fused image, extract intensity components of the low-resolution multispectral image and the fused image, respectively, input the intensity components of the low-resolution multispectral image and the fused image and the panchromatic image into a multimodal texture correction model, and optimize and solve the multimodal texture correction model by an optimization method to obtain a texture-corrected image, wherein the multimodal texture correction model is constructed based on a variational optimization model; Extract details and apply edge protection to the texture-corrected image to obtain details of a first image; extract details and apply edge protection to the upsampled low-resolution multispectral image to obtain details of a second image; adaptively fuse the details of the first image and the details of the second image to obtain detail information, and add the detail information to the upsampled low-resolution multispectral image to obtain a final high-resolution multispectral image; The multimodal texture correction model is: Among them, T C is the texture correction image, D represents the downsampling matrix, H represents the degradation filter, I0 represents the intensity component of the low-resolution multispectral image, α, β, γ, δ, θ represent the penalty parameters corresponding to different items, is the Laplace operator, P represents the full-color image, I net represents the intensity component of the fused image, ||·|| F represents the Frobenius norm, ||·||1 represents the 1 norm; The degradation filter H is obtained by an adaptive degradation filter algorithm, wherein the degradation filter H adopts a Gaussian filter H A , the adaptive degradation filter algorithm is: Among them, DH A T C =DF -1 (H A (u,v)F(T C )); F(.) represents FFT operation, F -1 (.) indicates IFFT operation; Gaussian filter H A The frequency domain expression H A (u,v) is: Among them, D C (u, v) represents the distance from point (u, v) to the center of the frequency domain, σ represents the standard deviation, and σ is optimally obtained based on the correlation and similarity indicators. The optimal value of σ is σ best : Among them, ρ(DH A T C ,I0) is DH A T C CC index between I0 and S(DH A T C ,I0) is DH A T C The SSIM indicator between I0 and I1.
2. The method according to claim 1, characterized in that Performing linear weighted summation on each band image of the low-resolution multispectral image and performing linear weighted summation on each band image of the fused image, and extracting intensity components of the low-resolution multispectral image and the fused image.
3. The method according to claim 1, characterized in that The upsampled low-resolution multispectral image is fused with the panchromatic image through a panchromatic sharpening model based on a target adaptive convolutional neural network to obtain a fused image.
4. The method according to claim 1, characterized in that The multimodal texture correction model is optimized and solved by the ADMM model.
5. The method according to claim 1, characterized in that The process of extracting details from the texture-corrected image includes: Among them, D TC Image details for texture-corrected images, T C represents the texture correction image, T CL A low-resolution version of the texture-corrected image, T CL =χ1I UP +(1-χ1)T CD s.t.0<χ1<1; Among them, χ1 represents the weight coefficient, I UP represents the intensity component of the upsampled low-resolution multispectral image, T CD Represents the image after the texture correction image is processed by the Gaussian filter; Among them, x3 represents the normalized weight, x1 represents I UP The influence coefficient of T CD The influence coefficient of x1 is T C and I UP The mean of the correlation and similarity, x2 value is T C and T CD The mean of the correlation and similarity.
6. The method according to claim 1, characterized in that The process of adaptively fusing the first image details and the second image details includes: The second image details are enhanced to the same level as the first image details according to the scaling factor ξ: Wherein, F2 represents the second image detail, F3 represents the enhanced second image detail, and the superscript or subscript i represents the band number corresponding to the image; The enhanced second image details are fused with the first image details to obtain detail information F: Among them, χ2 is the weight coefficient, Among them, x1 represents I UP The influence coefficient of x1 is T C and I UP The mean of the correlation and similarity, F1 represents the first image detail.
7. The method according to claim 1, characterized in that The process of adding the detail information to the upsampled low-resolution multispectral image includes: Where g represents the scaling factor of the injected details, M UP is the upsampled low-resolution multispectral image, B represents the total number of bands, i represents the band number, the superscript or subscript i represents the band number corresponding to the image, F represents the detail information, and M HR High-resolution multispectral images.
8. The method according to claim 1, characterized in that: The scaling factor g for injected detail is: Among them, cov(.) is the covariance function, σ 2 is the variance function, T C represents the texture correction image, M UP is the upsampled low-resolution multispectral image, and the superscript or subscript i represents the band number corresponding to the image.
Citation Information
Patent Citations
Video frame processing method, device and equipment and storage medium
CN110191340A
Remote sensing image fusion method and system combining spectrum and space dual-scale detail injection
CN118505527A