Pan-sharpening method based on progressive expansion framework
Through the full-color sharpening method based on the progressive expansion framework, it is decomposed into a three-stage task, combined with variational optimization and deep learning, the problems of spectral feature distortion and spatial detail loss are solved, and high-quality multi-spectral image fusion is achieved to adapt to complex scenes and reduce feature loss.
Patent Information
- Application Number
- CN202510838799.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing full-color sharpening technology has problems of spectral feature distortion and spatial details loss. Traditional methods lack flexibility. Deep learning methods rely on training data and are complex in calculations. Direct upsampling leads to feature loss and it is difficult to adapt to local transformation.
The full-color sharpening method based on the progressive expansion framework is adopted to decompose the full-color sharpening task into a three-stage progressive multi-spectral image recovery task. Combined with variational optimization and deep learning, the multi-scale local cross-attention module is used to guide the details recovery and spectral feature retention of multi-spectral images.
It improves the quality of the fusion image, enhances the interpretability and transparency of the network, can adapt to complex scenarios, effectively balance spectral information and spatial details, and reduces the information loss caused by direct upsampling.
Smart Images

Figure CN120339123B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of panchromatic sharpening, and in particular relates to a panchromatic sharpening method based on a progressive expansion framework. Background Art
[0002] With the advancement of remote sensing technology, multispectral images have been widely used in tasks such as land cover classification, target identification, and target detection. However, due to inherent physical hardware limitations, high-resolution multispectral images cannot be directly acquired. In order to balance spectral and spatial characteristics, satellites are usually equipped with two different types of sensors, one to capture a low-resolution multispectral image and the other to capture a high-resolution single-band panchromatic image of the same scene. The panchromatic sharpening task has emerged to fuse low-resolution multispectral images with high-resolution panchromatic images and generate a high-resolution multispectral image with rich spatial details and accurate spectral information. This high-quality multispectral image can significantly improve the accuracy of downstream tasks.
[0003] At present, the methods for solving the panchromatic sharpening task are mainly divided into four categories, including three traditional methods and one deep learning method, as follows:
[0004] The first category is component replacement methods, whose main idea is to map the upsampled low-resolution multispectral image to a transform domain, replace the intensity component with the panchromatic image, and then use the inverse transform to obtain a high-resolution multispectral image. This method effectively recovers spatial information, but is prone to severe spectral distortion.
[0005] The second category is multiresolution analysis methods. The main idea is to use multiscale decomposition or spatial filtering to obtain spatial details, and then inject these details into the upsampled low-resolution multispectral image to obtain the target image. This method tends to preserve color information, but some spatial details may be lost due to repeated transformations.
[0006] The third category is variational optimization methods. Their main idea is to treat the pan-sharpening task as an ill-posed problem, using a degradation model and prior constraints to construct an energy equation to iteratively optimize the quality of the target image. This method is conducive to finding global optimality, but its priors are often artificially set, making it difficult to adapt to various complex scenarios.
[0007] The fourth category is deep learning methods, which utilize frameworks such as convolutional neural networks, Transformers, and generative adversarial networks to adaptively mine features from source images and generate high-quality target images. These methods offer powerful nonlinear fitting and fast inference capabilities, but lack interpretability and are highly dependent on the quantity and quality of training data.
[0008] Despite significant advances in models in the field of pan-sharpening, the state-of-the-art still suffers from three major shortcomings.
[0009] First, traditional methods suffer from varying degrees of spectral distortion and loss of spatial detail. While deep learning methods can better balance spectral and spatial information, they suffer from black-box characteristics and rely heavily on the quality and quantity of training data. If the distribution of test images does not match the training data, the model will experience severe performance degradation.
[0010] Secondly, both traditional and deep learning methods typically employ a preprocessing step of directly upsampling low-resolution multispectral images by a factor of four to match the spatial resolution of the panchromatic image before subsequent fusion. However, this oversimplification overlooks the significant feature loss during direct upsampling of low-resolution multispectral images, which in turn impacts subsequent operations.
[0011] Finally, when using panchromatic images to supplement the details of multispectral images, existing deep learning models typically use fixed convolution kernels and global attention mechanisms. Fixed convolution kernels result in a lack of flexibility in the model's processing of varying spatial details, potentially failing to adapt to local transformations and making it difficult to accurately capture detailed information from different regions within the image. While global attention mechanisms can improve the utilization of contextual information and enhance global relationships, they are computationally complex and can lead to excessive attention to irrelevant or redundant regions. Using global attention alone makes it difficult to capture the features and details of small regions. Therefore, existing technologies still have room for improvement in adaptive local fusion.
[0012] To solve the above problems, the present invention proposes a panchromatic sharpening method based on a progressive unfolding framework, which decomposes the panchromatic sharpening task into a three-stage progressive multispectral image restoration task guided by panchromatic images. Summary of the Invention
[0013] To solve the above technical problems, the present invention proposes a panchromatic sharpening method based on a progressive expansion framework, which can solve the problems of spectral feature distortion and insufficient spatial details of the fused image, and improve the quality of the fused image.
[0014] The present invention provides a panchromatic sharpening method based on a progressive expansion framework, comprising:
[0015] Get the image to be optimized;
[0016] The image to be optimized is input into a variational optimization model to obtain a clear multispectral image, wherein the variational optimization model is obtained by training a training set, and the training set includes a multispectral image and a panchromatic image. The variational optimization model is used to utilize multispectral images at different scales to perform mutual modulation with corresponding panchromatic images, supplement the detailed information of the multispectral image and retain the spectral characteristics, and obtain the clear multispectral image.
[0017] Optionally, obtaining the training set includes:
[0018] Acquire original multispectral images and panchromatic images;
[0019] Downsampling the original multispectral image and the panchromatic image to obtain a processed image;
[0020] The processed image is segmented to obtain the training set.
[0021] Optionally, the variational optimization model includes: a plurality of prior learning modules and a plurality of image update modules;
[0022] The prior learning module is used to limit the restoration direction of the multispectral image, use the panchromatic image to guide the high-frequency detail restoration and low-frequency structure enhancement of the multispectral image, and denoise the enhanced multispectral image;
[0023] The image updating module is used to update the multispectral image.
[0024] Optionally, before obtaining a clear multispectral image, the following steps may also be performed:
[0025] Acquire a full-color image, and perform spectral super-resolution on the full-color image using a convolution and activation function to obtain a first feature map;
[0026] Performing wavelet sampling on the first feature map to obtain a second feature map and a first high-frequency image;
[0027] Perform wavelet sampling on the second feature map to obtain a third feature map and a second high-frequency image.
[0028] Optional, clear multispectral image acquisition includes:
[0029] acquiring a first blurred multispectral image;
[0030] Performing a first scale transformation on the first blurred multispectral image in combination with the third feature map to obtain a first clear multispectral image;
[0031] The first clear multispectral image is combined with the second high-frequency image to perform high-frequency modulation and inverse wavelet transform to obtain a second blurred multispectral image;
[0032] Performing a second scale transformation on the second blurred multispectral image in combination with the second feature map to obtain a second clear multispectral image;
[0033] The second clear multispectral image is combined with the first high-frequency image to perform high-frequency modulation and inverse wavelet transform to obtain a third blurred multispectral image;
[0034] The third blurred multispectral image is combined with the first feature map to perform a third scale transformation to obtain the clear multispectral image.
[0035] Optionally, performing a scale transformation on the blurred multispectral image combined with the feature map includes:
[0036] The blurred multispectral image is sequentially input into the prior learning module and the image updating module for multiple iterations, with the number of iterations in each stage set to 3, to obtain a clear multispectral image.
[0037] Optionally, the blurred multispectral image is sequentially input into a priori learning module and an image updating module to obtain a clear multispectral image as follows:
[0038] ;
[0039] ;
[0040] in, For the iterations, For the Scaling process ( =1, 2, 3), Auxiliary variables to help update multispectral images, For the prior The proximal operation of , here the prior learning module is used to implicitly implement the whole process, and is the step length, For clear multispectral images, is the blur kernel, is the transpose of the blur kernel, is a blurred multispectral image, is the scale parameter.
[0041] Compared with the prior art, the present invention has the following advantages and technical effects:
[0042] 1. This invention cleverly combines the advantages of traditional variational optimization methods and deep learning methods, not only enhancing the interpretability and transparency of the network, but also improving the algorithm's fusion performance. The invention designs a progressive fusion strategy based on wavelet transform, which gradually increases the details of the multispectral image at three spatial resolutions, alleviating the information loss problem caused by direct upsampling of low-resolution multispectral images. The invention proposes a priori learning module with multi-scale local cross-attention as its core, which uses full-color images to effectively supplement the detailed information of the multispectral image and preserve the spectral characteristics.
[0043] 2. This invention effectively improves the quality of fused images on simulated datasets, demonstrating superior spectral information preservation and spatial detail enhancement. Furthermore, it demonstrates sufficient generalization in practical applications of real-world images, adapting to complex real-world scenarios and effectively balancing spectral and spatial information in diverse scenes, including vegetation, roads, oceans, and buildings. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0045] Figure 1 is a flow chart of a pan-sharpening method based on a progressive expansion framework according to an embodiment of the present invention;
[0046] Figure 2 Schematic diagram of a network framework according to an embodiment of the present invention;
[0047] Figure 3 Schematic diagram of experimental results of an embodiment of the present invention. DETAILED DESCRIPTION
[0048] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0049] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0050] This embodiment proposes a pan-sharpening method based on a progressive expansion framework, such as Figure 1 As shown, the specific steps include:
[0051] Get the image to be optimized;
[0052] The image to be optimized is input into the variational optimization model to obtain a clear multispectral image. The variational optimization model obtains the optimal parameters through training with a training set. The training set includes multispectral images and panchromatic images. The variational optimization model is used to mutually modulate the multispectral images at different scales with the corresponding panchromatic images, supplement the detailed information of the multispectral image and retain the spectral characteristics to obtain a clear multispectral image.
[0053] Furthermore, obtaining a training set includes:
[0054] Acquire original multispectral images and panchromatic images;
[0055] Downsampling the original multispectral image and the panchromatic image to obtain a processed image;
[0056] Segment the processed image to obtain the training set.
[0057] Specifically, this example uses the WorldView-3 dataset. The original multispectral and panchromatic images are downsampled fourfold to generate simulated data. The original multispectral images serve as reference images for loss calculation. The low-resolution multispectral images are then randomly sliced into 16×16 images, and the corresponding panchromatic and reference images are sliced into 64×64 images to generate training data.
[0058] After obtaining the processed training data set, it is input into the network in batches for training and parameter update. The network loss function uses the mean absolute error (MAE), which can reduce the sensitivity to noise. The formula is as follows:
[0059] (1);
[0060] in, is the mean absolute error loss function, This is the final clear multispectral image output at the third scale. is the corresponding reference image.
[0061] When the loss no longer decreases, the converged network parameters are saved, the fusion test is performed on unknown data, and a reasonable evaluation is performed through quantitative indicators and qualitative visual evaluation.
[0062] Furthermore, the variational optimization model includes: several prior learning modules and several image update modules;
[0063] A priori learning module is used to limit the restoration direction of the multispectral image, use the panchromatic image to guide the high-frequency detail restoration and low-frequency structure enhancement of the multispectral image, and denoise the enhanced multispectral image;
[0064] The image updating module is used to update the multispectral image.
[0065] Specifically, the network framework constructed in this embodiment is based on the iterative solution process of the variational optimization model. Other model-driven deep learning methods assume that low-resolution multispectral images are obtained by blurring and downsampling high-resolution multispectral images. Different from the existing assumptions, in order to reduce the information loss caused by upsampling of low-resolution multispectral images, this embodiment decomposes the fusion task at a single spatial resolution into image restoration tasks at three spatial resolutions. It is assumed that the blurred multispectral image at each scale is restored to a clear multispectral image under the guidance of the panchromatic image. Assume that different scales are represented as , in On a scale, the full-color image is represented as , the blurred multispectral image is represented as , a clear multispectral image is represented as , the Gaussian blur kernel is expressed as , we can get the following variational optimization energy equation:
[0066] (2);
[0067] in, represents the denoising prior for multispectral images guided by panchromatic images, In traditional variational optimization methods, this prior is manually designed and relies on limited domain prior knowledge, making it difficult to adapt to complex real-world scenarios. Therefore, in order to improve the prior representation capability, this embodiment uses a deep learning method to adaptively learn complex prior knowledge from training data. The energy equation can be solved by the semi-quadratic splitting method, by introducing auxiliary variables , decompose the equation into the auxiliary variable estimation process and the multispectral image update process, where the auxiliary variable estimation process is also called the prior learning process, and the obtained equation is as follows:
[0068] ;
[0069] in, represents the scale parameter, Indicates the iterations, assuming and Represents the step size, and the specific iteration formula is as follows:
[0070] ;
[0071] ;
[0072] Furthermore, before obtaining a clear multispectral image, the following steps are also required:
[0073] Obtain a full-color image, perform spectral super-resolution on the full-color image using convolution and activation functions, and obtain a first feature map;
[0074] Perform wavelet sampling on the first feature map to obtain a second feature map and a first high-frequency image;
[0075] The second feature map is subjected to wavelet sampling to obtain a third feature map and a second high-frequency image.
[0076] Furthermore, obtaining clear multispectral images includes:
[0077] acquiring a first blurred multispectral image;
[0078] Performing a first scale transformation on the first blurred multispectral image in combination with the third feature map to obtain a first clear multispectral image;
[0079] The first clear multispectral image is combined with the second high-frequency image to perform high-frequency modulation and inverse wavelet transform to obtain a second blurred multispectral image;
[0080] Performing a second scale transformation on the second blurred multispectral image in combination with the second feature map to obtain a second clear multispectral image;
[0081] The second clear multispectral image is combined with the first high-frequency image to perform high-frequency modulation and inverse wavelet transform to obtain a third blurred multispectral image;
[0082] The third blurred multispectral image is combined with the first feature map to perform a third scale transformation to obtain a clear multispectral image.
[0083] Furthermore, performing scale transformation on the blurred multispectral image combined with the feature map includes:
[0084] The blurred multispectral image is sequentially input into the prior learning module and the image updating module for multiple iterations to obtain a clear multispectral image.
[0085] Specifically, the network invented in this embodiment is divided into three major stages according to three scales: Each major stage includes The multispectral image is iteratively restored using small stages. Each stage contains a prior learning module and a target image update module, corresponding to formulas (3) and (4) respectively. The network is a bidirectional network. The high-resolution full-color image is first spectrally super-resolved using convolution and activation functions. The obtained feature map is used to guide Image restoration, the feature map is downsampled twice by wavelet transform, retaining the high-frequency feature maps of the middle two scales, and the low-frequency feature maps of the two scales are used for and Prior guidance learning. At the end of each large stage, after obtaining the multispectral image with enhanced details at that scale, the details of the image are used to interactively modulate the high frequencies of the three panchromatic images at that scale, so that the four feature maps are flexibly aligned and then up-sampled by inverse wavelet transform, thereby improving the ability to retain the spectrum and reducing detail redundancy. In order to better capture local details, this embodiment proposes a priori module with multi-scale local cross attention as the core, generating a local adaptive kernel from the fusion features of the multispectral image and the panchromatic image, which is used to supplement the detail information of the multispectral image and retain the spectral features, and realizes multi-scale attention enhancement by combining adaptive kernels of different sizes. The specific network framework is as follows: Figure 2 shown.
[0086] The following is combined with Figure 3 This embodiment is described in detail:
[0087] This example is trained and tested on the WorldView-3 dataset, with a training batch size of 8, a learning rate of 0.0005, and 9714 training image pairs. The experimental environment is Windows 11, a 3.4GHz AMD Ryzen 5950X, an NVIDIA GeForce RTX 3090 GPU, and Python 3.9. It is compared with nine other full color sharpening algorithms, and the experimental results are shown in the figure below. Figure 3 As shown in FIG. 1 , it can be seen that only this embodiment successfully retains the red color and clear edge details in the reference image, indicating the effectiveness of this embodiment in improving the quality of the fused image.
[0088] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A panchromatic sharpening method based on a progressive expansion framework, characterized in that: include: Get the image to be optimized; Inputting the image to be optimized into a variational optimization model to obtain a clear multispectral image, wherein the variational optimization model is obtained by training a training set, the training set includes a multispectral image and a panchromatic image, and the variational optimization model is used to mutually modulate the multispectral images at different scales with the corresponding panchromatic image to supplement the detail information of the multispectral image and retain the spectral characteristics to obtain the clear multispectral image; Before obtaining a clear multispectral image, the following steps must be taken: Acquire a full-color image, and perform spectral super-resolution on the full-color image using a convolution and activation function to obtain a first feature map; Performing wavelet sampling on the first feature map to obtain a second feature map and a first high-frequency image; Perform wavelet sampling on the second feature map to obtain a third feature map and a second high-frequency image; Acquiring clear multispectral images includes: acquiring a first blurred multispectral image; Performing a first scale transformation on the first blurred multispectral image in combination with the third feature map to obtain a first clear multispectral image; The first clear multispectral image is combined with the second high-frequency image to perform high-frequency modulation and inverse wavelet transform to obtain a second blurred multispectral image; Performing a second scale transformation on the second blurred multispectral image in combination with the second feature map to obtain a second clear multispectral image; The second clear multispectral image is combined with the first high-frequency image to perform high-frequency modulation and inverse wavelet transform to obtain a third blurred multispectral image; The third blurred multispectral image is combined with the first feature map to perform a third scale transformation to obtain the clear multispectral image.
2. The method for pan-sharpening based on a progressive expansion framework according to claim 1, wherein: Obtaining the training set includes: Acquire original multispectral images and panchromatic images; Downsampling the original multispectral image and the panchromatic image to obtain a processed image; The processed image is segmented to obtain the training set.
3. The method for pan-sharpening based on a progressive expansion framework according to claim 1, wherein: The variational optimization model includes: a plurality of prior learning modules and a plurality of image update modules; The prior learning module is used to limit the restoration direction of the multispectral image, use the panchromatic image to guide the high-frequency detail restoration and low-frequency structure enhancement of the multispectral image, and denoise the enhanced multispectral image; The image updating module is used to perform an iterative update on the multispectral image.
4. The method for pan-sharpening based on a progressive expansion framework according to claim 1, wherein: The scale transformation of the blurred multispectral image combined with the feature map includes: The blurred multispectral image is sequentially input into the prior learning module and the image updating module for multiple iterations, with the number of iterations in each stage set to 3, to obtain a clear multispectral image.
5. The method for pan-sharpening based on a progressive expansion framework according to claim 4, wherein: The blurred multispectral image is sequentially input into the prior learning module and the image updating module to obtain a clear multispectral image as follows: ; ; in, For the iterations, For the Scaling process ( =1, 2, 3), Auxiliary variables to help update multispectral images, For the prior The proximal operation of , here the prior learning module is used to implicitly implement the whole process, and is the step length, For clear multispectral images, is the blur kernel, is the transpose of the blur kernel, is a blurred multispectral image, is the scale parameter.
Citation Information
Patent Citations
Multi-spectral image sharpening method, device and equipment and storage medium
CN111275632A
Remote sensing image fusion method based on NSST and parameter adaptive PCNN
CN114897757A