A method for panchromatic sharpening based on deep coupled feedback network

By deeply fusing features from PAN and MS images through a deep coupled feedback network (PSCF-Net), the problem of insufficient preservation of spectral and spatial information in existing technologies is solved, generating high-quality hyperspectral and spatial resolution images.

CN116309115BActive Publication Date: 2026-07-24ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2023-02-06
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously preserve the spectral and spatial information of satellite images. Traditional methods suffer from spectral or spatial domain distortion, while deep learning-based methods fail to fully utilize image features, resulting in limited fusion performance.

Method used

A deep coupled feedback network (PSCF-Net) is adopted. By constructing two CFB sub-networks, low-level and high-level features of PAN image are extracted respectively and fused with MS image features. Multiple sets of CFB modules are used for feature fusion, and finally a fused image with high spectral and spatial resolution is output.

Benefits of technology

It achieves effective fusion of features from PAN and MS images in the feature domain, preserves spectral information and accurately extracts spatial information, and generates high-quality hyperspectral and spatial resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309115B_ABST
    Figure CN116309115B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on depth coupling feedback network's full color sharpening method, it is related to image fusion technical field, including the following steps: obtaining PAN image and MS image;Depth coupling feedback network is constructed;PAN image and MS image are input to depth coupling feedback network, and finally output is fused image;Depth coupling feedback network includes: input module, receives PAN image and MS image;Feature extraction module includes three FEB modules, for extracting the low-level feature and high-level feature of PAN image and the feature of MS image;Two CFB sub-networks include multiple groups of symmetrical CFB modules, for the multiple features extracted are fused;Output module, output fused image.PSCF-Net of the application can realize the deep feature fusion of PAN image and MS image in feature domain, can obtain the image with high spectral and spatial resolution simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image fusion technology, and in particular to a full-color sharpening method based on a deep coupled feedback network. Background Technology

[0002] Earth observation satellites accurately observe and monitor Earth's resources and environment, and the large-scale observational data they acquire can be used in many fields such as agriculture, forestry, water conservancy, surveying and mapping, and transportation. Although the spatial and spectral resolution of optical images obtained by satellites is constantly improving, due to the advent of satellite sensors, we cannot simultaneously obtain images with high spectral and spatial resolution. Therefore, we need a method to address this problem, a method that generates images with both high spectral and spatial resolution. Panchromatic sharpening fuses the spatial information of panchromatic (PAN) images with the spectral information of multispectral (MS) images to generate an image containing both types of information. Therefore, panchromatic sharpening is a feasible solution that can bridge the gap between practical needs and technical limitations.

[0003] Pancolor sharpening is mainly divided into traditional methods and deep learning (DL)-based methods. Traditional methods mainly include component substitution (CS), multi-resolution analysis (MRA), and variational optimization (VO). CS methods typically replace the spatial components of the MS image with the spatial components of the PAN image, which has rich spatial information. MRA methods generally use multi-resolution methods to inject the spatial information extracted from the PAN image into the MS image. VO methods use some prior knowledge to represent the relationship between PAN, MS, and high-resolution MS (HRMS) images using a model, and then use variational optimization to process the objective function defined by the model. DL-based methods mainly train a model on a large amount of remote sensing data to obtain a model that can map MS and PAN images to ideal HRMS images. These models mainly include convolutional neural networks (CNNs) and generative adversarial networks (GANs).

[0004] While CS methods typically generate MS images with accurate spatial information, they sometimes suffer from severe spectral distortion. Compared to CS methods, the MRA series performs better in spectral preservation but is more sensitive to registration errors, which can lead to severe spatial domain distortion. Therefore, CS- and MRA-based methods cannot simultaneously preserve spectral and spatial information. VO methods, compared to these two approaches, are more time-consuming and require precise prior knowledge. In general, traditional methods have insufficient representational power, and their fusion performance is often limited. Furthermore, when assumptions are inconsistent with a specific dataset, the aforementioned algorithms can lead to severe quality degradation. Existing DL-based methods do not fully utilize the different levels of features in PAN images to achieve deep feature domain fusion, and the loss of spatial information still needs to be reduced. Summary of the Invention

[0005] This invention provides a full-color sharpening method based on a deeply coupled feedback network, which can solve the problems existing in the prior art.

[0006] This invention provides a pancolor sharpening method based on a deeply coupled feedback network, comprising the following steps:

[0007] Acquire PAN and MS images;

[0008] Construct a deeply coupled feedback network;

[0009] The PAN image and MS image are input into a deep coupled feedback network, and the final output is a fused image.

[0010] The deeply coupled feedback network includes:

[0011] The input module receives PAN and MS images;

[0012] The feature extraction module includes three FEB modules for extracting low-level and high-level features of PAN images and features of MS images;

[0013] Two CFB subnetworks are used. One CFB subnetwork is used to fuse the low-level features of the PAN image with injected spectral information with the features of the MS image. The other CFB subnetwork is used to fuse the low-level features of the PAN image with injected spectral information with the high-level features of the PAN image.

[0014] The output module outputs the fused image.

[0015] Preferably, the three FEB modules are expressed by the following formula:

[0016] F ms =f FEB (UPMS)

[0017] F pan_H =f FEB_H (PAN)

[0018] F pan_L =f FEB_L (PAN)

[0019] In the formula, F ms F represents the features of an MS image. pan_H F represents the high-level features of a PAN image. pan_L f represents the low-level features of the PAN image. FEB f represents feature extraction from MS images. FEB_H This represents high-level feature extraction from PAN images, f FEB_LThis represents low-level feature extraction from the PAN image.

[0020] Preferred, f FEB and f FEB_L It consists of two convolutional layers with PReLU activation. The first layer is a convolutional layer with 128 channels and a kernel size of 3×3, and the second layer is a convolutional layer with 64 channels and a kernel size of 1×1.

[0021] Preferred, f FEB_H It consists of two convolutional layers with PReLU activation. The first layer is a convolutional layer with 128 channels and a kernel size of 7×7, and the second layer is a convolutional layer with 64 channels and a kernel size of 1×1.

[0022] Preferably, the PAN image and MS image are input into a deep coupled feedback network, and the final output is a fused image, specifically including the following steps:

[0023] The PAN and MS images are input into a deep coupled feedback network;

[0024] via f FEB_H and f FEB_L Extracting high-level and low-level features from the PAN image yields F. pan_H and F pan_L , through f FEB Extracting features F from MS images ms ;

[0025] Inject spectral information into F pan_L In, and combined with F ms The first group of CFB modules in a CFB subnetwork is input to inject spectral information F pan_L With F pan_H Input to the first group of CFB modules in another CFB subnetwork;

[0026] Each subsequent group of CFB modules, upon receiving the output from the previous group of CFB modules, will also receive F again. ms and F pan_H Perform feature fusion;

[0027] After feature fusion of multiple CFB modules, the output of the last CFB module is finally integrated across channels to form the final fused image.

[0028] Preferably, each CFB module includes:

[0029] The input module is used to receive the corresponding CFB output features and corresponding image features from the previous group of CFB modules;

[0030] The convolution module is used to combine the two input features mentioned above;

[0031] Multiple projection groups, each of which includes upsampling and downsampling, wherein the upsampling is a deconvolution layer and the downsampling is a convolution layer;

[0032] The output module is used to output the fused features.

[0033] Preferably, when performing feature fusion, a CFB module in group t includes the following steps:

[0034] Receive the output of the CFB module corresponding to group t-1 and F pan_H ;

[0035] The two input features are combined using a 1×1 convolutional layer to obtain refined features from the two inputs.

[0036] Will The input is fed into multiple projection groups, and upsampling and downsampling operations are repeated to obtain high-resolution and low-resolution feature maps.

[0037] Collect the LR feature maps of all projection groups, fuse them using a set of 1×1 filters, and form the output of the t-th CFB module.

[0038] Preferably, the convolution module combines the two input features through a 1×1 convolutional layer, as shown in the following equation:

[0039]

[0040] It is a refined feature based on two inputs, M in It is a set of 1×1 convolutions. It is a chain of internal elements.

[0041] Preferably, high-resolution and low-resolution feature maps are obtained using the following formula:

[0042]

[0043]

[0044] in D represents the nth high-resolution feature map. n C represents the upsampling operation of the nth projection group. n This indicates the downsampling operation in the nth projection group.

[0045] Preferably, the output of the t-th CFB is as follows:

[0046]

[0047] Where M out This represents a 1×1 convolution operation.

[0048] Compared with the prior art, the beneficial effects of the present invention are:

[0049] The deep coupled feedback network PSCF-Net of this invention can achieve deep feature fusion of PAN and MS images in the feature domain, can well preserve the spectral information of MS images, and can also accurately extract the spatial information of PAN images, ultimately obtaining high-quality images with both high spectral and spatial resolution. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart of a full-color sharpening method based on a deep coupled feedback network according to the present invention;

[0052] Figure 2 This is a schematic diagram of the structure of the deep coupled feedback network of the present invention;

[0053] Figure 3 This is a schematic diagram of the structure of the t-th group of CFBs in the deep coupled feedback network of the present invention;

[0054] Figure 4 (a) is the MS image of the IKONOS dataset after bicubic interpolation;

[0055] Figure 4 (b) is an image from the IKONOS dataset PAN;

[0056] Figure 4 (c) is the HRMS image obtained from the IKONOS dataset using PSCF-Net;

[0057] Figure 4 (d) is the residual image between the HRMS image obtained from the IKONOS dataset and the source image;

[0058] Figure 5 (a) is the MS image of the WORLDVIEW-2 dataset after bicubic interpolation;

[0059] Figure 5(b) is a PAN image from the WORLDVIEW-2 dataset;

[0060] Figure 5 (c) is the HRMS image obtained from the WORLDVIEW-2 dataset using PSCF-Net;

[0061] Figure 5 (d) is the residual image between the HRMS image obtained from the WORLDVIEW-2 dataset and the source image. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] In this invention, PAN image refers to panchromatic image, MS image refers to multispectral image, and PSCF-Net refers to deep coupled feedback network.

[0064] Reference Figure 1 This invention provides a panchromatic sharpening method based on a deeply coupled feedback network, comprising the following steps:

[0065] Step 1: Acquire PAN and MS images;

[0066] Step 2: Construct a deeply coupled feedback network.

[0067] Reference Figure 2 This invention proposes a Deep Coupled Feedback Network (PSCF-Net) for pan-color sharpening. PSCF-Net has two CFB sub-networks. First, low-level and high-level features are extracted from the PAN image using two different Feature Extraction Blocks (FEBs). Then, spectral information is injected into the low-level features and fed together with the features extracted from the MS image into one CFB sub-network for feature fusion. The other CFB sub-network receives the low-level features with injected spectral information and the high-level features from the PAN image. Multiple CFBs are used in each sub-network, and the CFBs at corresponding positions in two sub-networks are considered a group. Each group of modules after the first group receives the output of the previous group of CFBs and also receives the previously extracted MS image features and high-level features from the PAN image. After feature fusion through multiple groups of CFBs, the output of the last group of CFBs is finally integrated across channels to form the final fused image.

[0068] The features extracted by the three FEBs can be expressed by the following formulas:

[0069] F ms =f FEB (UPMS) (1)

[0070] F pan_H =f FEB_H (PAN) (2)

[0071] F pan_L =f FEB_L (PAN) (3)

[0072] Where F ms F pan_H F pan_L These represent the features of the MS image, and the high-level and low-level features of the PAN image, respectively. FEB f FEB_H f FEB_L These represent feature extraction for MS images, and high-level and low-level feature extraction for PAN images, respectively. FEB and f FEB_L It consists of two convolutional layers with PReLU activation: the first layer is a 128-channel convolutional layer with a 3×3 kernel, and the second layer is a 64-channel convolutional layer with a 1×1 kernel. The first layer extracts basic features, while the second layer further integrates cross-channel features to make them more compact. FEB_H It also consists of two convolutional layers with PReLU activation, except that the first layer is replaced with a convolutional layer with 128 channels and a kernel size of 7×7, extracting the feature F. ms F pan_H and F after injecting spectral information pan_L As input to the subsequent two CFB subnetworks.

[0073] CFB is the most important module of PSCF-Net, designed to achieve function fusion and super-resolution through complex network connections. This coupled feedback mechanism has proven to bring significant benefits to image super-resolution and image fusion tasks. The feedback mechanism allows the network to carry the concept of the output to correct previous states. The coupling mechanism is mainly used to ensure the network's ability to integrate features at different levels.

[0074] Reference Figure 3 There are multiple CFBs in two subnets, but the structure of each group of CFBs (corresponding to the positions of the CFBs in the two subnets) is the same. Therefore, taking group t of CFBs in the network as an example, we will introduce its interaction with other modules and its internal structure.

[0075] The t-th group of CFB receives two inputs. and The first group of CFBs was generated, and the second group was the previously extracted F. ms and Fpan_H Specifically, and It is feedback from the same subnet, so its main function is to correct low-level features after injecting spectral information to improve the performance of feature extraction. In contrast, F ms and F pan_H These are features derived from the MS image and the high-level features of the PAN image, respectively. Their main function is to provide spatial information from the PAN image and spectral information from the MS image to improve the fusion performance of CFBs. Since the structures of the two CFBs above and below a set are identical, this section mainly introduces... Figure 2 The CFB structure above. Before further feature fusion, the CFB combines the two input features through a 1×1 convolutional layer, as shown below:

[0076]

[0077] It is a refined feature based on two inputs, M in It is a set of 1×1 convolutions. It's a chain of internal elements. After that, The upsampling and downsampling operations are repeatedly performed on the input to extract high-level features that refine the input and improve the quality of feature fusion. Each upsampling and downsampling operation is treated as a projection group, in which upsampling is achieved through deconvolutional layers and downsampling through convolutional layers. Upsampling yields a high-resolution feature map, while downsampling transforms it into a low-resolution feature map with the same resolution as the input. To ensure that the information of each feature is preserved during the fusion process, dense connections are used to consider the previous feature maps at different resolutions.

[0078]

[0079] in Let D represent the nth high-resolution feature map, corresponding to which D is... n This represents the upsampling operation for the nth projection group, which is performed by a set of 1×1 convolutions and a deconvolution layer with a kernel size of 6×6 and a stride of 2. As can be seen from Equation 5, all previous low-resolution feature maps are considered before performing the deconvolution. The process of obtaining low-resolution feature maps is similar; the nth low-resolution feature map... All previous high-resolution feature maps will be considered before downsampling:

[0080]

[0081] Where C nThis represents the convolution operation in the nth projection group. Whether it's a deconvolution or convolution operation, the upsampling and downsampling ratios can be adjusted by modifying the convolution kernel and stride.

[0082] As the number of projection groups increases, the feedback characteristics of the corresponding subnets change. The reduced influence of [the material] led to an unsatisfactory final fusion result. To maintain [the desired effect] during the fusion process... In CFB, the input of each CFB is embedded into the projection group to re-stimulate the network's memory, thereby improving the fusion effect of CFB. Considering the module structure, the selection... Location as a stimulus:

[0083]

[0084] This formula indicates the new feature map. In The addition replaced Finally, after 3 projection groups, all LR feature maps are collected from each projection group and fused together using a set of 1×1 filters to form the output of the t-th CFB, as shown below:

[0085]

[0086] Where M out This represents a 1×1 convolution operation.

[0087] PSCF-Net was tested on the IKONOS and WorldView-2 satellite datasets. For both datasets, a 128×128 MS image and a 512×512 PAN image were used as source images. To simulate the degradation process of the MS image, the pixel resolution of the original MS image was reduced from 128×128 to 32×32 using a Wald protocol modulation transfer function (MTF) filter. Simultaneously, the pixel resolution of the original PAN image was reduced from 512×512 to 128×128 using a bicubic interpolation method. Then, the downsized MS image (32×32) was bicubic interpolated to match the resolution of the PAN image (128×128). Finally, this pair of MS and PAN images were used as training samples input into the network, with the original MS image serving as the reference image.

[0088] The network was trained on two separate datasets. In the IKONOS dataset, 24,823 pairs of preprocessed images were used as training samples, and the test set contained 30 pairs of 256×256 images. In the WorldView-2 dataset, 13,459 pairs of preprocessed images were used as training samples, and the test set contained 24 pairs of 256×256 images. During training, the Adam optimizer was used with an epoch of 100, a learning rate of 0.00001, and a batch size of 8. Furthermore, to enhance the network's robustness, SmoothL1 was used as the loss function, and Structural Similarity (SSIM) was incorporated to improve the preservation of image brightness, contrast, and structural information.

[0089] Experimental results in the IKONOS dataset and WorldView-2 are as follows: Figure 4 and Figure 5 As shown, Figure 4 (a) and Figure 5 (a) is the MS image after bicubic interpolation. Figure 4 (b) Figure 5 (b) is the PAN image. Figure 4 (c) and Figure 5 (c) is the HRMS image obtained by PSCF-Net. Figure 4 (d) and Figure 5 (d) shows the residual plots of the obtained HRMS image and the source image. It can be seen that the interpolated MS image has very little spatial information, while the PAN image, although possessing rich spatial information, lacks spectral information. From Figure 4 (c) and Figure 5 (c) It can be seen that PSCF-Net effectively preserves spectral information while improving the resolution of MS images. From Figure 4 (d) and Figure 5 As can be seen from (d), the structural information of the residual map is not obvious, which also proves that the method preserves the structural information of the image very well.

[0090] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0091] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A panchromatic sharpening method based on a deeply coupled feedback network, characterized in that, Includes the following steps: Acquire PAN and MS images; Construct a deeply coupled feedback network; The PAN image and MS image are input into a deep coupled feedback network, and the final output is a fused image. The deeply coupled feedback network includes: The input module receives PAN and MS images; The feature extraction module includes three FEB modules for extracting low-level and high-level features of PAN images and features of MS images; Two CFB subnetworks, each containing multiple sets of symmetrical CFB modules, with multiple CFBs used in each subnetwork. CFBs at corresponding positions in the two subnetworks are considered as a group. One CFB subnetwork is used to fuse the low-level features of the PAN image with injected spectral information with the features of the MS image, and the other CFB subnetwork is used to fuse the low-level features of the PAN image with injected spectral information with the high-level features of the PAN image. Each module after the first group receives the previously extracted MS image features and PAN image high-level features while receiving the CFB output of the previous group; after feature fusion of multiple groups of CFB, the output of the last group of CFB is finally integrated across channels into the final fused image. The output module outputs the fused image.

2. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 1, characterized in that, The three FEB modules are expressed by the following formula: In the formula, Features representing MS images High-level features representing PAN images, Represents low-level features of the PAN image. This represents feature extraction from MS images. This represents high-level feature extraction from the PAN image. This represents low-level feature extraction from the PAN image.

3. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 2, characterized in that, and It consists of two convolutional layers with PReLU activation. The first layer is a convolutional layer with 128 channels and a kernel size of 3×3, and the second layer is a convolutional layer with 64 channels and a kernel size of 1×1.

4. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 2, characterized in that, It consists of two convolutional layers with PReLU activation. The first layer is a convolutional layer with 128 channels and a kernel size of 7×7, and the second layer is a convolutional layer with 64 channels and a kernel size of 1×1.

5. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 2, characterized in that, The PAN image and MS image are input into a deep coupled feedback network, and the final output is a fused image. The specific steps include: The PAN and MS images are input into a deep coupled feedback network; pass and Extracting high-level and low-level features from the PAN image yields... and ,pass Extracting features from MS images ; Inject spectral information In, and combined The input is given to the first CFB module in a CFB subnetwork, which injects spectral information. and Input to the first CFB module in another CFB subnetwork; Each subsequent group of CFB modules, upon receiving the output from the previous group of CFB modules, will also receive... and Perform feature fusion; After feature fusion of multiple CFB modules, the output of the last CFB module is finally integrated across channels to form the final fused image.

6. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 2, characterized in that, Each CFB module includes: The input module is used to receive the output features and corresponding image features of the corresponding CFB module in the previous group of CFB modules; The convolution module is used to combine the two input features mentioned above; Multiple projection groups, each of which includes upsampling and downsampling, wherein the upsampling is a deconvolution layer and the downsampling is a convolution layer; The output module is used to output the fused features.

7. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 6, characterized in that, The output of a CFB module in group t to the corresponding CFB module in group (t-1) Feature fusion includes the following steps: Receive the output of the CFB module corresponding to group t-1 and ; The two input features are combined using a 1×1 convolutional layer. and Combining the two input features yields refined features. ; Will The input is fed into multiple projection groups, and upsampling and downsampling operations are repeated to obtain high-resolution and low-resolution feature maps. Collect the LR feature maps of all projection groups, fuse them using a set of 1×1 filters to form the output of the t-th CFB module.

8. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 7, characterized in that, The convolutional module combines the two input features through a 1×1 convolutional layer, as shown in the following equation: It is a refined feature based on two input features. It is a set of 1×1 convolutions. It is a chain of internal elements.

9. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 8, characterized in that, High-resolution and low-resolution feature maps are obtained using the following formula: in This represents the nth high-resolution feature map. This represents the upsampling operation for the nth projection group. This represents the downsampling operation in the nth projection group. This represents the nth low-resolution feature map.

10. The panchromatic sharpening method based on a deeply coupled feedback network as described in claim 9, characterized in that, The output of the t-th CFB is shown in the following formula: in This represents a 1×1 convolution operation. Indicates used to replace New feature map, .