Real scene remote sensing image super-resolution reconstruction method and system based on progressive feature aggregation network, and storage medium

By adopting a progressive feature aggregation network, key information perception module, local feature enhancement module and second-order shuffled random degradation model in the super-resolution reconstruction of remote sensing images, the problems of insufficient performance of remote sensing image resolution improvement and difficulty in fitting the degradation model in the prior art are solved, and high-quality super-resolution reconstruction of remote sensing images are achieved.

CN120013762APending Publication Date: 2025-05-16FUZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510102991.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing remote sensing image super-resolution reconstruction methods have problems such as insufficient performance, long training time, large resource consumption, and the difficulty of fitting low-resolution remote sensing images in improving the resolution of remote sensing images in remote sensing images of the existing remote sensing images in improving the resolution of remote sensing images of remote sensing images of remote sensing images in remote sensing images of the simple degradation model.

Method used

A super-resolution reconstruction method for real scene remote sensing images based on a progressive feature aggregation network is proposed. By designing a key information perception module and a local feature enhancement module, combining a second-order shuffled random degradation model and a super-resolution reconstruction loss function for real scenes, a more realistic degradation model and more effective feature extraction are achieved.

Benefits of technology

It improves the resolution of remote sensing images in real scenes, restores more high-frequency details and texture features, improves image reconstruction quality, and reduces the complexity and training time of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013762A_ABST
    Figure CN120013762A_ABST
Patent Text Reader

Abstract

The invention provides a real scene remote sensing image super-resolution reconstruction method and system based on a progressive feature aggregation network, and a storage medium. The method comprises the following steps: establishing the progressive feature aggregation network PFAN for real scene remote sensing image super-resolution reconstruction; establishing a second-order shuffling random degradation model, and performing second-order shuffling random degradation on the remote sensing image serving as a training sample to obtain a low-resolution remote sensing image simulating the random degradation process of a real scene; inputting the remote sensing image before degradation and the corresponding low-resolution remote sensing image into a PFAN network for training; and obtaining a real remote sensing image to be reconstructed, and inputting the real remote sensing image to the trained PFAN network for image reconstruction to obtain a super-resolution image. According to the method, the edge details and texture features of the remote sensing image are aggregated through a progressive strategy, and the second-order shuffling random degradation model is utilized to assist the network in learning the complex degradation process of the real scene remote sensing image, so that higher-precision super-resolution reconstruction of the real scene remote sensing image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a real scene remote sensing image super-resolution reconstruction method, system and storage medium based on a progressive feature aggregation network. Background Art

[0002] High-resolution remote sensing images are widely used in military reconnaissance, environmental monitoring, agricultural management, disaster relief and other fields. However, due to the hardware performance limitations of the imaging equipment and the influence of the imaging distance, the acquired remote sensing images usually have low resolution. In addition, during the image acquisition and transmission process, they will be interfered by factors such as system noise, motion blur, and image compression, resulting in impaired resolution, which further reduces the imaging quality of remote sensing images. Therefore, it is of great significance to study methods to improve the resolution of remote sensing images.

[0003] From the perspective of post-processing, image super-resolution reconstruction technology can improve the resolution of remote sensing images without upgrading hardware equipment, and has the advantages of low cost, short cycle and high efficiency. In recent years, with the development of deep learning technology, super-resolution reconstruction methods based on convolutional neural networks have achieved remarkable results in improving the resolution of remote sensing images. However, existing methods often use very deep network architectures to improve performance, resulting in long training time, high difficulty and high resource consumption. In terms of performance, existing methods are still insufficient in the recovery of high-frequency details and local texture features, and reconstructed images are prone to blurring and loss of texture structure information. In addition, most remote sensing image super-resolution reconstruction methods often use a single degradation method or a relatively simple degradation model to simulate the generation of low-resolution remote sensing images. However, the degradation process of remote sensing images in real scenes is interfered by many factors, and the degradation order and occurrence are not fixed. Simple degradation models are often difficult to fit the generation of low-resolution remote sensing images in real scenes. The above problems limit the application scope of super-resolution reconstruction methods in the field of remote sensing to a certain extent. Summary of the invention

[0004] The purpose of the present invention is to propose a real-scene remote sensing image super-resolution reconstruction method, system and medium based on a progressive feature aggregation network, which can restore more high-frequency details and texture features through a more realistic degradation model and a more effective feature extraction method, thereby improving the resolution of remote sensing images in real scenes.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] In a first aspect, the present invention proposes a real scene remote sensing image super-resolution reconstruction method based on a progressive feature aggregation network, the method comprising the following steps:

[0007] S100, establishing a progressive feature aggregation network (PFAN) for super-resolution reconstruction of real-scene remote sensing images, the PFAN network comprising a shallow feature extraction module, a deep feature extraction module and a reconstruction module; wherein the deep feature extraction module comprises a plurality of progressive feature aggregation modules, each of which comprises a key information perception module and a local feature enhancement module, the key information perception module adopts a dual-path fusion structure, the two paths respectively capture high-frequency information from the channel dimension and the spatial dimension, the local feature enhancement module adopts a multi-branch structure and has different convolutional layer sizes, and is used to globally interact the image features output by the key information perception module;

[0008] S200, establishing a second-order shuffle random degradation model, subjecting the remote sensing image used as a training sample to second-order shuffle random degradation, and obtaining a low-resolution remote sensing image that simulates a random degradation process of a real scene;

[0009] S300, inputting the remote sensing image before degradation and the corresponding low-resolution remote sensing image into the PFAN network for training;

[0010] S400, obtaining a real remote sensing image to be reconstructed, inputting the trained PFAN network to perform image reconstruction, and obtaining a super-resolution image.

[0011] Preferably, the shallow feature extraction module is a 3×3 convolutional layer; the shallow feature extraction module is described as:

[0012] F0=f SFE (I LR )

[0013] where f SFE (·) represents the shallow feature extraction module, I LR represents a low-resolution remote sensing image, and F0 represents a shallow feature map;

[0014] The deep feature extraction module includes several progressive feature aggregation modules and a 3×3 convolutional layer to aggregate deeper high-level features from shallow features. The process is expressed as follows:

[0015] F i =f PFABi (F i-1 ),i=1,2,...,n

[0016] F=f Conv3×3 (F n )

[0017] Among them, f PFABi(·) represents the i-th progressive feature aggregation module, n represents the number of progressive feature aggregation modules, and F i-1 、F i 、F n represents the intermediate feature map of the i-1th, i-th, and n-th progressive feature aggregation modules, f Conv3×3 (·) represents the convolution operation with a kernel size of 3×3, and F represents the deep feature map output by the deep feature extraction module;

[0018] The reconstruction module includes a sub-pixel convolution layer and a 3×3 convolution layer. A sub-pixel convolution layer is used to upsample the deep features extracted by the deep feature extraction module, and the information flow is reorganized into a feature map with a specified upsampling ratio. The process is described as follows:

[0019] I SR =f IR (F+F0)+Bilinear(I LR )

[0020] Among them I SR represents the super-resolution image after super-resolution reconstruction, f IR (·) represents the reconstruction module, and Bilinear(·) represents the bilinear interpolation of the low-resolution image to the target resolution.

[0021] Preferably, the progressive feature aggregation module includes a key information perception module (Key Information Perception Module, KIPM), a local feature enhancement module (Local Feature Enhancement Module, LFEM), two feed-forward networks (Feed-Forward Network, FFN) and residual connections, wherein the key information perception module adopts a dual-path fusion structure, the two paths respectively capture high-frequency information across channels and spaces to restore edge details, the local feature enhancement module further globally interacts the captured high-frequency information to enhance texture features, and the feed-forward network gradually aggregates rich edge details and texture features to improve the reconstruction quality of remote sensing images. The specific calculation formula is expressed as follows:

[0022] Y=FFN(LFEM(FFN(KIPM(X))))+X

[0023] Where X represents the input features of the progressive feature aggregation module, KIPM(·) represents the key information perception module, FFN(·) represents the forward propagation network, LFEM(·) represents the local feature enhancement module, and Y represents the output features of the progressive feature aggregation module.

[0024] Preferably, the key information perception module includes a layer normalization operation, two 1×1 convolutional layers, channel splitting, a Low Energy Channel Attention (LECA), a Dual Low Energy Spatial Attention (DLESA), and a 7×7 depth separable convolutional layer, and the specific operations are as follows:

[0025] X1,X2=split(f Conv1×1 (LN(X)))

[0026] Y K =X·Sig(f Conv1×1 (f DWConv7×7 (LECA(X1)+DLESA(X2))))

[0027] Among them, X represents the input features of the key information perception module, LN(·) represents the layer normalization operation, and f Conv1×1 (·) represents a 1×1 convolution operation, split(·) represents a channel splitting operation, X1 and X2 represent the feature layers after channel splitting, LECA(·) represents the lowest energy channel attention operation, DLESA(·) represents the dual lowest energy spatial attention operation, and f DWConv7×7 (·) represents a 7×7 depthwise separable convolution operation, Sig(·) represents the Sigmoid function, · represents the channel-wise multiplication operation, and Y K Represents the output features of the key information perception module;

[0028] The minimum energy channel attention uses global average pooling to extract features, and uses the energy function minimization method to enhance features and integrate important channel information. The specific operations are as follows:

[0029] X 11 =GP(X1)

[0030]

[0031] Y K1 =X1·Sig(X 12 )

[0032] Among them, X1 represents the input feature of the lowest energy channel attention, GP(·) represents the global average pooling operation, and X 11 ∈R C×1×1 represents the channel vector obtained after the pooling operation, μ represents the average value of the channel vector, σ 2 represents the variance of the channel vector, λ represents the stability parameter of the formula, X 12 ∈R C×1×1represents the weighted attention map, Sig(·) represents the Sigmoid function, and Y K1 Output features representing the lowest energy channel attention;

[0033] The dual minimum energy spatial attention enhances spatial features by minimizing the energy function. The specific operations are as follows:

[0034]

[0035]

[0036] Y K2 =X2·Sig(X 21 +X 22 )

[0037] Among them, X2 represents the input features of the dual minimum energy spatial attention, μ1∈R C×H×1 represents the average value calculated along the spatial width W, σ1 2 ∈R C×H×1 represents the variance calculated along the spatial width W, μ2∈R C×1×W represents the average value calculated along the spatial height H, σ2 2 ∈R C×1×W represents the variance calculated along the spatial height H, λ represents the stability parameter of the formula, X 21 ,X 22 ∈R C×H×W represents the weighted attention map along different spatial directions, Sig(·) represents the Sigmoid function, and Y K2 Output features representing the lowest energy channel attention.

[0038] Preferably, the local feature enhancement module adopts a multi-branch structure and has different convolutional layer sizes, and is used to globally interact the image features output by the key information perception module. The local feature enhancement module includes a layer normalization, a local feature aggregation branch, a context-aware local enhancement branch, a large core attention branch, channel merging and a 1×1 convolutional layer;

[0039] The local feature aggregation branch extracts local features by a 9×9 depthwise separable convolution;

[0040] The context-aware local enhancement branch includes a 5×5 depthwise separable convolutional layer, a 5×5 dilated convolutional layer, a 7×7 dilated convolutional layer, Hardswish, Tanh, and two 1×1 convolutional layers;

[0041] The large core attention branch includes a 7×7 depthwise separable convolutional layer, a 9×9 dilated convolutional layer, and a 1×1 convolutional layer; the specific operations are as follows:

[0042] V=f DWConv9×9 (LN(X L ))

[0043] Q=f DWDConv5×5 (f DWConv5×5 (LN(X L )))

[0044] K=f DWDConv7×7 (f DWConv5×5 (LN(X L )))

[0045] X L1 =Tanh(f Conv1×1 (Hardswish(f Conv1×1 (Q·K))))·V

[0046] X L2 =f Conv1×1 (f DWDConv9×9 (f DWConv7×7 (LN(X L ))))·LN(X L )

[0047] Y L =f Conv1×1 (concat(X L1 ,X L2 ))+X L

[0048] Among them, X L represents the input features of the local feature enhancement module, LN(·) represents the layer normalization operation, and f DWConv9×9 (·) represents a 9×9 depthwise separable convolution operation, f DWConv5×5 (·) represents a 5×5 depthwise separable convolution operation, f DWDConv5×5 (·) represents a 5×5 dilated convolution operation, f DWDConv7×7 (·) represents a 7×7 dilated convolution operation, f Conv1×1 (·) represents a 1×1 convolution operation, Hardswish(·) represents the Hardswish function, Tanh(·) represents the Tanh function, and f DWConv7×7 (·) represents a 7×7 depthwise separable convolution operation, f DWDConv9×9 (·) represents a 9×9 dilated convolution operation, concat(·) represents a channel merging operation, and Y L Represents the output features of the local feature enhancement module.

[0049] Preferably, the forward propagation network includes a layer normalization operation, two 1×1 convolutional layers and channel splitting, and the specific operations are as follows:

[0050] X F1 ,X F2 = split(f Conv1×1 (LN(X F )))

[0051] Y F =f Conv1×1 (X F1 ·X F2 )+X F

[0052] Among them, X F represents the input features of the forward propagation network, LN(·) represents the layer normalization operation, and f Conv1×1 (·) represents a 1×1 convolution operation, split(·) represents a channel split operation, and X F1 , X F2 Represents the feature layer after channel splitting, Y F Represents the output features of the forward propagation network.

[0053] Preferably, PFAN network training adopts a remote sensing image super-resolution reconstruction loss function for real scenes that combines pixel loss, perceptual loss, and structural loss:

[0054]

[0055] Among them, L represents the comprehensive loss, L1 represents the pixel loss, and L per represents the perceptual loss, L s Indicates structural loss;

[0056] The pixel loss calculation formula is as follows:

[0057] L1=||I HR -f(I LR ,θ)||1

[0058] Among them, I HR represents high-resolution remote sensing images, I LR represents the input low-resolution remote sensing image to be reconstructed, f(·,θ) represents the super-resolution reconstruction network with parameters θ, and ||·||1 represents the L1 norm;

[0059] The perceptual loss calculation formula is as follows:

[0060]

[0061] in, represents the pre-trained VGG network, ||·||1 represents the L1 norm;

[0062] The formula for calculating structural loss is as follows:

[0063]

[0064] L s =1-TSSIM(D·f(I LR ,θ),f(D·I LR ,θ),D·I HR )

[0065] Among them, μ x , μ y , μ z and σ x , σ y , σ z Represent the mean and standard deviation of x, y, and z, σ xy , σ yz , σ xz Respectively represent the correlation coefficients between (x, y), (y, z), (x, z), C1, C2 represent hyperparameters, D·f(I LR ,θ) represents the deformed super-resolution reconstructed image, f(D·I LR ,θ) represents the reconstructed image of the deformed low-resolution remote sensing image, D·I HR represents the deformed high-resolution remote sensing image, and TSSIM(·) represents the deformation function.

[0066] Preferably, the second-order shuffled random degradation model includes two first-order shuffled random degradation models, and the first-order shuffled random degradation model includes four degradation categories shuffled, and the four degradation categories are blur, noise, compression and scaling. Each degradation type occurs randomly with a certain probability P. The specific operation is as follows:

[0067]

[0068] DR∈{DR1,DR2,DR j ...DR 24}

[0069] DR1={A,B,C,D}1,DR2={A,C,B,D}2,DR j ={B,D,A,C} j …DR 24 ={C,A,D,B} 24

[0070]

[0071] X t ~N(μ t ,σ t 2 )

[0072] Among them, ILR represents low-resolution remote sensing images, I HR Represents high-resolution remote sensing images, DR first-order shuffle random degradation model, DR j represents one of the four degradation category shuffle combinations, A, B, C, D represent blur, noise, compression, scaling operation or no operation, P is the probability threshold, X t Obey the standard normal distribution μ t =0,σ t =1.

[0073] In a second aspect, the present invention proposes a real scene remote sensing image super-resolution reconstruction system based on a progressive feature aggregation network, the system comprising:

[0074] At least one computing device, the computing device comprising at least one processor and at least one memory, the memory being used to store at least one program;

[0075] at least one display, the display being used to display the operating results of the computing device;

[0076] When the at least one program is executed by the computing device, the computing device implements the above-mentioned real-scene remote sensing image super-resolution reconstruction method based on a progressive feature aggregation network.

[0077] In a third aspect, the present invention proposes a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute the above-mentioned real-scene remote sensing image super-resolution reconstruction method based on a progressive feature aggregation network.

[0078] Compared with the prior art, the present invention has the following beneficial effects:

[0079] (1) The present invention proposes an efficient and lightweight progressive feature aggregation network that can extract rich edge details and texture features through a progressive strategy to achieve high-quality super-resolution reconstruction of remote sensing images;

[0080] (2) The present invention designs a key information perception module and a local feature enhancement module to capture high-frequency information and internal self-similar information in the image. Specifically, the key information perception module contains two parameter-free attention designs, LECA and DLESA, which capture high-frequency information in the channel and spatial dimensions respectively, and help to restore the edge details of the image. Compared with traditional methods, the key information perception module adopts a parameter-free attention design to ensure the lightweight of the module. The local feature enhancement module uses shared and context-aware weights to capture the self-similar information inside the image, which helps to restore the complex texture features of the image. Compared with the traditional Transformer, it gives full play to the ability of multi-scale convolution to extract local features and the Transformer global information interaction;

[0081] (3) The present invention proposes a second-order shuffle random degradation model for fitting the generation of low-resolution remote sensing images of real scenes, which effectively improves the quality of super-resolution reconstruction of remote sensing images in real scenes. Compared with traditional methods, the shuffle strategy is used to randomly add degradation categories such as blur, noise, compression, and scaling to simulate the imaging process subject to many interferences, such as system noise, motion blur, transmission compression, etc., which can more fully fit the degradation of remote sensing images of real scenes;

[0082] (4) The present invention proposes a super-resolution reconstruction loss function for remote sensing images in real scenes, which fully considers the advantages of pixel loss, perception loss and structure loss, avoids some problems that may exist in super-resolution reconstruction of remote sensing images in real scenes, such as erroneous edges and distorted structures, and ensures that the network can better meet the application requirements of real scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 It is a flowchart of a method for super-resolution reconstruction of remote sensing images of real scenes based on a progressive feature aggregation network in an embodiment of the present invention;

[0084] Figure 2 is a diagram of the overall structure of a progressive feature aggregation network in an embodiment of the present invention;

[0085] Figure 3 is a schematic diagram of the structure of a key information perception module in an embodiment of the present invention;

[0086] Figure 4 is a schematic diagram of the lowest energy channel attention in an embodiment of the present invention;

[0087] Figure 5 is a schematic diagram of dual minimum energy spatial attention in an embodiment of the present invention;

[0088] Figure 6 is a schematic diagram of the structure of a local feature enhancement module in an embodiment of the present invention;

[0089] Figure 7 is a schematic diagram of a second-order shuffle random degradation model in an embodiment of the present invention;

[0090] Figure 8 It is a structural schematic diagram of a real scene remote sensing image super-resolution reconstruction system based on a progressive feature aggregation network in an embodiment of the present invention. DETAILED DESCRIPTION

[0091] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects of the present invention, so as to fully understand the purpose, scheme and effect of the present invention. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.

[0092] Super-resolution reconstruction technology based on deep learning is an important means to improve the quality of remote sensing images. However, the current mainstream method uses a relatively single CNN network, which makes it impossible to utilize global self-similar information and is insufficient in restoring texture features. The Transformer-based method can model global information, thereby effectively utilizing the self-similar information inside the image, which is helpful for texture feature recovery, but at the same time faces problems such as insufficient extraction of high-frequency information and high complexity of the network model, and is not suitable for remote sensing equipment. The present invention effectively combines the advantages of CNN and Transformer, and adopts a progressive strategy to aggregate the high-frequency edge details and local texture features of remote sensing images, while improving the quality of remote sensing image reconstruction and keeping the model lightweight. In addition, few existing methods have studied the degradation of low-resolution remote sensing images. Most of them use a relatively simple method to simulate the degradation of remote sensing images, which is not enough to fit the degradation process of remote sensing images in real scenes. Therefore, the present invention also proposes a second-order shuffle random degradation model to fit the random degradation of real-scene remote sensing images, effectively improving the practical application of remote sensing super-resolution reconstruction methods.

[0093] Figure 1 The present invention provides a real scene remote sensing image super-resolution reconstruction method based on a progressive feature aggregation network, the method comprising the following steps:

[0094] S100, establishing a progressive feature aggregation network (PFAN) for super-resolution reconstruction of real-scene remote sensing images, wherein the PFAN network includes a shallow feature extraction module, a deep feature extraction module and a reconstruction module; wherein the deep feature extraction module includes a plurality of progressive feature aggregation modules, each of which includes a key information perception module and a local feature enhancement module, wherein the key information perception module adopts a dual-path fusion structure, wherein the two paths capture high-frequency information from the channel dimension and the spatial dimension respectively, and the local feature enhancement module adopts a multi-branch structure and has different convolutional layer sizes, and is used to globally interact the image features output by the key information perception module;

[0095] S200, establishing a second-order shuffle random degradation model, subjecting the remote sensing image used as a training sample to second-order shuffle random degradation, and obtaining a low-resolution remote sensing image simulating a random degradation process of a real scene;

[0096] S300, inputting the remote sensing image and the low-resolution remote sensing image into a PFAN network for training;

[0097] S400, obtaining a real remote sensing image to be reconstructed, and inputting it into a trained PFAN network for image reconstruction to obtain a super-resolution image.

[0098] To solve the problem of super-resolution reconstruction of remote sensing images, this paper proposes a method based on progressive feature aggregation network (PFAN), which combines the advantages of CNN network and Transformer network to achieve efficient and high-quality image reconstruction.

[0099] Specifically, first, a key information perception module (KIPM) is designed to capture high-frequency details across spatial and channel dimensions to restore edge features. Among them, KIPM only uses two 1×1 convolutional layers, one 7×7 depth-separable convolutional layer and two parameter-free attention designs to achieve efficient information perception while ensuring the lightweight of the module. Secondly, a local feature enhancement module (LFEM) is designed based on Transformer and CNN, which uses multi-scale convolution to extract local features and the global information interaction ability of Transformer to restore complex texture features. Among them, LFEM uses depth-separable convolution, dilated convolution and 1×1 convolution design to reduce the number of model parameters and improve the efficiency of the model. Finally, by gradually aggregating rich edge details and complex texture features, PFAN effectively improves the quality of super-resolution reconstruction of remote sensing images.

[0100] In addition, the present invention also proposes a second-order shuffling random degradation model, which uses a shuffling strategy to randomly add degradation categories such as blur, noise, compression, and scaling to simulate the imaging process being subject to many interferences, such as system noise, motion blur, transmission compression, etc., effectively improving the practical application of remote sensing super-resolution reconstruction methods.

[0101] In a preferred embodiment, in S100, a PFAN network for super-resolution reconstruction of real scene remote sensing images is established, such as Figure 1 As shown, including:

[0102] S110, shallow feature extraction is performed on the input image. The specific operations are as follows:

[0103] F0=f SFE (I LR )

[0104] where f SFE (·) indicates that the shallow feature extraction module consists of a 3×3 convolution. LR represents a low-resolution remote sensing image, and F0 represents a shallow feature map;

[0105] S120, extract and fuse shallow features through a deep feature extraction module composed of n cascaded progressive feature aggregation modules and global residual connections to obtain richer and deeper features. In this example, n=10 is taken to effectively balance network performance while ensuring the lightweight of the network. The specific operations are as follows:

[0106] F i =f PFABi (F i-1 ),i=1,2,...,n

[0107] F=f Conv3×3 (F n )

[0108] in, represents the i-th progressive feature aggregation module, n represents the number of progressive feature aggregation modules, F i-1 、F i 、F n represents the intermediate feature map of the i-1th, i-th, and n-th progressive feature aggregation modules, f Conv3×3 (·) represents the convolution operation with a kernel size of 3×3, and F represents the deep feature map output by the deep feature extraction module;

[0109] like Figure 2As shown in the figure, the progressive feature aggregation module includes a key information perception module (KIPM), a local feature enhancement module (LFEM), and two feed-forward networks (FFN). The two paths of the key information perception module capture high-frequency information from the channel and spatial dimensions to restore edge detail features. The local feature enhancement module further interacts the captured information globally to enhance texture features. The feed-forward network gradually integrates rich edge details and texture features to improve the reconstruction quality of remote sensing images. The specific operations are as follows:

[0110] Y=FFN(LFEM(FFN(KIPM(X))))+X

[0111] Where X represents the input features of the progressive feature aggregation module, KIPM(·) represents the key information perception module, FFN(·) represents the forward propagation network, LFEM(·) represents the local feature enhancement module, and Y represents the feature map output by the progressive feature aggregation module.

[0112] like Figure 3 As shown in the figure, the dual-path design of the key information perception module can perceive high-frequency information from two different dimensions, channel and space, and combine them to obtain a global important feature weight map, and use the gating mechanism to more accurately restore edge details. The key information perception module includes two 1×1 convolutional layers, a Low Energy Channel Attention (LECA), a Dual Low Energy Spatial Attention (DLESA), and a 7×7 depth-separable convolutional layer. LECA and DLESA are two attention designs that do not introduce parameters, which reduce the number of model parameters while improving efficiency. The specific operations are as follows:

[0113] X1,X2=split(f Conv1×1 (LN(X)))

[0114] Y K =X·Sig(f Conv1×1 (f DWConv7×7 (LECA(X1)+DLESA(X2))))

[0115] Among them, X represents the input features of the key information perception module, LN(·) represents the layer normalization operation, and f Conv1×1(·) represents a 1×1 convolution operation, split(·) represents a channel splitting operation, X1 and X2 represent the feature layers after channel splitting, LECA(·) represents the lowest energy channel attention operation, DLESA(·) represents the dual lowest energy spatial attention operation, and f DWConv7×7 (·) represents a 7×7 depthwise separable convolution operation, Sig(·) represents the Sigmoid function, · represents the channel-wise multiplication operation, and Y K Represents the output features of the key information perception module.

[0116] like Figure 4 As shown in FIG. 1 , the minimum energy channel attention is a channel attention design without introducing parameters, which can effectively reduce the number of parameters to ensure the lightweight of the model. It uses a global average pooling operation to extract features, enhances features based on the method of minimizing the energy function, and integrates important channel information using a gating mechanism. The stable parameter λ of the formula is set according to the application scenario of the method. In the real scenario of this embodiment, λ is set to 10. -4 , the specific operations are as follows:

[0117] X 11 =GP(X1)

[0118]

[0119] Y K1 =X1·Sig(X 12 )

[0120] Among them, X1 represents the input feature of the lowest energy channel attention, GP(·) represents the global average pooling operation, and X 11 ∈R C×1×1 represents the channel vector obtained after the pooling operation, μ represents the average value of the channel vector, σ 2 represents the variance of the channel vector, λ represents the stability parameter of the formula, X 12 ∈R C×1×1 represents the weighted attention map, Sig(·) represents the Sigmoid function, and Y K1 Output features representing the lowest energy channel attention.

[0121] like Figure 5 As shown in FIG. 1 , the dual minimum energy spatial attention is a spatial attention design without introducing parameters, which can effectively reduce the number of parameters to ensure the lightweight of the model. It adopts global information interaction in different directions and uses the method of minimizing the energy function to enhance the spatial features, thereby maximally capturing the spatial high-frequency information. The stable parameter λ of the formula is set according to the application scenario of the method. In the real scenario of this embodiment, λ is set to 10. -4 , the specific operations are as follows:

[0122]

[0123]

[0124] Y K2 =X2·Sig(X 21 +X 22 )

[0125] Among them, X2 represents the input features of the dual minimum energy spatial attention, μ1∈R C×H×1 represents the average value calculated along the spatial width W, σ1 2 ∈R C×H×1 represents the variance calculated along the spatial width W, μ2∈R C×1×W represents the average value calculated along the spatial height H, σ2 2 ∈R C×1×W represents the variance calculated along the spatial height H, λ represents the stability parameter of the formula, X 21 ,X 22 ∈R C×H×W represents the weighted attention map along different spatial directions, Sig(·) represents the Sigmoid function, and Y K2 Output features representing the lowest energy channel attention.

[0126] The high-frequency information extracted by the key information perception module is transmitted to the local feature enhancement module by the forward propagation network, and the local feature enhancement module is used to further restore the complex texture features, thereby gradually aggregating high-frequency edge details and texture features. The forward propagation network includes a layer normalization operation and two 1×1 convolution layers. The specific operations are as follows:

[0127] X F1 ,X F2 = split(f Conv1×1 (LN(X F )))

[0128] Y F =f Conv1×1 (X F1 ·X F2 )+X F

[0129] Among them, X F represents the input features of the forward propagation network, LN(·) represents the layer normalization operation, and f Conv1×1 (·) represents a 1×1 convolution operation, split(·) represents a channel split operation, and X F1 , X F2 Represents the feature layer after channel splitting, Y F Represents the output features of the forward propagation network.

[0130] like Figure 6As shown in the figure, the local feature enhancement module (LFEM) adopts a multi-branch structure with different convolutional layer sizes, which is used to globally interact the image features output by the key information perception module. The local feature enhancement module combines the advantages of multi-scale convolution to extract local features and Transformer long-distance information interaction, and better utilizes the multi-scale features and self-similar information inside the image to help restore texture features. LFEM includes a layer normalization, a local feature aggregation branch, a context-aware local enhancement branch, a large kernel attention branch and a 1×1 convolution layer. The local feature enhancement branch includes a 9×9 depth-separable convolution to extract local features; the context-aware local enhancement branch consists of a 5×5 depth-separable convolution layer, a 5×5 dilated convolution layer, a 7×7 dilated convolution layer, Hardswish, Tanh and two 1×1 convolution layers; the large kernel attention branch includes a 7×7 depth-separable convolution layer, a 9×9 dilated convolution layer and a 1×1 convolution layer. It should be noted that the local feature enhancement module uses depthwise separable convolution, dilated convolution and 1×1 convolution design to reduce the number of model parameters and ensure the lightweight of the module. The specific operations are as follows:

[0131] V=f DWConv9×9 (LN(X L ))

[0132] Q=f DWDConv5×5 (f DWConv5×5 (LN(X L )))

[0133] K=f DWDConv7×7 (f DWConv5×5 (LN(X L )))

[0134] X L1 =Tanh(f Conv1×1 (Hardswish(f Conv1×1 (Q·K))))·V

[0135] X L2 =f Conv1×1 (f DWDConv9×9 (f DWConv7×7 (LN(X L ))))·LN(X L )

[0136] Y L =f Conv1×1 (concat(X L1 ,X L2 ))+X L

[0137] Among them, X Lrepresents the input features of the local feature enhancement module, LN(·) represents the layer normalization operation, and f DWConv9×9 (·) represents a 9×9 depthwise separable convolution operation, f DWConv5×5 (·) represents a 5×5 depthwise separable convolution operation, f DWDConv5×5 (·) represents a 5×5 dilated convolution operation, f DWDConv7×7 (·) represents a 7×7 dilated convolution operation, f Conv1×1 (·) represents a 1×1 convolution operation, Hardswish(·) represents the Hardswish function, Tanh(·) represents the Tanh function, and f DWConv7×7 (·) represents a 7×7 depthwise separable convolution operation, f DWDConv9×9 (·) represents a 9×9 dilated convolution operation, concat(·) represents a channel merging operation, and Y L Represents the output features of the local feature enhancement module.

[0138] S130, reconstructing the deep feature map. The reconstruction module includes a sub-pixel convolution layer and a 3×3 convolution. A sub-pixel convolution layer is used to upsample the deep features extracted by the deep feature extraction module, and the information stream is reorganized into a feature map with a specified upsampling ratio. The specific operations are as follows:

[0139] I SR =f IR (F+F0)+Bilinear(I LR )

[0140] Among them I SR represents the super-resolution image after super-resolution reconstruction, f IR (·) represents the reconstruction module, and Bilinear(·) represents the bilinear interpolation of the low-resolution image to the target resolution.

[0141] In a preferred embodiment, reference Figure 7 In S200, the second-order shuffle random degradation model is established, including two first-order shuffle random degradation models, the first-order shuffle random degradation model includes four degradation categories shuffled, the four degradation categories are blur, noise, compression and scaling, each degradation type occurs randomly with a certain probability P, and the specific operations are as follows:

[0142]

[0143] DR∈{DR1,DR2,DR j ...DR 24}

[0144] DR1={A,B,C,D}1,DR2={A,C,B,D}2,DR j={B,D,A,C} j …DR 24 ={C,A,D,B} 24

[0145]

[0146] X t ~N(μ t ,σ t 2 )

[0147] Among them, I LR represents low-resolution remote sensing images, I HR Represents high-resolution remote sensing images, DR first-order shuffle random degradation model, DR j represents one of the four degradation category shuffle combinations, A, B, C, D represent blur, noise, compression, scaling operation or no operation, P is the probability threshold, X t Obey the standard normal distribution μ t =0,σ t =1.

[0148] In a preferred embodiment, in S300, since remote sensing images in real scenes usually involve complex ground structures, and the degradation process is complex and diverse, it is necessary not only to consider pixel loss, but also to fully consider the important role played by perception loss and structure loss in the reconstruction process, otherwise the reconstructed super-resolution image is likely to have some blurred edges and distorted structures. In summary, weighing the advantages of the three, a super-resolution reconstruction loss function for remote sensing images in real scenes is proposed for training, including:

[0149] S310, the calculation formula of the remote sensing image super-resolution reconstruction loss function for real scenes is as follows:

[0150]

[0151] Among them, L represents the comprehensive loss, L1 represents the pixel loss, and L per represents the perceptual loss, L s Indicates structural loss.

[0152] S320, the pixel loss calculation formula is as follows:

[0153] L1=||I HR -f(I LR ,θ)||1

[0154] Among them, L1 represents pixel loss, I HR represents high-resolution remote sensing images, I LRrepresents the input low-resolution remote sensing image to be reconstructed, f(·,θ) represents the super-resolution reconstruction network with parameters θ, and ||·||1 represents the L1 norm;

[0155] S330, the calculation formula of perception loss is as follows:

[0156]

[0157] Among them, L per represents the perceived loss, I HR represents high-resolution remote sensing images, I LR represents the input low-resolution remote sensing image to be reconstructed, f(·,θ) represents the super-resolution reconstruction network with parameters θ, represents the pre-trained VGG network, ||·||1 represents the L1 norm;

[0158] S340, the structural loss calculation formula is as follows:

[0159]

[0160] L s =1-TSSIM(D·f(I LR ,θ),f(D·I LR ,θ),D·I HR )

[0161] Among them, μ x , μ y , μ z and σ x , σ y , σ z Represent the mean and standard deviation of x, y, and z, σ xy , σ yz , σ xz They represent the correlation coefficients between (x, y), (y, z), and (x, z), respectively. C1 and C2 represent the hyperparameters set to 10 -1 With stable training, L s represents the structural loss, D·f(I LR ,θ) represents the deformed super-resolution reconstructed image, f(D·I LR ,θ) represents the reconstructed image of the deformed low-resolution remote sensing image, D·I HR represents the deformed high-resolution remote sensing image, TSSIM(·) represents the deformation function;

[0162] In order to better illustrate the effectiveness of the present invention, the embodiment of the present invention uses a comparative experiment to compare the reconstruction effects.

[0163] Specifically, the embodiment of the present invention uses the AID data set to divide the training set and the test set into 7:3, wherein the training set obtains a low-resolution remote sensing image that simulates the degradation of the real scene remote sensing image through a second-order shuffle random degradation model, and the obtained low-resolution remote sensing image and the high-resolution remote sensing image in the training set are used as training pairs, and the BSRGAN, Real-ESRGAN and PFAN networks are trained according to unified rules to obtain a trained network. Three common remote sensing image test sets, namely AIDData Set, RSCCN7 Data Set and UC MercedLand-Use Data Set, are selected to test the trained network. The peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and non-reference image quality evaluation index (NIQE) are used to evaluate the model performance. As shown in Table 1, BIC represents bicubic interpolation. In the four-fold super-resolution reconstruction experiment, the PFAN method proposed in the present invention achieves the best test results in the three test sets. Specifically, the average PSNR and SSIM on the three test sets are 0.93dB and 0.0317 higher than BSRGAN, while the NIQE score is 2.58 lower than BSRGAN; the average PSNR and SSIM on the three test sets are 1.02dB and 0.0360 higher than Real-ESRGAN, while the NIQE score is 1.35 lower than Real-ESRGAN. The above results show that the present invention can effectively improve the resolution of remote sensing images of real scenes, retain more image structure information and detail features, and significantly improve its application in real scenes.

[0164] Table 1. Comparison of average PSNR, SSIM and NIQE values ​​on three remote sensing test sets

[0165]

[0166] In summary, the present invention has the following advantages and effects compared with the prior art:

[0167] (1) An efficient and lightweight progressive feature aggregation network is proposed, which can extract rich edge details and texture features through a progressive strategy to achieve high-quality super-resolution reconstruction of remote sensing images;

[0168] (2) The present invention designs a key information perception module and a local feature enhancement module to capture high-frequency information and internal self-similar information in the image. Specifically, the key information perception module contains two parameter-free attention designs, LECA and DLESA, which capture high-frequency information in the channel and spatial dimensions respectively, and help to restore the edge details of the image. Compared with traditional methods, the key information perception module adopts a parameter-free attention design to ensure the lightweight of the module. The local feature enhancement module uses shared and context-aware weights to capture the self-similar information inside the image, which helps to restore the complex texture features of the image. Compared with the traditional Transformer, it gives full play to the ability of multi-scale convolution to extract local features and the Transformer global information interaction;

[0169] (3) A second-order shuffle random degradation model is proposed to fit the generation of low-resolution remote sensing images of real scenes, which effectively improves the quality of super-resolution reconstruction of remote sensing images in real scenes. Compared with traditional methods, the shuffle strategy is used to randomly add degradation categories such as blur, noise, compression, and scaling to simulate the imaging process subject to many interferences, such as system noise, motion blur, and transmission compression, which can more fully fit the degradation of remote sensing images of real scenes;

[0170] (4) A loss function for super-resolution reconstruction of remote sensing images in real scenes is proposed. It fully considers the advantages of pixel loss, perceptual loss and structural loss, avoids some problems that may exist in super-resolution reconstruction of remote sensing images in real scenes, such as erroneous edges and distorted structures, and ensures that the network can better meet the application requirements of real scenes.

[0171] and Figure 1 The method corresponds to and refers to Figure 8 The embodiment of the present invention provides a real scene remote sensing image super-resolution reconstruction system based on a progressive feature aggregation network, comprising:

[0172] At least one computing device, the computing device comprising at least one processor and at least one memory, the memory being used to store at least one program;

[0173] At least one display for displaying the operating results of the computing device;

[0174] When the above program is executed by a processor, the processor can implement the above method.

[0175] It can be seen that the content of the above method embodiment is also applicable to the present system embodiment. The functions implemented by the present system embodiment are the same as those of the above method embodiment, and the beneficial effects achieved are also consistent with those achieved by the method embodiment.

[0176] In addition, an embodiment of the present invention further discloses a computer program product or a computer program, which is stored in a computer-readable storage medium. The processor of the computer device can load and execute the program from the storage medium, so that the computer device completes the operation of the above method. Similarly, the content of the above method embodiment is also applicable to the storage medium embodiment, and the functions implemented by the storage medium embodiment are consistent with those of the method embodiment, and the beneficial effects achieved are also the same as those of the method embodiment.

[0177] It will be appreciated by those skilled in the art that all or part of the methods disclosed above, as well as the system, may be implemented in the form of software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components may be implemented by software executed by a processor (such as a central processing unit, a digital signal processor, or a microprocessor), or by hardware, or as an integrated circuit (such as an application-specific integrated circuit). Such software may be distributed on a computer-readable medium, which includes a computer storage medium (non-transitory medium) and a communication medium (transitory medium). As is well known to those skilled in the art, a computer storage medium refers to various methods or techniques used in storing information (such as computer-readable instructions, data structures, program modules, or other data), including volatile and non-volatile, removable and non-removable media. Examples of computer storage media include, but are not limited to: RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, DVD or other optical disk storage media, magnetic cassettes, tapes, disk storage or other magnetic storage devices, or any other medium that can store the required information and be accessed by a computer. In addition, communication media generally includes computer readable instructions, data structures, program modules, or data in a modulated data signal such as a carrier wave or other transport mechanism and may encompass any media for transmitting information.

[0178] The preferred embodiments of the present disclosure are described in detail above, but the present disclosure is not limited to the above specific embodiments. Those skilled in the art may make various equivalent changes or substitutions without departing from the spirit and scope of the present invention, and these equivalent changes or substitutions also fall within the scope defined by the claims of the present disclosure.

Claims

1. A real scene remote sensing image super-resolution reconstruction method based on a progressive feature aggregation network, characterized in that: The method comprises the following steps: S100, establish a progressive feature aggregation network PFAN for super-resolution reconstruction of real scene remote sensing images; S200, establishing a second-order shuffle random degradation model, subjecting the remote sensing image used as a training sample to second-order shuffle random degradation, and obtaining a low-resolution remote sensing image that simulates a random degradation process of a real scene; S300, inputting the remote sensing image before degradation and the corresponding low-resolution remote sensing image into the PFAN network for training; S400, obtaining a real remote sensing image to be reconstructed, inputting the trained PFAN network to perform image reconstruction, and obtaining a super-resolution image.

2. According to claim 1, a real scene remote sensing image super-resolution reconstruction method based on progressive feature aggregation network is characterized in that: The PFAN network includes a shallow feature extraction module, a deep feature extraction module and a reconstruction module. The shallow feature extraction module is a 3×3 convolutional layer. The shallow feature extraction module is described as: F0=f SFE (IN LR ) where f SFE (·) represents the shallow feature extraction module, I LR represents a low-resolution remote sensing image, and F0 represents a shallow feature map; The deep feature extraction module includes several progressive feature aggregation modules and a 3×3 convolutional layer to extract high-level features from shallow features. The process is expressed as follows: F=f Conv3×3 (F n ) in, represents the i-th progressive feature aggregation module, n represents the number of progressive feature aggregation modules, F i-1 、F i 、F n represents the intermediate feature map of the i-1th, i-th, and n-th progressive feature aggregation modules, f Conv3×3 (·) represents the convolution operation with a kernel size of 3×3, and F represents the deep feature map; The reconstruction module includes a sub-pixel convolution layer and a 3×3 convolution layer. A sub-pixel convolution layer is used to upsample the deep features extracted by the deep feature extraction module, and the information flow is reorganized into a feature map with a specified upsampling ratio. The process is described as follows: I SR =f IR (F+F0)+Bilinear(I LR ) Among them I SR represents the super-resolution image after super-resolution reconstruction, f IR (·) represents the reconstruction module, and Bilinear(·) represents the bilinear interpolation of the low-resolution image to the target resolution.

3. A real scene remote sensing image super-resolution reconstruction method based on a progressive feature aggregation network according to claim 2, characterized in that: The progressive feature aggregation module includes a key information perception module, a local feature enhancement module, two forward propagation networks and residual connections. The specific calculation formula is as follows: Y=FFN(LFEM(FFN(KIPM(X))))+X Where X represents the input features of the progressive feature aggregation module, KIPM(·) represents the key information perception module, FFN(·) represents the forward propagation network, LFEM(·) represents the local feature enhancement module, and Y represents the output features of the progressive feature aggregation module.

4. The real scene remote sensing image super-resolution reconstruction method based on progressive feature aggregation network according to claim 3 is characterized in that: The key information perception module includes a layer normalization operation, two 1×1 convolutional layers, channel splitting, a minimum energy channel attention, a double minimum energy spatial attention, and a 7×7 depth-separable convolutional layer. The specific operations are as follows: X1,X2=split(f Conv1×1 (LN(X))) Y K =X·Sig(f Conv1×1 (f DWConv7×7 (LECA(X1)+DLESA(X2)))) Among them, X represents the input features of the key information perception module, LN(·) represents the layer normalization operation, and f Conv1×1 (·) represents a 1×1 convolution operation, split(·) represents a channel splitting operation, X1 and X2 represent the feature layers after channel splitting, LECA(·) represents the lowest energy channel attention operation, DLESA(·) represents the dual lowest energy spatial attention operation, and f DWConv7×7 (·) represents a 7×7 depthwise separable convolution operation, Sig(·) represents the Sigmoid function, · represents the channel-wise multiplication operation, and Y K Represents the output features of the key information perception module; The minimum energy channel attention uses global average pooling to extract features, and uses the energy function minimization method to enhance features and integrate important channel information. The specific operations are as follows: X 11 =GP(X1) AND K1 =X1·Sig(X 12 ) Among them, X1 represents the input feature of the lowest energy channel attention, GP(·) represents the global average pooling operation, and X 11 ∈R C ×1×1 represents the channel vector obtained after the pooling operation, μ represents the average value of the channel vector, σ 2 represents the variance of the channel vector, λ represents the stability parameter of the formula, X 12 ∈R C×1×1 represents the weighted attention map, Sig(·) represents the Sigmoid function, and Y K1 Output features representing the lowest energy channel attention; The dual minimum energy spatial attention enhances spatial features by minimizing the energy function. The specific operations are as follows: Y K2 =X2·Sig(X 21 +X 22 ) Among them, X2 represents the input features of the dual minimum energy spatial attention, μ1∈R C×H×1 represents the average value calculated along the spatial width W, σ1 2 ∈R C×H×1 represents the variance calculated along the spatial width W, μ2∈R C×1×W represents the average value calculated along the spatial height H, σ2 2 ∈R C×1×W represents the variance calculated along the spatial height H, λ represents the stability parameter of the formula, X 21 ,X 22 ∈R C×H×W represents the weighted attention map along different spatial directions, Sig(·) represents the Sigmoid function, and Y K2 Output features representing the lowest energy channel attention.

5. The real scene remote sensing image super-resolution reconstruction method based on progressive feature aggregation network according to claim 3 is characterized in that: The local feature enhancement module adopts a multi-branch structure with different convolutional layer sizes, and is used to globally interact the image features output by the key information perception module. The local feature enhancement module includes a layer normalization, a local feature aggregation branch, a context-aware local enhancement branch, a large core attention branch, channel merging and a 1×1 convolutional layer; The local feature aggregation branch extracts local features by a 9×9 depthwise separable convolution; The context-aware local enhancement branch includes a 5×5 depthwise separable convolutional layer, a 5×5 dilated convolutional layer, a 7×7 dilated convolutional layer, Hardswish, Tanh, and two 1×1 convolutional layers; The large core attention branch includes a 7×7 depthwise separable convolutional layer, a 9×9 dilated convolutional layer, and a 1×1 convolutional layer; the specific operations are as follows: V=f DWConv9×9 (LN(X L )) Q=f DWDConv5×5 (f DWConv5×5 (LN(X L ))) K=f DWDConv7×7 (f DWConv5×5 (LN(X L ))) X L1 =Tanh(f Conv1×1 (Hardswish(f Conv1×1 (Q·K))))·V X L2 =f Conv1×1 (f DWDConv9×9 (f DWConv7×7 (LN(X L ))))·LN(X L ) Y L =f Conv1×1 (concat(X L1 ,X L2 ))+X L Among them, X L represents the input features of the local feature enhancement module, LN(·) represents the layer normalization operation, and f DWConv9×9 (·) represents a 9×9 depthwise separable convolution operation, f DWConv5×5 (·) represents a 5×5 depthwise separable convolution operation, f DWDConv5×5 (·) represents a 5×5 dilated convolution operation, f DWDConv7×7 (·) represents a 7×7 dilated convolution operation, f Conv1×1 (·) represents a 1×1 convolution operation, Hardswish(·) represents the Hardswish function, Tanh(·) represents the Tanh function, and f DWConv7×7 (·) represents a 7×7 depthwise separable convolution operation, f DWDConv9×9 (·) represents a 9×9 dilated convolution operation, concat(·) represents a channel merging operation, and Y L Represents the output features of the local feature enhancement module.

6. The real scene remote sensing image super-resolution reconstruction method based on progressive feature aggregation network according to claim 3 is characterized in that: The forward propagation network includes layer normalization operations, two 1×1 convolutional layers and channel splitting, and the specific operations are as follows: X F1 ,X F2 =split(f Conv1×1 (LN(X F ))) Y F =f Conv1×1 (X F1 ·X F2 )+X F Among them, X F represents the input features of the forward propagation network, LN(·) represents the layer normalization operation, and f Conv1×1 (·) represents a 1×1 convolution operation, split(·) represents a channel split operation, and X F1 , X F2 Represents the feature layer after channel splitting, Y F Represents the output features of the forward propagation network.

7. The real scene remote sensing image super-resolution reconstruction method based on progressive feature aggregation network according to claim 1 is characterized in that: PFAN network training uses a remote sensing image super-resolution reconstruction loss function for real scenes that combines pixel loss, perceptual loss, and structural loss: Among them, L represents the comprehensive loss, L1 represents the pixel loss, and L per represents the perceptual loss, L s Indicates structural loss; The pixel loss calculation formula is as follows: L1=||I HR -f(I LR ,θ)||1 Among them, I HR represents high-resolution remote sensing images, I LR represents the input low-resolution remote sensing image to be reconstructed, f(·,θ) represents the super-resolution reconstruction network with parameters θ, and ||·||1 represents the L1 norm; The perceptual loss calculation formula is as follows: in, represents the pre-trained VGG network, ||·||1 represents the L1 norm; The formula for calculating structural loss is as follows: L s =1-TSSIM(D·f(I LR ,θ),f(D·I LR ,θ),D·I HR ) Among them, μ x , μ y , μ z and σ x , σ y , σ z Represent the mean and standard deviation of x, y, and z, σ xy , σ yz , σ xz Respectively represent the correlation coefficients between (x, y), (y, z), (x, z), C1, C2 represent hyperparameters, D·f(I LR ,θ) represents the deformed super-resolution reconstructed image, f(D·I LR ,θ) represents the reconstructed image of the deformed low-resolution remote sensing image, D·I HR represents the deformed high-resolution remote sensing image, and TSSIM(·) represents the deformation function.

8. The real scene remote sensing image super-resolution reconstruction method based on progressive feature aggregation network according to claim 1 is characterized in that: The second-order shuffle random degradation model includes two first-order shuffle random degradation models. The first-order shuffle random degradation model includes four degradation categories shuffled, the four degradation categories are blur, noise, compression and scaling, and each degradation type occurs randomly with a certain probability P. The specific operation is as follows: DR∈{DR1,DR2,DR j ...DR 24 } DR1={A,B,C,D}1,DR2={A,C,B,D}2,DR j ={B,D,A,C} j …DR 24 ={C,A,D,B} 24 X t ~N(μ t ,s t 2 ) Among them, I LR represents low-resolution remote sensing images, I HR Represents high-resolution remote sensing images, DR first-order shuffle random degradation model, DR j represents one of the four degradation category shuffle combinations, A, B, C, D represent blur, noise, compression, scaling operation or no operation, P is the probability threshold, X t Obey the standard normal distribution μ t =0,σ t =1.

9. A real scene remote sensing image super-resolution reconstruction system based on progressive feature aggregation network, characterized in that: The system comprises: At least one computing device, the computing device comprising at least one processor and at least one memory, the memory being used to store at least one program; at least one display, the display being used to display the operating results of the computing device; When the at least one program is executed by the computing device, the computing device implements the real-scene remote sensing image super-resolution reconstruction method based on a progressive feature aggregation network as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 8 when executed by the processor.

Citation Information

Cited By

  • Infrared image super-resolution reconstruction method and system

    CN120318075A

  • An infrared image super-resolution reconstruction method and system

    CN120318075B

  • Remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment

    CN120765933A