A Single Remote Sensing Image Super-Resolution Method Based on Deep Learning

By adopting deep learning methods in remote sensing image super-resolution technology, combining adversarial generation networks and multi-level residual blocks, and using pre-trained VGG networks for feature extraction, the problem of poor model complexity and perception effects in the prior art is solved, and a more efficient super-resolution reconstruction effect of remote sensing image is achieved.

CN116342392BActive Publication Date: 2025-05-30LANZHOU UNIV

Patent Information

Application Number
CN202310344467.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-03
Publication Date
2025-05-30
Estimated Expiration
2043-04-03

AI Technical Summary

Technical Problem

The existing remote sensing image super-resolution technology has problems such as over-complex models, high time costs, and poor generalization capabilities of models. The pixel-based optimization method results in the image being too smooth, lacking high-frequency texture details, and poor perception effect.

Method used

A single remote sensing image super-resolution method based on deep learning is used to generate low-resolution images using a bi-cubic kernel function, combined with the SRResNet super-resolution model, an adversarial generation structure and multi-level residual blocks are added, and a pre-trained 19-layer VGG network is used for feature extraction, and the perceived loss function is defined to improve the model reconstruction effect.

Benefits of technology

The model reconstruction effect is significantly improved, making the generated remote sensing image closer to the human eye perception effect, has stronger semantic information capture capabilities, and effectively reduces noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342392B_ABST
    Figure CN116342392B_ABST
Patent Text Reader

Abstract

The present invention discloses a single remote sensing image super-resolution method based on deep learning (Single Remote Sensing Image Super-Resolution, SRSISR). Compared with traditional methods and methods based on convolutional neural networks, this method can significantly improve the resolution of remote sensing images. On the basis of the pixel-based loss function, a loss function closer to perceptual similarity is added, making the result closer to the human eye perception effect and significantly improving the model reconstruction effect. The single remote sensing image super-resolution (SRSISR) based on deep learning improves the spatial resolution of low-resolution satellite images and provides an effective solution for applications such as image processing and denoising. Thus, it can be seen that using deep learning for remote sensing image super-resolution reconstruction has important research significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of remote sensing images and deep learning, and particularly to improving the quality of remote sensing images through an image super-resolution technology in computer vision. Background Art

[0002] Remote sensing image super-resolution reconstruction is a technology for processing low-resolution remote sensing images (LRRSI) with semantic information to obtain high-resolution remote sensing images (HRRSI). Remote sensing images are the data support and application basis in remote sensing technology, playing an important role in aspects such as pattern classification, target detection, and ground object extraction, providing a large amount of information for surface monitoring, and having broad application prospects in aspects such as urban expansion analysis, disaster monitoring and evaluation, and geological resource exploration. Due to the limitations of optical and sensor technologies and high costs, the spatial and spectral resolutions of earth observation satellites generally fail to meet the expected effects; therefore, it is of crucial significance to develop software-based algorithms to improve the spatial and spectral quality of satellite images. Single remote sensing image super-resolution (SRSISR) provides an effective solution for improving the spatial resolution of low-resolution satellite image inputs and enhancing image processing applications and denoising performance. Thus, the research on a remote sensing image super-resolution reconstruction technology based on software algorithms is of great importance.

[0003] For remote sensing images, super resolution (SR) can be regarded as an inverse solution process of blurring and downsampling. Various methods for image SR have been proposed. Traditional methods such as frequency domain methods and spatial domain methods, although achieving good results, still have great limitations, such as complex models, high time costs, and poor model generalization ability, etc., so they are no longer the mainstream research methods at present.

[0004] Sparse representation has also been introduced into super-resolution. By using a sparse coding algorithm to learn the correlation between HR and LR images from the HR-LR pair dataset, and then applying it to the restoration of HR images. The learning-based method significantly improves the performance; however, a large amount of computing resources are required during both the training of the dictionary and the reconstruction process of the image, resulting in high costs.

[0005] In recent years, deep learning has developed rapidly. In the field of natural image super-resolution, deep learning has significantly improved the reconstruction effect of models, demonstrating its superiority. Therefore, many scholars have borrowed the experience of natural image super-resolution reconstruction and applied it to the field of remote sensing images, achieving good results. Currently, in the field of deep learning, the commonly used remote sensing image super-resolution methods are mainly based on convolutional neural network methods. Compared with traditional methods, they are fast and have a simple structure, but there are also certain limitations. For example, pixel-based optimization methods can lead to overly smooth images, lacking high-frequency texture details and not having a better perceptual effect. Based on this, we propose a remote sensing image super-resolution method based on generative adversarial networks, adding a loss function closer to perceptual similarity to make the results closer to the human eye's perceptual effect and significantly improving the model's reconstruction effect. Summary of the Invention

[0006] Aiming at the problems in the related technologies, the present invention proposes a single remote sensing image super-resolution method based on deep learning to overcome the above-mentioned technical problems existing in the existing related technologies.

[0007] To this end, the specific technical solution adopted by the present invention is as follows: A single remote sensing image super-resolution method based on deep learning, comprising the following steps:

[0008] A. Downsample the high-resolution remote sensing sample image using a bicubic kernel function to generate a corresponding low-resolution image, and use the original high-resolution image as the corresponding label image. Use the pair of high- and low-resolution remote sensing images as training data;

[0009] B. Divide the pair of high- and low-resolution images into a training set and a validation set according to a ratio, and perform data augmentation on the training set;

[0010] C. Input the training set into a deep learning-based super-resolution model for training, and use the validation set to evaluate the fast detection model to obtain an optimal network parameter model; select the SRResNet super-resolution model as the network basic model, select MSE as the basic loss function; add a generative adversarial structure and adversarial loss to the network; replace the single residual block with a multi-level residual block to improve the model's reconstruction effect; and use a pre-trained 19-layer VGG network for feature extraction to capture high-level perceptual differences, and define the VGG loss as the Euclidean distance between the feature representations of the reconstructed image G net (LR) and the reference image HR, and use this as the perceptual loss;

[0011] The formula of the MSE loss function is as follows:

[0012]

[0013] where, HR i,j,cRepresents the reference high - resolution image, where i, j, and c represent the positions of the length, width, and height of the pixel in the image respectively. LR represents the low - resolution image, and G net Denotes the generator network, G net (LR) i,j,c Represents the high - resolution image output after the low - resolution image is processed by the generator network model. H represents the height of the low - resolution image, W represents the width of the low - resolution image, num_channel represents the channel value of the low - resolution image, and R represents the magnification factor;

[0014] The adversarial loss function is as follows:

[0015]

[0016] Among them, D net Denotes the discriminator network, D net (HR, G net (LR)) represents the probability that the real data is more real than the fake generated data. D net (G net (LR), HR) represents the probability that the fake generated data is more fake than the real data. N represents the number of samples; The original high - resolution image is the reference image (true high - resolution), and the super - resolution image is the high - resolution result (false high - resolution) output by the low - resolution image through the super - resolution network model. Both of them are fed into the discriminator network model, aiming to enable the discriminator network to identify as much as possible whether the generated data is more real than the real data.

[0017] The perceptual loss function is as follows:

[0018]

[0019] Among them, E net Denotes the feature extractor network, E net (HR) represents the output of the feature extractor network in the high - resolution HR branch. E net (G net (LR)) represents the output of the feature extractor network in the super - resolution SR branch. H n(p,q) , W n(p,q) and C n(p,q) respectively represent the height, width, and number of channels of the n - th high - level feature map in the high - level feature network; The feature extractor network here refers to the pre - trained VGG19 network. Both the "true" high - resolution image and the "false" high - resolution image (the super - resolution result after passing through the Gnet network) are fed into the feature extractor network to extract high - level features, and the Euclidean distance between them in the n - th feature map is calculated;

[0020] The total loss function is as follows:

[0021] loss = loss_MSE + loss_adv + loss_per;

[0022] D. Input the remote sensing images of the test set into the trained optimal network parameter model to obtain the final result.

[0023] The present invention uses a residual network as the backbone structure of the generator network, and uses multi-level residual blocks as the basic network units. Based on the MSE loss, the generator structure is trained to extract basic features. A discriminator structure is added to the network, and a relative rather than an absolute discriminator is used to capture more realistic texture details. In addition, a pre-trained 19-layer VGG network is used in the model for high-level feature extraction, optimizing the extraction method, improving the extraction effect, and effectively reducing noise. In addition to the improved network structure, considering that the network is deep and complex, before adding the residual to the main path, the residual is reduced by multiplying a constant less than 1, thereby reducing the training instability, and using a smaller initialization makes the residual architecture easier to train.

[0024] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0025] Based on the adversarial generative network structure, the present invention greatly improves the model reconstruction effect; uses a pre-trained 19-layer VGG network for feature extraction, and uses a loss function closer to perceptual similarity to capture high-level perceptual differences; the present invention uses multi-level residual blocks instead of a single residual block as the basic unit of the network generator, making it have a deeper and more complex network structure, enabling stronger semantic information capture ability for regular structures and effectively reducing noise; improvements are also made in the discriminator part of the model, using a relative discriminator instead of an absolute discriminator (the relative discriminator predicts the probability that a real image is more real than a fake image), and this modification to the discriminator helps to learn clearer edges and more detailed textures, further improving the model effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1 It is the network structure diagram of the present invention. In the figure, Conv: Convolutional layer, Dense Block1: Dense connection block, Upsampling: Upsampling.

[0028] Figure 2 It is the adversarial generative network structure adopted by the present invention.

[0029] Figure 3 It is a comparison of the proposed method with other methods on dataset (1). Dataset (1) comes from airplane62 in the UC Merced dataset, with the format of tif, and the size is 256×256.

[0030] Figure 4 It is a comparison of the proposed method with other methods on dataset (2). Dataset (2) comes from buildings26 in the NWPU-RESISC45 dataset, with the format of tif and the size is 256×256.

[0031] Note: ground truth (true value), Bicubic (bicubic interpolation), SRCNN (Super-Resolution Convolutional Neural Network), SRResNet (Super-Resolution Residual Network), SRGAN (Super-Resolution Generative Adversarial Network), DBNet (Dense Residual Network), DSRGAN (Improved Super-Resolution Generative Adversarial Network). Detailed implementation manners

[0032] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0033] Taking the publicly available land use dataset UC Merced dataset as an example in this embodiment, the process of super-resolution of remote sensing images using the present invention is described in detail.

[0034] According to Figure 1 the method flow chart shown, the super-resolution method based on a single remote sensing image in this embodiment includes the steps:

[0035] Step A: Downsample the high-resolution remote sensing sample image using the bicubic kernel function to generate a corresponding low-resolution image with a size of 64×64, and use the original high-resolution image as the corresponding label with a size of 256×256. The high- and low-resolution remote sensing image pairs are used as training data. The training data used includes 21 categories, and each category contains 100 images. Specifically, this step includes:

[0036] A1: Select remote sensing images with different category scenes as training samples. The training set images include 21 land use types, a total of 2100 images;

[0037] A2: The size of the original high-resolution image is 256×256, the size of the downsampled low-resolution image is 64×64, and the scaling ratio is 4;

[0038] A3: During the training stage, the high-resolution image is cropped to 128×128, which is the size of the corresponding label.

[0039] Step B: Divide the high- and low-resolution image pairs into a training set and a validation set at a ratio of 8:2, and perform data augmentation on the training set. Step B specifically includes:

[0040] B1: Divide the high- and low-resolution image pairs into a training set and a validation set at a ratio of 8:2;

[0041] B2: Perform image augmentation on the image pairs in the training set, including operations such as random flipping and rotation.

[0042] Step C: Input the training set into a deep learning-based super-resolution model for training, and use the validation set to evaluate the fast detection model to obtain an optimal network parameter model; select the SRResNet super-resolution model as the network basic model, and select MSE as the basic loss function; add an adversarial generation structure to the network to make the results more real and reliable, the adversarial loss; and replace the single residual block with a multi-level residual block to make the network have a deeper and more complex structure, thereby improving the model reconstruction effect; and use a pre-trained 19-layer VGG network for feature extraction to capture high-level perceptual differences, and define the VGG loss as the Euclidean distance between the feature map in the high-level feature space and the reference map, and use this as the perceptual loss.

[0043] The formula for the MSE loss function is as follows:

[0044]

[0045] Among them, HR i,j,c represents the reference high-resolution image, i, j, and c respectively represent the positions of the length, width, and height of the pixel in the image in the image, LR represents the low-resolution image, and G net represents the generator network, and G net (LR) i,j,c represents the high-resolution image output after the low-resolution image is processed by the generator network model. H represents the height of the low-resolution image, W represents the width of the low-resolution image, num_channel represents the channel value of the low-resolution image, and R represents the magnification factor.

[0046] The adversarial loss function is as follows:

[0047]

[0048] Among them, D net represents the discriminator network, and D net (HR,G net (LR)) represents the probability that the real data is more real than the fake generated data, and D net (G net (LR),HR) represents the probability that the fake generated data is more fake than the real data, and N represents the number of samples;

[0049] The perceptual loss function is as follows:

[0050]

[0051] Among them, E net represents the feature extractor network, and E net (HR) represents the output of the feature extractor network in the high-resolution HR branch, and E net (G net (LR)) represents the output of the feature extractor network in the super-resolution SR branch. H n(p,q) , W n(p,q) and C n(p,q) respectively represent the height, width, and number of channels of the high-level feature map of the nth layer in the high-level feature network (the feature map obtained from the qth convolutional layer before the nth maximum pooling layer in the 19-layer VGG network);

[0052] The total loss function is as follows:

[0053] loss = loss_MSE + loss_adv + loss_per

[0054] Specifically, step C is executed according to the following steps:

[0055] C1. The model training consists of two parts, namely the generator model and the discriminator model. First, train the generator part and construct the residual network generator model (SRResNet);

[0056] C2. Input the training dataset into the SRResNet generator model for training, use MSE as the loss function for iterative training and update the network parameters, and perform 1000 generations of training in total;

[0057] C3. Use the Peak Signal to Noise Ratio (PSNR) and the Structural Similarity Index (SSIM) as the evaluation metrics for the validation set to evaluate the model performance, and select the network parameters with the best performance on the validation set as the final generator model for saving.

[0058] C4. After the generator model training is completed, train the discriminator model to adjust the network generator, add perceptual loss and adversarial loss to improve the high-frequency texture details and reduce noise, use the Natural Image Quality Evaluator (NIQE) as the model evaluation metric, and select the model with the best performance on the validation set as the final model for saving.

[0059] Step D: Input the low-resolution remote sensing images of the test set into the trained optimal network parameter model to obtain the final result.

[0060] Specifically, Step D includes the following steps:

[0061] D1: Select the remote sensing images to be tested;

[0062] D2: Crop the remote sensing images to an appropriate size, and use the bicubic kernel function for downsampling to obtain the corresponding low-resolution images;

[0063] D3: Input the low-resolution images into the model for processing to obtain the super-resolved images, and use the original high-resolution images as a reference to obtain various index parameters.

[0064] In this embodiment, the proposed network is compared with Bicubic, SRCNN, ResNet, and SRGAN, and PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and NIQE (Natural Image Quality Evaluator) are selected as evaluation indicators.

[0065] Evaluation indicators:

[0066] Peak Signal-to-Noise Ratio (PSNR): An objective standard for evaluating images. The larger the value, the smaller the image distortion.

[0067] Structural Similarity Index (SSIM): An index for measuring the similarity between two images, with a value range of [0, 1]. The larger the value, the smaller the image distortion.

[0068] Natural Image Quality Evaluator (NIQE): A lower NIQE value corresponds to a higher overall naturalness.

[0069] Since the proposed network DBNet is optimized based on pixels, PSNR and SSIM are used to evaluate the model performance; while the proposed network DSRGAN adds perceptual loss and adversarial loss to the loss function, so the evaluation indicator NIQE, which is more in line with the overall naturalness, is used.

[0070] Table 1 is the quantitative comparison of various indicators of DBNet and other methods on dataset (1)

[0071] Table 1

[0072]

[0073] Table 2 is the quantitative comparison of indicators of DSRGAN and other methods on dataset (1)

[0074] Table 2

[0075]

[0076] Table 3 shows the quantitative comparison of various indicators between DBNet and other methods on dataset (2).

[0077] Table 3

[0078]

[0079] Table 4 shows the quantitative comparison of indicators between DSRGAN and other methods on dataset (2).

[0080] Table 4

[0081]

[0082] From the qualitative results, as Figure 3 and Figure 4 , for the proposed network DBNet, compared with Bicubic, SRCNN, and SRResNet, it has clearer textures and less noise. This is due to the deeper and more complex structure of DBNet, which has a stronger ability to extract semantic information. Especially for regular structures like the roof in Figure 4 ; for the proposed network DSRGAN, compared with all the other methods, it has a more realistic effect that is closer to human visual perception, and at the same time has richer textures and details. This is due to the addition of a perceptual loss function and an adversarial loss function that are more perceptually similar in the network.

[0083] From the quantitative results, for DBNet, as shown in Table 1 and Table 3, the two indicators of PSNR and SSIM on the two datasets are higher than those of the other methods. Replacing the single-connection network SRResNet with the densely connected network DBNet, on the two datasets, the PSNR is respectively 0.3409 and 0.5019 higher, and the SSIM is respectively 0.0338 and 0.0452 higher; for DSRGAN, the NIQE on the two datasets is lower than that of the other methods (a lower NIQE value represents higher overall naturalness). As shown in Table 2 and Table 4, compared with SRGAN, the proposed network DSRGAN on the two datasets is respectively 0.3540 and 0.6232 lower, a decrease of 7.9% and 12.77% respectively.

[0084] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A single remote sensing image super-resolution method based on deep learning, including the following steps: A. Downsample the high-resolution remote sensing sample image using the bicubic kernel function to generate the corresponding low-resolution image, and use the original high-resolution image as the corresponding label image. Use the pair of high- and low-resolution remote sensing images as training data; B. Divide the pair of high- and low-resolution images into a training set and a validation set according to a ratio, and perform data augmentation on the training set; C. Input the training set into the super-resolution model based on deep learning for training, and use the validation set to evaluate the fast detection model to obtain the optimal network parameter model; The network basic model selects the SRResNet super-resolution model, and selects MSE as the basic loss function; Add the adversarial generative network structure to the network, and replace the single residual block with multiple residual blocks to improve the reconstruction effect of the model generator; Use the pre-trained 19-layer VGG network for feature extraction to capture the high-level perceptual difference, and define the VGG loss as the Euclidean distance between the feature representations of the reconstructed image G net (LR) and the reference image HR, and use this as the perceptual loss; The formula of the MSE loss function is as follows: Among them, HR i,j,c represents the reference high-resolution image, where i, j, and c represent the positions of the length, width, and height of the pixel in the image respectively, LR represents the low-resolution image, Gnet represents the generator network, and G net (LR) i,j,c represents the high-resolution image output after the low-resolution image is processed by the generator network model. H represents the height of the low-resolution image, W represents the width of the low-resolution image, num_channel represents the channel value of the low-resolution image, and R represents the magnification factor; The adversarial loss function is as follows: Among them, D net represents the discriminator network, D net (HR, G net (LR)) represents the probability that the real data is more real than the fake generated data, D net (G net (LR), HR) represents the probability that the fake generated data is more fake than the real data, where N represents the number of samples; The perceptual loss function is as follows: Among them, E net represents the feature extractor network, and E net (HR) represents the output of the feature extractor network on the high-resolution HR branch, and E net (G net (LR)) represents the output of the feature extractor network on the super-resolution SR branch, H n(p,q) , W n(p,q) and C n(p,q) respectively represent the height, width, and number of channels of the high-level feature map of the nth layer in the high-level feature network, where the nth layer is the feature map obtained by the qth convolutional layer before the nth maximum pooling layer in the 19-layer VGG network; The total loss function is as follows: loss = loss_MSE + loss_adv + loss_per D. Input the remote sensing image of the test set into the trained optimal network parameter model to obtain the final result.

2. The single remote sensing image super-resolution method based on deep learning according to claim 1, characterized in that, In step A, the size of the original high-resolution image is 256×256, the size of the downsampled low-resolution image is 64×64, and the scaling ratio is 4; select remote sensing images with different category scenes as training samples. The training set images include 21 land use types, each category contains 100 images, for a total of 2100 images; during the training phase, the high-resolution images are cropped to 128×128, that is, the corresponding label size.

3. The single remote sensing image super-resolution method based on deep learning according to claim 1, characterized in that, The specific steps of step B include: B1. Divide the pair of high- and low-resolution images into a training set and a validation set according to a ratio; B2. Perform image augmentation on the image pairs in the training set, including random flipping and rotation operations.

4. The single remote sensing image super-resolution method based on deep learning according to claim 1, characterized in that, The specific steps of step C include: C1. The model training includes two parts, namely the generator model and the discriminator model. First, train the generator model part and construct the residual network generator model SRResNet; C2. Input the training data set into the residual network generator model SRResNet for training, use MSE as the loss function, use the Adam optimization algorithm for iterative training, set the initial learning rate to 0.001, the batch size to 32, and perform 1000 generations of training in total; C3. Use the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) as the evaluation indicators for the validation set to evaluate the model performance, and select the network parameters with the best performance on the validation set as the final generator model for saving; C3. After the generator model training is completed, train the discriminator model to adjust the network generator, add perceptual loss and adversarial loss to improve the high-frequency texture details and reduce noise, use the natural image quality evaluator NIQE as the model evaluation indicator, and select the model with the best performance on the validation set as the final model for saving.

5. The single remote sensing image super-resolution method based on deep learning according to claim 1, characterized in that, The specific steps of step D include: D1. Select the remote sensing image to be tested; D2. Crop the remote sensing image to an appropriate size, and perform downsampling using the bicubic kernel function to obtain the corresponding low-resolution image; D3. Input the low-resolution image into the model for processing to obtain the super-resolved image, and obtain various index parameters with the original high-resolution image as a reference.

6. A single remote sensing image super-resolution method based on deep learning as claimed in claim 3, characterized in that, in step B1, the ratio of the divided training set to the test set is 8:

2.

7. A single remote sensing image super-resolution method based on deep learning as claimed in claim 4, characterized in that, in step C3, the model is constructed based on the Python environment, and Pytorch is used as the deep learning framework; the pre-trained generator model is added to the discriminator model to further adjust the generator model; the perceptual loss is defined using the ReLU activation layer in the pre-trained 19-layer VGG network, and n(p,q) represents the feature map obtained from the q-th convolutional layer before the n-th max pooling layer in the 19-layer VGG network after activation.

8. A single remote sensing image super-resolution method based on deep learning as claimed in claim 5, characterized in that, in step D3, the RGB channel is selected for index evaluation.

Citation Information

Patent Citations

  • Method for automatically and optimally selecting remote sensing image segmentation parameters based on regional inconsistency evaluation

    CN106651861A

  • A satellite image super-resolution method based on adversarial network and aerial image a priori

    CN109035142A

Cited By

  • Lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel cavity coordinate attention

    CN120876231A