Remote sensing image super-resolution high-power lifelike high-definition image reconstruction method

By combining approximate difference multi-level networking and deep recursive convolution networking and split generation networking for supervision and learning, the problem of excessive parameters and smooth generation results in the super-resolution reconstruction of remote sensing images is solved, and efficient and accurate remote sensing image reconstruction is achieved.

CN120339066APending Publication Date: 2025-07-18吴山丹
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510404692.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the remote sensing image super-resolution reconstruction algorithm has the problem of excessive convolutional neural networking parameters, high training difficulty, and using mean square error as the loss model leads to smooth generation results, lack of high-frequency information, and poor reconstruction effect.

Method used

Combining approximate difference multi-level networking and deep recursive convolution networking, an image super-resolution reconstruction network is constructed, and the intermediate layer output results are used as a loss model on the pre-trained VGG networking, and supervised learning is performed in combination with split generation networking, and the feature differences between the reconstruction results and the real image are extracted.

Benefits of technology

The reconstruction efficiency and quality of remote sensing images are improved, the generated results are closer to the actual high-resolution images, and the performance of high-frequency information is better, and the smoothness of the generated results is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339066A_ABST
    Figure CN120339066A_ABST
Patent Text Reader

Abstract

According to the remote sensing image super-resolution high-power lifelike high-definition image reconstruction method, approximation difference multi-level networking and deep recursion convolution networking are fused, the super-resolution reconstruction network of the image is constructed, and in the pre-trained VGG networking, the output result of the middle layer is used to reconstruct the super-resolution reconstruction network of the image. The feature level difference between the super-resolution reconstruction result and a real high-resolution image is extracted, and the super-resolution reconstruction of the image is carried out by adopting a split generation network, so that the problem that the generation result is missing in image high-frequency information and detail information due to a loss model based on a mean square error is solved; semi-supervised split learning is changed into supervised learning, so that the super-resolution reconstructs the network, and feature differences between corresponding low-resolution images and high-resolution images are gradually learned in split learning; the super-resolution reconstruction algorithm based on supervised split generation networking is good in high-multiple reconstruction effect, high in remote sensing image reconstruction efficiency and less in distortion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to an image super-resolution reconstruction method, in particular to a remote sensing image super-resolution high-magnification realistic high-definition image reconstruction method, belonging to the technical field of image super-resolution reconstruction. Background Technique

[0002] Image super-resolution reconstruction refers to the process of reconstructing a higher-resolution image SR from one or more input low-resolution images. Image super-resolution reconstruction is a fundamental and important issue in the field of computer vision. Image super-resolution reconstruction has important application values in many application fields such as medical image analysis, video surveillance, high-definition TV conversion, remote sensing mapping, astronomical observation, and star image positioning.

[0003] Due to the limitations of the actual usage scenarios, the size of the pixels of the image acquisition device affects the size of the acquired image. The influence of the surrounding environment light and shadow during image acquisition will cause optical blurring of the image. The jitter and movement that may occur during acquisition will cause motion blurring of the image. The above usage scenarios will all result in poor quality of the acquired image. In addition, for the video surveillance field, due to the need for long-term storage and transmission of videos, the huge amount of data brought by high-definition videos, as well as the high traffic costs, video surveillance often uses low-definition video acquisition devices, compresses the acquired videos to a certain extent, and in order to obtain a larger visible range at the same video surveillance point, surveillance videos often use a wide-angle mode to shoot the surveillance scene, which will cause a large distortion of the image, and the resolution of the image gradually decreases starting from the center of the image, resulting in the fact that it is often difficult to obtain effective information about the monitored target.

[0004] The problem of image super-resolution reconstruction is an ill-posed problem. Especially for high-degree image super-resolution reconstruction, almost all the texture detail information of the super-resolved image does not exist. The optimization objectives of supervised super-resolution reconstruction algorithms mostly lie in minimizing the mean square error between the reconstructed result and the corresponding high-resolution image. While minimizing the mean square error, the peak signal-to-noise ratio is maximized, and the peak signal-to-noise ratio is also used to evaluate and compare super-resolution algorithms. However, the mean square error and the peak signal-to-noise ratio cannot effectively obtain the texture information of the image because these methods are calculated based on the differences between corresponding pixels of the image.

[0005] The mapping between the low-resolution image space and the high-resolution image space in the super-resolution reconstruction problem is a one-to-many mapping relationship. That is, there may be multiple corresponding high-resolution images for the same low-resolution image, but it is non-trivial to determine which corresponding high-resolution image is correct. Many super-resolution reconstruction techniques have an important assumption: a lot of high-frequency information is redundant, so it can be correctly reconstructed from the corresponding low-frequency information. The super-resolution reconstruction problem can, therefore, be regarded as an inference problem, and the reconstruction result depends on the statistical result of the low-resolution image information by the model of the super-resolution reconstruction network.

[0006] The problems to be solved in the super-resolution reconstruction of remote sensing images in the prior art and the key technical difficulties of this application include:

[0007] (1) There are still many problems in the super-resolution reconstruction algorithm based on deep learning in the prior art. First of all, various computer vision tasks in the deep convolutional neural network usually require a very large receptive field. For the super-resolution reconstruction of images, the larger the receptive field of the network, the more neighborhood information can participate in the super-resolution reconstruction of the image. Among many methods to expand the receptive field of the neural network, the most commonly used method is to increase the depth of the convolutional neural network. However, for increasing the convolutional layer, more network parameters will be added, and the training difficulty of the convolutional neural network will also increase accordingly. Secondly, the mean square error between pixels is the loss model optimized by many current super-resolution reconstruction networks. Although a very good peak signal-to-noise ratio index can be obtained with this as the loss model, using only the mean square error as the optimization function will result in a relatively smooth generated result, lacking the most critical high-frequency information part for the super-resolution reconstruction of images, and the super-resolution reconstruction effect of remote sensing images is poor.

[0008] (2) There are problems in the super-resolution reconstruction task of images in the prior art: to increase the receptive field of the super-resolution reconstruction network of images, increasing the convolutional layer will lead to more network parameters being added, and the training difficulty of the convolutional neural network will also increase accordingly; using the mean square error as the optimization function will result in a relatively smooth generated result, lacking the most critical high-frequency information part for the super-resolution reconstruction of images. It is urgent to improve the super-resolution reconstruction network for these two problems. The idea of combining the approximation difference multi-level network and the deep recursive convolutional network is lacking, and the convolutional layer and network parameters are increased. The problem that using the mean square error as the loss model results in a relatively smooth generated result is poor in objective evaluation indicators and visual effects. The characteristics of the difference between the generated data and the real data obtained by combining the split generation network are lacking, and there is a lack of a super-resolution reconstruction algorithm based on supervised split generation network, and the reconstruction effect at high multiples is poor.

[0009] (3) The problem of super-resolution reconstruction of images is an ill-posed problem. Especially for high-degree super-resolution reconstruction of images, almost all the texture detail information of the super-resolved images does not exist. The optimization objectives of supervised super-resolution reconstruction algorithms mostly lie in minimizing the mean square error between the reconstructed result and the corresponding high-resolution image. However, the mean square error and peak signal-to-noise ratio cannot effectively obtain the texture information of the image because these methods are calculated based on the differences between corresponding pixels of the image. When the prior art increases the receptive field of the super-resolution reconstruction network, adding more network parameters and recurrent neural networks leads to problems such as gradient explosion or gradient disappearance. Using only the mean square error as the optimization function will result in a relatively smooth generated result, lacking the most crucial high-frequency information part for image reconstruction, being unable to extract the differences at the feature level between the reconstructed result and the real image, with low remote sensing image reconstruction efficiency and serious distortion. Summary of the Invention

[0010] In view of the fact that in the prior art, to increase the receptive field of the super-resolution reconstruction network of an image, adding convolutional layers will lead to an increase in more network parameters, and the training difficulty of the convolutional neural network will also increase accordingly; using the mean square error as the optimization function will result in a relatively smooth generated result, lacking the most crucial high-frequency information part for the super-resolution reconstruction of the image, this application improves the super-resolution reconstruction network based on the above problems. By combining the ideas of approximation difference multi-level networking and deep recursive convolutional networking, the proposed multi-level approximation difference image super-resolution reconstruction network avoids adding convolutional layers and network parameters while increasing the receptive field of the super-resolution reconstruction network of the image. Aiming at the problem that using the mean square error as the loss model results in a relatively smooth generated result, by using the output result of the middle layer of the pre-trained VGG network as the loss model, it is used to extract the differences at the feature level between the reconstructed result and the real image, so that the generated result has a better performance in the high-frequency part. Combining split generative networking, which can learn the features of the differences between the generated data and the real data, a super-resolution reconstruction algorithm based on supervised split generative networking is proposed, improving the unsupervised split generative networking algorithm, making the generated result of the super-resolution reconstruction generative network G approach the actual high-resolution image. The super-resolution reconstruction algorithm based on supervised split generative networking has good reconstruction effects at high multiples, high remote sensing image reconstruction efficiency, and less distortion.

[0011] To achieve the above technical effects, the technical solutions adopted in this application are as follows:

[0012] Remote sensing image super-resolution high-fidelity high-definition image reconstruction method, which combines the approximation difference multi-level networking and the deep recursive convolution networking, constructs an image super-resolution reconstruction network, and on the pre-trained VGG network, uses the output results of its intermediate layers to extract the differences at the feature level between the super-resolution reconstruction results and the real high-resolution images, and adopts a split generation networking for image super-resolution reconstruction to solve the problem that the generation results lack high-frequency information and detail information in the image due to the mean square error-based loss model, changes the semi-supervised split learning to a supervised learning method, so that the super-resolution reconstruction network gradually learns the feature differences between the corresponding low-resolution images and high-resolution images during split learning;

[0013] 1) Image super-resolution reconstruction based on multi-level approximation difference networking: Based on the approximation difference multi-level networking and the deep recursive convolution networking, a multi-level approximation difference image super-resolution reconstruction network is established, and an approximation difference module, a recursive module, a transposed convolution layer and a networking output layer are constructed in the super-resolution reconstruction network. On the pre-trained VGG network, the output results of the reconstruction results and the real high-resolution images at the intermediate layer of the VGG network are used as the loss model to extract the differences at the feature level between the reconstruction results and the real images;

[0014] 2) Image super-resolution reconstruction based on supervised split networking: Combine the split generation networking to obtain the differences between the generated data and the real data during the training process, and establish a corresponding discriminative networking according to the correlation features between the two, and improve the unsupervised split generation networking algorithm to make the generation results of the super-resolution reconstruction generation networking G approach the actual high-resolution images.

[0015] Preferably, the approximation difference structure: Based on the intermediate layer strategy of the repeated networking of VGG and ResNets, and at the same time considering the split-deform-merge strategy in the Inception model, the two are combined;

[0016] For the input color RGB image, the 3-channel image needs to first pass through the first convolutional layer to convert the 3-channel color image into a 256-channel tensor. In each approximation difference networking module, the number of channels of the input tensor remains 256 channels, and each sub-structure performs convolutional processing on the input tensor, and the form of each sub-structure is:

[0017] (1) Convolutional layer 1, a convolutional layer with a kernel size of 1×1 and a convolutional stride of 1 is used to convolve and downsample the 256-channel input tensor to a 4-channel tensor, and batch normalization processing is performed on the convolved data through the batch normalization layer, and then the convolved tensor is output through the ReLU activation function;

[0018] (2) Convolution layer 2: The tensor after downsampling to 4 channels is processed by a convolution layer with a convolution kernel size of 3×3 and a convolution stride of 1 to extract features and map them to a 4-channel tensor. The data after convolution is batch-normalized through a batch normalization layer, and then the output of the convolution tensor is output through a ReLU activation function;

[0019] (3) Convolution layer 3: The tensor after feature extraction and mapping to 4 channels is upsampled to a 256-channel tensor by a convolution layer with a convolution kernel size of 1×1 and a convolution stride of 1. The data after convolution is batch-normalized through a batch normalization layer, and then the output of the convolution tensor is output through a ReLU activation function;

[0020] For each 256-channel tensor input to the approximation difference module, it is processed separately by 3 of the above substructures. The 256-channel tensors output by the 3 substructures are added together to obtain the output result of the current structure, and then added to the input 256-channel tensor through tensor addition to implement the approximation difference structure and obtain the output of the current approximation difference module.

[0021] Preferably, recursive structure: The super-resolution reconstruction network expands the receptive field that can process the input image. A deep recursive convolution network is used, and the same convolution layer is repeatedly applied as needed. In the middle part of the super-resolution reconstruction network based on multi-level approximation difference, the approximation difference module is looped 15 times in the way of a recursive neural network, and the receptive field of the network reaches 31×31.

[0022] Preferably, transposed convolution layer: A transposed convolution layer is added to the backend of the super-resolution reconstruction network to save the computing resources consumed by the network while enlarging the image size. For a network with a magnification factor of 2, the stride parameter of the transposed convolution layer is set to 2; for a network with a magnification factor of 4, the stride parameter of the transposed convolution layer is set to 4.

[0023] Preferably, network output layer and loss model: For 4-channel remote sensing images, the data of 4 channels are sequentially selected and the data of one of the channels is put into the reconstruction network, and the output results are merged. The input and output channels of the network are both 1 channel;

[0024] For 3-channel visible light remote sensing images, the 3-channel image data are placed side by side in the image for processing, and the output channel parameter is adjusted to 3 channels in the output layer of the network;

[0025] For each input remote sensing image data, the data is normalized through preprocessing, and the activation function of the last layer of the entire network is set to sigmoid to ensure that the output value is within a reasonable range;

[0026] Set up a loss model to calculate the mean square error between pixels. Let m and n represent the super-resolution reconstruction result and the corresponding high-resolution image respectively. Using the output results of the intermediate layers on the pre-trained VGG network, extract the differences at the feature level between the reconstruction result and the real image. The output results of the intermediate layers of the VGG network are regarded as the features extracted from the image. Put the result reconstructed by the super-resolution reconstruction network of this application and the real high-resolution image into the VGG network to obtain their eigenvalue. For these two eigenvalues, calculate the mean square error as the loss model of the super-resolution reconstruction network, calculate the differences at the feature level of the image, and optimize the super-resolution reconstruction network; fuse the mean square error between pixels and the feature error of the output of the intermediate layer of the VGG network as the loss model for super-resolution reconstruction.

[0027] Preferably, for the supervised split generation network: After each gradient update of the split generation network, clamp the weight w to a small window, thereby generating a compact parameter space W, and obtain its lower and upper limits to maintain continuity;

[0028] Modify the behavior of limiting the weight value within a small window. By randomly selecting a ratio from a uniform distribution between 0 and 1, add the generated result and the real sample in this ratio. Obtain the result of the discriminative network through the discriminative network, and calculate the gradient of this mixed result with respect to the output result of the discriminative network. Use the gap between this gradient and a set certain value as a part of the loss model and put it into the loss model of the discriminative network.

[0029] Preferably, for the super-resolution reconstruction split network G: In the supervised split generation network of this application, the super-resolution reconstruction generation network adopts an approximation difference neural network structure to ensure that during the split learning process, the super-resolution reconstruction network can effectively perform gradient backpropagation and optimize the parameters of the super-resolution reconstruction generation network;

[0030] Set 3 sequentially connected approximation difference modules in the super-resolution reconstruction split network G to ensure that there is enough receptive field during the feature extraction process of the convolutional layer of the super-resolution reconstruction network.

[0031] Preferably, for the network discriminator D: Discriminate the differences between the generated result of the super-resolution reconstruction network and the real high-resolution image, and establish a corresponding discriminative network according to the correlation features between the two;

[0032] In the discriminative network D, use LeakyReLU as the activation function for the output of each convolutional layer, where the parameter is set to 0.2 to fix the possible training instability problem during the training process of the ReLU activation function. The LeakyReLU activation function sets a very small gradient for the part less than 0;

[0033] Pooling layers are avoided throughout the discrimination network D to prevent information loss caused by pooling layers in the discrimination network D. Instead, a convolutional layer with a stride of 2 is used to downsample the tensors in the network nodes. The discrimination network consists of a total of 8 convolutional networks. Each convolutional network has a convolutional kernel size of 3×3 and 256 convolutional channels. After each convolutional layer, the size of the tensor is reduced by a factor of 4;

[0034] The size of the input image patch is set to 256×256. After passing through 8 convolutional layers, the size of the tensor output by the network nodes is 4×4×256. At the end of the discrimination network, this part of the tensor is converted into a one-dimensional vector and fed into the final fully connected layer. The output channel of this fully connected layer is 1, which is used to represent the corresponding distance between the discrimination results.

[0035] Preferably, the training process of the image super-resolution reconstruction of the supervised splitting network: The supervised splitting generation network uses the Rmsprop optimization algorithm, and the initial learning rate is set to 0.0004. During the training process of the supervised splitting generation network, the learning rate is set to decay exponentially by a factor of 0.9 every 1000 times to alleviate the impact of too large a learning rate on network training in the later stage of network training. m represents the number of images selected in each training:

[0036] Step 1: Initialize the parameters of the super-resolution reconstruction network and the discrimination network;

[0037] Step 2: Randomly select m low-resolution LR images and the corresponding m high-resolution HR images;

[0038] Step 3: Put the low-resolution images into the super-resolution generation network G to obtain the super-resolution reconstruction result SR;

[0039] Step 4: Put the super-resolution reconstruction result SR and the corresponding high-resolution image into the discrimination network respectively to obtain the corresponding discrimination network output results, D SR and D HR ;

[0040] Step 5: Calculate the mean square error between D SR and D HR :

[0041]

[0042] Step 6: Randomly select a ratio α from a uniform distribution between 0 and 1, randomly select m low-resolution LR images and the corresponding m high-resolution HR images, and weight the two with the ratio α:

[0043] AR = α * LR + (1 - α) * HR Equation 2

[0044] Step 7: Put the weighted image into the discrimination network to calculate the gradient value of the discrimination network with respect to the hybrid image;

[0045] Step 8: Use the calculation results of Step 5 and Step 7 as the loss model to optimize the weights of the discrimination network;

[0046] Step 9: Randomly select m low-score LR images and the corresponding m high-score HR images;

[0047] Step 10: Put the low-score image into the super-resolution generation network G to obtain the super-resolution reconstruction result SR;

[0048] Step 11: Put the super-resolution reconstruction result SR and the corresponding high-score image into the discrimination network respectively to obtain the corresponding discrimination network output results, D SR and D HR ;

[0049] Step 12: Calculate the mean square error between D SR and D HR :

[0050]

[0051] Step 13: Use the calculation result of Step 12 as the loss model to optimize the super-resolution reconstruction network G;

[0052] Step 14: Loop from Step 2 to Step 13 until the super-resolution reconstruction network G converges.

[0053] Compared with the prior art, the innovation points and advantages of this application are:

[0054] (1) In view of the problems in the prior art that to increase the receptive field of the super-resolution reconstruction network of an image, adding convolutional layers will lead to an increase in more network parameters and a corresponding increase in the training difficulty of the convolutional neural network; using the mean square error as the optimization function will result in a relatively smooth generated result, lacking the most crucial high-frequency information part for the super-resolution reconstruction of an image, this application improves the super-resolution reconstruction network based on the above problems. By combining the ideas of multi-level approximation difference networking and deep recursive convolutional networking, a multi-level approximation difference image super-resolution reconstruction network is proposed, which can increase the receptive field of the super-resolution reconstruction network of an image while avoiding adding convolutional layers and network parameters. Regarding the problem that using the mean square error as the loss model leads to a relatively smooth generated result, by using the output result of the intermediate layer of the pre-trained VGG network as the loss model, the difference in the feature level between the reconstructed result and the real image is extracted, so that the generated result has a better performance in the high-frequency part. Combining the split generation network that can learn the features of the difference between the generated data and the real data, a super-resolution reconstruction algorithm based on supervised split generation networking is proposed, which improves the unsupervised split generation networking algorithm, making the generated result of the super-resolution reconstruction generation network G approach the actual high-resolution image. The super-resolution reconstruction algorithm based on supervised split generation networking has a good reconstruction effect at high magnification, high efficiency in remote sensing image reconstruction and less distortion.

[0055] (2) To solve the problem of excessive parameters caused by multi-level neural networking, this application combines the ideas of multi-level approximation difference networking and deep recursive convolutional networking to construct a super-resolution reconstruction network for images. Regarding the problem that the mean square error between pixels will result in a relatively smooth generated result and lack of high-frequency information part, this application uses the output result of the intermediate layer of the pre-trained VGG network to extract the difference in the feature level between the reconstructed result and the real image as the loss model of the super-resolution reconstruction network, and optimizes the super-resolution reconstruction network. It has a poor reconstruction effect at high magnification and good super-resolution reconstruction quality of remote sensing images.

[0056] (3) The split-generation network can learn the differences between the generated data and the real data during the training process of the split network and has good generation ability. For the super-resolution reconstruction task of images, the split-generation network can optimize the generation network by identifying the feature differences between the generated result and the original high-resolution image, making the generated result gradually approach the high-resolution image and realizing the process of super-resolution reconstruction. However, for the super-resolution reconstruction network, it is not very difficult to obtain appropriate low-resolution image-high-resolution image pairs based on the semi-supervised split learning method. Compared with the traditional split-generation learning, unsupervised learning is not very suitable for super-resolution reconstruction. Even under the condition of unsupervised learning, it is very easy to have a situation where the reconstructed result is very different from the actual high-definition image. Based on such problems, while utilizing the ability of the split network to extract image features, this application changes the split learning to a supervised learning method, enabling the super-resolution reconstruction network to gradually learn the features of the corresponding high-resolution image during split learning and converge faster, which is suitable for application to remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a structural diagram of a super-resolution reconstruction network based on multi-level approximation difference.

[0058] Figure 2 It is an example diagram of 4 bands of a remote sensing image of the Beijing area taken by IKONOS.

[0059] Figure 3 It is a schematic diagram of the statistical results of various indicators of different super-resolution algorithms.

[0060] Figure 4 It is a schematic diagram of the super-resolution reconstruction split network G of the supervised split-generation network.

[0061] Figure 5 It is a schematic diagram of the network discriminator D of the supervised split-generation network.

[0062] Figure 6 It is a flowchart of the alternating training of the super-resolution generation network and the discriminator network in split learning.

[0063] Figure 7 It is a schematic diagram of the statistical results of image indicators of the UC Merced land use dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following further describes the technical solution of the method for reconstructing a high-fidelity high-definition map of remote sensing image super-resolution provided by this application with reference to the accompanying drawings, so that those skilled in the art can better understand this application and be able to implement it.

[0065] There are still many problems in the super-resolution reconstruction algorithm based on deep learning. First, various computer vision tasks in deep convolutional neural networks usually require a very large receptive field. For the super-resolution reconstruction of images, the larger the receptive field of the network, the more neighborhood information can be involved in the super-resolution reconstruction of images. Among many methods for expanding the receptive field of neural networks, the most commonly used method is to increase the depth of the convolutional neural network. However, for adding convolutional layers, more network parameters will be added, and the training difficulty of the convolutional neural network will also increase accordingly. Second, the mean square error between pixels is the loss model optimized by many current super-resolution reconstruction networks. Although a very good peak signal-to-noise ratio index can be obtained with this as the loss model, using only the mean square error as the optimization function will result in a relatively smooth generated result, lacking the high-frequency information part that is most critical for the super-resolution reconstruction of images.

[0066] To address the above two problems, this application studies the image super-resolution reconstruction algorithm based on deep learning, combines the approximation difference multi-level network and the deep recursive convolutional network, constructs an image super-resolution reconstruction network, and uses the output results of its intermediate layer on the pre-trained VGG network to extract the differences at the feature level between the super-resolution reconstruction result and the real high-resolution image. In addition, a split generative network is used for the super-resolution reconstruction of images to solve the problem of the lack of high-frequency information and detail information in the generated result due to the mean square error-based loss model, and the semi-supervised split learning is changed to a supervised learning method, enabling the super-resolution reconstruction network to gradually learn the feature differences between the corresponding low-resolution image and high-resolution image during split learning.

[0067] I. Image Super-Resolution Reconstruction Based on Multi-Level Approximation Difference Network

[0068] Super-resolution reconstruction is an ill-posed inverse problem. Collecting more domain pixels and analyzing them can always provide more clues to complement the missing information for the downsampled low-resolution image.

[0069] To increase the receptive field of the super-resolution reconstruction network while avoiding adding more network parameters and the problems of gradient explosion or gradient disappearance caused by recursive neural networks, this application combines the approximation difference multi-level network and the deep recursive convolutional network to construct an image super-resolution reconstruction network.

[0070] The mean square error between pixels is the loss model optimized by many current super-resolution reconstruction networks. Using only the mean square error as the optimization function will result in a relatively smooth generated result, lacking the most crucial high-frequency information part for image reconstruction. Therefore, in this application, on the pre-trained VGG network, the mean square error between the reconstruction result and the output result of the real high-resolution image in the middle layer of the VGG network is used as the loss model to extract the feature-level difference between the reconstruction result and the real image.

[0071] (1) Super-resolution reconstruction network structure based on multi-level approximation difference

[0072] Based on the multi-level approximation difference network and the deep recursive convolution network, a multi-level approximation difference super-resolution image reconstruction network is constructed. The framework is as Figure 1 shown. The super-resolution reconstruction network includes an approximation difference module, a recursive module, a transposed convolution layer, and a network output layer.

[0073] 1. Approximation difference structure

[0074] Based on the intermediate layer strategy of repeated networking of VGG and ResNets, and considering the split-deform-merge strategy in the Inception model at the same time, the two are combined;

[0075] For the input color RGB image, as Figure 1 shown, the 3-channel image needs to first pass through the first convolutional layer to convert the 3-channel color image into a 256-channel tensor. In each approximation difference networking module, the number of channels of the input tensor remains 256 channels, and each sub-structure performs convolutional processing on the input tensor. The form of each sub-structure is:

[0076] (1) Convolutional layer 1: A convolutional layer with a kernel size of 1×1 and a convolutional stride of 1 convolves and downsamples the 256-channel input tensor to a 4-channel tensor. The data after convolution is batch-normalized through the batch normalization layer, and then the output of the convolutional tensor is output through the ReLU activation function;

[0077] (2) Convolutional layer 2: A convolutional layer with a kernel size of 3×3 and a convolutional stride of 1 extracts features from the tensor downsampled to 4 channels and maps it to a 4-channel tensor. The data after convolution is batch-normalized through the batch normalization layer, and then the output of the convolutional tensor is output through the ReLU activation function;

[0078] (3) Convolutional layer 3: A convolutional layer with a kernel size of 1×1 and a convolutional stride of 1 maps the feature-extracted tensor to a 4-channel tensor, and then upsamples it to a 256-channel tensor through the convolutional layer. The data after convolution is batch-normalized through the batch normalization layer, and then the output of the convolutional tensor is output through the ReLU activation function;

[0079] For the 256-channel tensor input to each approximation error module, the input tensor is processed separately by 3 of the above sub-structures. The 256-channel tensors output by the 3 sub-structures are added tensorially to obtain the output result of the current structure, and then added tensorially to the input 256-channel tensor to implement the approximation error structure and obtain the output of the current approximation error module.

[0080] 2. Recursive structure

[0081] To expand the receptive field that the super-resolution reconstruction network can process for the input image, a deep recursive convolution network is used, and the same convolutional layer is repeatedly applied as needed; in the middle part of the constructed super-resolution reconstruction network based on multi-level approximation error, the approximation error module is cycled 15 times in the way of a recursive neural network, and the receptive field of the network reaches 31×31.

[0082] 3. Transposed convolution layer

[0083] A transposed convolution layer is added to the backend of the super-resolution reconstruction network to save the computing resources consumed by the network while magnifying the image size; for the network with a magnification factor of 2, the stride parameter of the transposed convolution layer is set to 2; for the network with a magnification factor of 4, the stride parameter of the transposed convolution layer is set to 4.

[0084] 4. Network output layer and loss model

[0085] For the 4-channel remote sensing image, the data of the 4 channels are sequentially selected and the data of one of the channels is put into the reconstruction network, and the output results are merged. The input and output channels of the network are both 1 channel;

[0086] For the 3-channel visible light remote sensing image, the 3-channel image data are placed side by side in the image for processing, and the output channel parameter is adjusted to 3 channels in the output layer of the network;

[0087] For each input remote sensing image data, the data is normalized through preprocessing, and the activation function of the last layer of the entire network is set to sigmoid to ensure that the output value is within a reasonable range;

[0088] A loss model is set to calculate the mean square error between pixels. m and n respectively represent the super-resolution reconstruction result and the corresponding high-resolution image. Although a very good PSNR index can be obtained with this as the loss model, using only the mean square error as the loss model will result in a relatively smooth generated result, lacking the most critical high-frequency information part for image reconstruction.

[0089] Therefore, on the pre-trained VGG network, the present application uses the output results of the intermediate layer to extract the differences at the feature level between the reconstructed result and the real image. The output results of the intermediate layer of the VGG network are regarded as the features extracted from the image. The result reconstructed by the super-resolution reconstruction network of the present application and the real high-resolution image are put into the VGG network to obtain their feature values. For these two feature values, the mean square error is calculated as the loss model of the super-resolution reconstruction network, and the differences at the feature level of the image are calculated to optimize the super-resolution reconstruction network; the mean square error between pixels and the feature error of the output of the intermediate layer of the VGG network are fused as the loss model for super-resolution reconstruction.

[0090] (2) Experiments and evaluations on the multi-level approximation difference super-resolution reconstruction network structure

[0091] 1. Experimental data

[0092] Based on the multi-level approximation difference super-resolution reconstruction network of the present application, experiments on super-resolution reconstruction are respectively carried out for the 4-channel IKONOS remote sensing satellite image dataset and the 3-channel remote sensing image dataset of the UCMerced land use dataset.

[0093] Among them, the 4-channel IKONOS remote sensing satellite image dataset consists of 10 remote sensing images of Beijing area taken by a 4-channel (blue band, green band, infrared band, red band) IKONOS satellite with a size of 1135×2066, and the spatial resolution of the images is 1 meter. For the remote sensing images taken by IKONOS, during the experimental process, after Gaussian blurring, the images are downsampled to low-resolution remote sensing images with a resolution of 2 meters by using the nearest neighbor interpolation as the input data. Examples of the 4 bands of the remote sensing images of Beijing area taken by IKONOS are as Figure 2 shown.

[0094] The remote sensing image dataset of the 3-channel UCMerced land use dataset is the images of associated ground objects selected from the large remote sensing image datasets obtained by the US Geological Survey in major urban areas. Among them, the size of each image is 256×256, there are 21 types of ground objects in total, and each type of ground object includes 100 pictures, and its spatial resolution is 1 foot. For the images in this dataset, during the experimental process of this section, the images are first processed by Gaussian blurring, and then downsampled by using the nearest neighbor interpolation method to obtain low-resolution remote sensing images as the input data.

[0095] 2. Experimental results based on IKONOS satellite remote sensing images

[0096] As Figure 3As shown in the figure, for the 4-channel IKONOS satellite remote sensing image, four indicators, namely the mean square error, peak signal-to-noise ratio, ERGAS applicable to the 4-band remote sensing image, and UIQI, are used to evaluate the reconstruction results. Among them, MSE and ERGAS are negatively correlated with the similarity between images, while PSNR and UIQI are positively correlated with the similarity between images.

[0097] On a dataset of 10 remote sensing images of the Beijing area taken by a 4-channel (blue band, green band, infrared band, red band) IKONOS satellite with a size of 1135×2066, 4 of the remote sensing images were randomly selected as the training dataset during the algorithm implementation process, and the remaining 6 remote sensing images were used as the test dataset to calculate the average values of various indicators between the reconstruction results of the super-resolution reconstruction network based on multi-level approximation difference proposed in this application and the corresponding real high-resolution images. It can be seen that the mean square error of the algorithm in this application is significantly better than the cubic interpolation and SRCNN methods. The ERGAS index is reduced by 0.5 and 0.2 compared to the Cubic method and the SRCNN method respectively. The PSNR index is increased by 4dB compared to Cubic and by 1.3dB compared to the SRCNN algorithm. The UIQI index is also better than the Cubic and SRCNN methods.

[0098] To facilitate the comparison of the reconstruction algorithm proposed in this application and other methods from the visual effect, two regions with obvious ground object features were selected from the test dataset in the 4-channel IKONOS remote sensing image dataset to show the comparison of the super-resolution reconstruction results proposed in this application, the reconstruction results of the Cubic interpolation method, and the SRCNN network. It can be seen from the visual effect that the reconstruction results of this application are significantly improved compared to the Cubic and SRCNN methods.

[0099] Based on the approximation difference multi-level network and the deep recursive convolution network, this application establishes a multi-level approximation difference image super-resolution reconstruction network, and constructs an approximation difference module, a recursive module, a transposed convolution layer, and a network output layer in the super-resolution reconstruction network. Aiming at the problem that using the mean square error as the loss model will lead to a relatively smooth generated result, the output result of the middle layer of the pre-trained VGG network is used as the loss model to extract the difference at the feature level between the reconstruction result and the real image. Finally, experiments on super-resolution reconstruction are carried out for the IKONOS satellite remote sensing image and the remote sensing image of the UCMerced land use dataset respectively. The experimental results show that the multi-level approximation difference super-resolution reconstruction network proposed in this application has a certain degree of improvement in both the objective evaluation index and the subjective visual effect of the reconstruction. Using the output result of the middle layer of the pre-trained VGG network as the loss model is superior to the super-resolution reconstruction result using the mean square error as the loss model in terms of visual effect.

[0100] II. Image Super-Resolution Reconstruction Based on Supervised Split Networking

[0101] In the work of super-resolution reconstruction based on deep learning, the mean square error is mostly used as the loss model. However, this will result in a relatively smooth generated result, lacking the most critical high-frequency information part for image reconstruction. Through experiments comparing the loss model based on the mean square error and the loss model based on the feature differences extracted by VGG networking, this application also demonstrates the disadvantages of the mean square error as a loss model.

[0102] Using split generative networking for image super-resolution reconstruction can solve the problem of the lack of high-frequency information and detail information in the generated result caused by the mean square error as a loss model.

[0103] However, for semi-supervised split learning, it is very unstable to train, often resulting in meaningless outputs from the generative networking. For the super-resolution reconstruction network, it is not very difficult to obtain suitable low-resolution image-high-resolution image pairs. Compared with split generative learning, unsupervised learning is not very suitable for super-resolution reconstruction. Even under the condition of unsupervised learning, it is very easy to have a situation where the reconstructed result is very different from the actual high-definition image. Based on such problems, while this application extracts the feature differences between the low-resolution image and the high-resolution image using split networking, it changes split learning to a supervised learning method, enabling the super-resolution reconstruction network to gradually learn the feature differences between the corresponding low-resolution image and high-resolution image during split learning and converge faster.

[0104] (I) Supervised Split Generative Networking

[0105] After each gradient update, the split generative networking clamps the weight w to a small window, thereby generating a compact parameter space W, obtaining its lower and upper limits to maintain continuity.

[0106] However, restricting the weight value within a fixed window range not only is not conducive to the learning of the networking but also limits the expressive ability of the networking. Therefore, this application modifies the behavior of restricting the weight value within a small window. By randomly selecting a ratio from a uniform distribution between 0 and 1, the generated result and the real sample are added with this ratio. The result of the discriminative networking is obtained through the discriminative networking, and the gradient between this mixed result and the output result of the discriminative networking is calculated. The gap between this gradient and a set certain value is used as a part of the loss model and put into the loss model of the discriminative networking.

[0107] 1. Super-Resolution Reconstruction Split Networking G

[0108] Due to the instability of the training of the split-generation networking itself, the same series of convolutional layers used repeatedly make the coupling between convolutional layers too tight, increasing the difficulty of training. Therefore, in the supervised split-generation networking of this application, the super-resolution reconstruction generation networking does not use a structure based on a recurrent neural network, but adopts an approximation-difference neural networking structure to ensure that during the split learning process, the super-resolution reconstruction network can effectively perform gradient backpropagation and optimize the parameters of the super-resolution reconstruction generation networking.

[0109] As Figure 4 shown in the super-resolution reconstruction generation networking G, three sequentially connected approximation-difference modules are set in the super-resolution reconstruction split networking G to ensure that there is an adequate receptive field during the feature extraction process of the convolutional layers of the super-resolution reconstruction network.

[0110] 2. Networking discriminator D

[0111] Discriminate the difference between the generation result of the super-resolution reconstruction network and the real high-resolution image, and establish a corresponding discrimination networking based on the correlation features between the two, as Figure 5 shown in the discrimination networking.

[0112] In the discrimination networking D, LeakyReLU is used as the activation function for the output of each convolutional layer, where the parameter is set to 0.2 to fix the possible training instability problem of the ReLU activation function. The LeakyReLU activation function sets a very small gradient for the part less than 0, such as 0.01.

[0113] Pooling layers are avoided throughout the discrimination networking D to avoid information loss caused by pooling layers in the discrimination networking D. Instead, convolutional layers with a stride of 2 are used to downsample the tensors in the networking nodes. The discrimination networking altogether includes 8 convolutional networkings, the convolutional kernel size of each convolutional networking is 3×3, and the convolutional channels are 256. After passing through each convolutional layer, the size of the tensor is reduced by 4 times;

[0114] The size of the input image patch is set to 256×256. After passing through 8 convolutional layers, the size of the tensor output by the networking node is 4×4×256. In the last part of the discrimination networking, this part of the tensor is converted into a one-dimensional vector and put into the last fully connected layer. The output channel of this fully connected layer is 1, which is used to represent the corresponding distance between the discrimination results.

[0115] 3. Image super-resolution reconstruction training process of the supervised split networking

[0116] The supervised splitting generation network uses the Rmsprop optimization algorithm, with the initial learning rate set to 0.0004. During the training process of the supervised splitting generation network, the learning rate is set to decrease exponentially by 0.9 every 1000 times, alleviating the impact of an overly large learning rate on network training in the later stage of network training. m represents the number of images selected in each training:

[0117] Step 1: Initialize the super-resolution reconstruction network and discriminative network parameters;

[0118] Step 2: Randomly select m low-resolution LR images and, correspondingly, m high-resolution HR images;

[0119] Step 3: Input the low-resolution images into the super-resolution generation network G to obtain the super-resolution reconstruction result SR;

[0120] Step 4: Input the super-resolution reconstruction result SR and the corresponding high-resolution images into the discriminative network respectively to obtain the corresponding discriminative network output results, D SR and D HR ;

[0121] Step 5: Calculate the mean square error between D SR and D HR :

[0122]

[0123] Step 6: Randomly select a ratio α from a uniform distribution between 0 and 1, randomly select m low-resolution LR images and, correspondingly, m high-resolution HR images, and weight the two with the ratio α:

[0124] AR = α * LR + (1 - α) * HR Equation 2

[0125] Step 7: Input the weighted images into the discriminative network and calculate the gradient value of the discriminative network with respect to the mixed image;

[0126] Step 8: Use the calculation results of Step 5 and Step 7 as the loss model to optimize the weights of the discriminative network;

[0127] Step 9: Randomly select m low-resolution LR images and, correspondingly, m high-resolution HR images;

[0128] Step 10: Input the low-resolution images into the super-resolution generation network G to obtain the super-resolution reconstruction result SR;

[0129] Step 11: Input the super-resolution reconstruction result SR and the corresponding high-resolution images into the discriminative network respectively to obtain the corresponding discriminative network output results, D SR and D HR ;

[0130] Step 12: Calculate DSR and D HR The mean square error between

[0131]

[0132] Step 13: Use the calculation result of Step 12 as the loss model to optimize the super-resolution reconstruction network G;

[0133] Step 14: Loop from Step 2 to Step 13 until the super-resolution reconstruction network G converges;

[0134] According to the above split learning algorithm process, the flow chart of the alternating training of the super-resolution generation network and the discrimination network in split learning is as Figure 6 shown.

[0135] (2) Remote sensing image experiment based on supervised split generation network

[0136] As Figure 7 shown, in the experiment, the MSE index, PSNR index, ERGAS index, UIQI index and SSIM index were used to evaluate the reconstruction results. Since the loss model of the split generation network is not based on the mean square error, it can be seen that the PSNR index is slightly lower than the result of SRCNN. However, for the indexes evaluating the image quality based on the correlation information such as the texture and contrast of the image, the results of the SSIM index, ERAGS index and UIQI index are similar to the result of SRCNN. In addition to these objective evaluation indexes, the visual effect of the details of the reconstruction result of the supervised split generation network proposed in this application is better than the reconstruction result of SRCNN. This also proves that the traditional indexes based on the difference between pixels such as PSNR are not very suitable as the quality evaluation indexes for super-resolution reconstruction.

[0137] In terms of visual effect, the result reconstructed by SRCNN is darker in color and has a certain degree of distortion. While the reconstruction result of this application is more consistent with the real high-resolution image in color. In the detail parts of the image, such as the contour line area, the visual effect of the reconstruction result of the method in this application is better than the reconstruction results of the Bicubic interpolation method and the SRCNN algorithm.

[0138] This application establishes a super-resolution reconstruction algorithm based on a supervised split generation network, combines the differences between the generated data and the real data obtained during the training process of the split generation network, improves the unsupervised split generation network algorithm, and makes the generation result of the super-resolution reconstruction generation network G approach the actual high-resolution image. Finally, an experiment on the remote sensing image of the UCMerced land use dataset was carried out with the supervised split generation network. The results show that the reconstruction result of the supervised split generation network in this application has a great improvement in visual effect.

Claims

1. A method for reconstructing a super-resolution, highly realistic, high-definition image from remote sensing images, characterized in that, Fusing the multi-level network with approximation difference and the deep recursive convolutional network, constructing a super-resolution reconstruction network for images, and using the output results of the intermediate layers of a pre-trained VGG network to extract the differences at the feature level between the super-resolution reconstruction results and the real high-resolution images. Adopting a split generative network for super-resolution reconstruction of images to solve the problem that the generation results lack high-frequency information and detail information in images due to the mean square error as the loss model. Changing the semi-supervised split learning to a supervised learning method, enabling the super-resolution reconstruction network to gradually learn the feature differences between the corresponding low-resolution images and high-resolution images during split learning; 1) Image super-resolution reconstruction based on the multi-level network with approximation difference: Based on the multi-level network with approximation difference and the deep recursive convolutional network, a multi-level approximation difference image super-resolution reconstruction network is established. An approximation difference module, a recursive module, a transposed convolution layer, and a network output layer are constructed in the super-resolution reconstruction network. Using the output results of the intermediate layers of the VGG network after pre-training as the loss model for the reconstruction results and the real high-resolution images, and extracting the differences at the feature level between the reconstruction results and the real images; 2) Image super-resolution reconstruction based on the supervised split network: Combining the split generative network to obtain the differences between the generated data and the real data during the training process, and establishing a corresponding discriminative network based on the correlation features between the two. Improving the unsupervised split generative network algorithm to make the generation results of the super-resolution reconstruction generative network G approach the actual high-resolution images.

2. The method for reconstructing a super-resolution, highly realistic, high-definition map of remote sensing images according to claim 1, wherein Approximation difference structure: Based on the intermediate layer strategy of the repeated network of VGG and ResNets, and at the same time considering the split-deform-merge strategy in the Inception model, combining the two; For the input color RGB image, the 3-channel image needs to first convert the 3-channel color image into a 256-channel tensor through the first convolutional layer. In each approximation difference network module, the number of channels of the input tensor remains 256 channels, and each sub-structure performs convolutional processing on the input tensor. The form of each sub-structure is as follows: (1) Convolutional layer 1: A convolutional layer with a kernel size of 1×1 and a stride of 1 convolves and downsamples the 256-channel input tensor to a 4-channel tensor. The data after convolution is batch-normalized through the batch normalization layer, and then the output of the convolutional tensor is passed through the ReLU activation function; (2) Convolutional layer 2: A convolutional layer with a kernel size of 3×3 and a stride of 1 extracts features from the tensor downsampled to 4 channels and maps it to a 4-channel tensor. The data after convolution is batch-normalized through the batch normalization layer, and then the output of the convolutional tensor is passed through the ReLU activation function; (3) Convolutional layer 3: A convolutional layer with a kernel size of 1×1 and a stride of 1 upsamples the tensor with the extracted features mapped to 4 channels to a tensor with a size of 256 channels through the convolutional layer. The data after convolution is batch-normalized through the batch normalization layer, and then the output of the convolutional tensor is passed through the ReLU activation function; For the 256-channel tensor input to each approximation difference module, the input tensor is processed separately by 3 of the above sub-structures. The 256-channel tensors output by the 3 sub-structures are added tensorially to obtain the output result of the current structure, and then added tensorially to the input 256-channel tensor to implement the approximation difference structure and obtain the output of the current approximation difference module.

3. The method for reconstructing a super-resolution, highly realistic, high-definition image of remotely sensed imagery according to claim 1, wherein Recursive structure: The receptive field that the super-resolution reconstruction network can process for the input image is expanded. A deep recursive convolution network is used, and the same convolutional layer is repeatedly applied as needed. In the middle part of the super-resolution reconstruction network based on multi-level approximation differences, the approximation difference module is cycled 15 times in the way of a recursive neural network, and the receptive field of the network reaches 31×31.

4. The method for reconstructing a super-resolution, highly realistic, high-definition map of remote sensing images according to claim 1, wherein, Transposed convolution layer: A transposed convolution layer is added to the backend of the super-resolution reconstruction network to save the computational resources consumed by the network while magnifying the image size. For a network with a magnification factor of 2, the stride parameter of the transposed convolution layer is set to 2; for a network with a magnification factor of 4, the stride parameter of the transposed convolution layer is set to 4.

5. The method for reconstructing a super-resolution, highly realistic, high-definition map of remotely sensed images according to claim 1, wherein Network output layer and loss model: For the 4-channel remote sensing image, the data of 4 channels are sequentially selected and the data of one of the channels is put into the reconstruction network, and the output results are merged. The input and output channels of the network are both 1 channel; For the 3-channel visible light remote sensing image, the 3-channel image data are placed side by side in the image for processing, and the output channel parameter is adjusted to 3 channels in the output layer of the network; For each input remote sensing image data, the data is normalized through preprocessing, and the activation function of the last layer of the entire network is set to sigmoid to ensure that the output value is within a reasonable value range; The loss model is set to calculate the mean square error between pixels. m and n respectively represent the super-resolution reconstruction result and the corresponding high-resolution image. On the pre-trained VGG network, the output results of the intermediate layer are used to extract the differences at the feature level between the reconstruction result and the real image. The output results of the intermediate layer of the VGG network are regarded as the features extracted for the image. The result reconstructed by the super-resolution reconstruction network of the present application and the real high-resolution image are put into the VGG network to obtain their eigenvalue. For these two eigenvalues, the mean square error is calculated as the loss model of the super-resolution reconstruction network, and the differences at the feature level of the image are calculated to optimize the super-resolution reconstruction network; The mean square error between pixels and the feature error of the output of the intermediate layer of the VGG network are fused as the loss model for super-resolution reconstruction.

6. The method for reconstructing a super-resolution, highly realistic, high-definition map of remote sensing images according to claim 1, wherein Supervised split generation network: After each gradient update, the split generation network clamps the weight w to a small window, thereby generating a compact parameter space W, and obtaining its lower and upper bounds to maintain continuity; Modify the behavior of limiting the weights within a small window. By randomly selecting a ratio from a uniform distribution between 0 and 1, add the generated result and the real sample with this ratio. Obtain the result of the discrimination network through the discrimination network, and calculate the gradient between this mixed result and the output result of the discrimination network. Use the gap between this gradient and a set fixed value as part of the loss model and put it into the loss model of the discrimination network.

7. The method for reconstructing a super-resolution, highly realistic, high-definition map from remote sensing images according to claim 1, wherein Super-resolution reconstruction split network G: In the supervised split generation network of this application, the super-resolution reconstruction generation network adopts an approximation difference neural network structure to ensure that during the split learning process, the super-resolution reconstruction network can effectively perform gradient backpropagation and optimize the parameters of the super-resolution reconstruction generation network. Set 3 sequentially connected approximation difference modules in the super-resolution reconstruction split network G to ensure that there is enough receptive field during the feature extraction process of the convolutional layer of the super-resolution reconstruction network.

8. The method for reconstructing a super-resolution, highly realistic, high-definition map of remote sensing images according to claim 1, wherein Network discriminator D: Discriminate the difference between the generated result of the super-resolution reconstruction network and the real high-resolution image, and establish a corresponding discrimination network based on the correlation features between the two. Adopt LeakyReLU as the activation function for the output of each convolutional layer in the discrimination network D, where the parameter is set to 0.2 to fix the possible training instability problem of the ReLU activation function during training. The LeakyReLU activation function sets a very small gradient for the part less than 0. Avoid using pooling layers throughout the entire discrimination network D to avoid information loss caused by pooling layers in the discrimination network D. Instead, use a convolutional layer with a stride of 2 to downsample the tensors in the network nodes. The discrimination network includes a total of 8 convolutional networks, the convolutional kernel size of each convolutional network is 3×3, and the convolutional channels are 256. After passing through each convolutional layer, the size of the tensor is reduced by 4 times. Set the size of the input image patch to 256×256. After passing through 8 convolutional layers, the size of the tensor output by the network node is 4×4×256. In the last part of the discrimination network, convert this part of the tensor into a one-dimensional vector and put it into the last fully connected layer. The output channel of this fully connected layer is 1, which is used to represent the corresponding distance between the discrimination results.

9. The method for reconstructing a super-resolution, highly realistic, high-definition map of remotely sensed images according to claim 1, wherein Image super-resolution reconstruction training process of the supervised split network: The supervised split generation network adopts the Rmsprop optimization algorithm, sets the initial learning rate to 0.0004, and during the training process of the supervised split generation network, set the learning rate to decay exponentially by 0.9 every 1000 times to alleviate the impact of too large learning rate on network training in the later stage of network training. m represents the number of images selected in each training. Step 1: Initialize the parameters of the super-resolution reconstruction network and the discrimination network. Step 2: Randomly select m low-resolution LR images and the corresponding m high-resolution HR images. Step 3: Put the low-resolution images into the super-resolution generation network G to obtain the super-resolution reconstruction result SR. Step 4: Put the super-resolution reconstruction result SR and the corresponding high-resolution image into the discrimination network respectively to obtain the corresponding discrimination network output results, D SR and D HR ; Step 5: Calculate D SR and D HR The mean square error between: Step 6: Randomly select a ratio α from a uniform distribution between 0 and 1, randomly select m low-resolution LR images and the corresponding m high-resolution HR images, and weight the two with the ratio α. AR = α * LR + (1 - α) * HR Equation 2 Step 7: Put the weighted image into the discrimination network and calculate the gradient value of the discrimination network with respect to the hybrid image; Step 8: Use the calculation results of Step 5 and Step 7 as the loss model to optimize the weights of the discrimination network; Step 9: Randomly select m low-resolution LR images and, correspondingly, m high-resolution HR images; Step 10: Put the low-resolution images into the super-resolution generation network G to obtain the super-resolution reconstruction result SR; Step 11: Put the super-resolution reconstruction result SR and the corresponding high-resolution image into the discrimination network respectively to obtain the corresponding discrimination network output results, D SR and D HR ; Step 12: Calculate D SR and D HR The mean squared error between: Step 13: Use the calculation result of Step 12 as the loss model to optimize the super-resolution reconstruction network G; Step 14: Loop from Step 2 to Step 13 until the super-resolution reconstruction network G converges.