A satellite image super-resolution reconstruction method assisted by drone images

By using drone image assistance methods, the training of the generative adversarial network for super-resolution reconstruction of satellite images is solved, the problem of insufficient resolution of satellite images is achieved, and the high-resolution reconstruction effect is achieved, saving hardware upgrade costs.

CN119151789BActive Publication Date: 2025-06-10ZHEJIANG GUOYAO GEOGRAPHIC INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411650807.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-06-10
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

The spatial resolution of satellite images is low, which is difficult to meet the needs of refined applications. The solution to improve sensor hardware performance is expensive and technically difficult to implement.

Method used

Using a super-resolution reconstruction method for satellite images based on drone image assisted, aerial images collected by drones are acquired for aerial triangulation and orthogonal correction, training data pairs of high-resolution and low-resolution images are constructed, and an adversarial network is trained to generate an adversarial network, and finally high-resolution satellite images are generated.

Benefits of technology

Super-resolution reconstruction of low-resolution satellite images is achieved without increasing or replacing hardware, saving costs, and technology implementation is relatively easy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119151789B_ABST
    Figure CN119151789B_ABST
Patent Text Reader

Abstract

An embodiment of this specification discloses a satellite image super-resolution reconstruction method assisted by drone images. The method includes obtaining aerial images collected by a drone, performing aerial triangulation based on image control points, and orthorectifying the aerial images; selecting first images in different terrain areas, downsampling the first images to obtain second images, and constructing first training data to train a generative adversarial network; obtaining a target image to be reconstructed, reading the target image in blocks as at least two sub-blocks, generating reconstruction results of each sub-block based on the trained generative adversarial network, and stitching the reconstruction results based on the edge gradient fusion method to obtain a reconstructed image. The embodiment of this specification can obtain higher-resolution images without upgrading or replacing hardware, saving costs and being easy to implement technically.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to image reconstruction technology, and in particular, to a method for super-resolution reconstruction of satellite images assisted by UAV images. Background Art

[0002] Remote sensing satellite images play an important role in fields such as land resource monitoring, urban planning, and agricultural production. However, limited by factors such as imaging equipment, atmospheric conditions, and orbital altitude, the spatial resolution of satellite images is relatively low and often difficult to meet the requirements of refined applications. Although higher-resolution images can be obtained by improving the hardware performance of sensors, this solution is costly and technically difficult to implement. Therefore, there is a need for a method to perform super-resolution reconstruction on low-resolution satellite images. Summary of the Invention

[0003] To solve the above problems, one or more embodiments of this specification describe a method for super-resolution reconstruction of satellite images assisted by UAV images.

[0004] According to a first aspect, there is provided a method for super-resolution reconstruction of satellite images assisted by UAV images, the method comprising:

[0005] Obtain aerial images collected by a UAV, perform aerial triangulation based on image control points set in the aerial images, and orthorectify the aerial images according to a digital elevation model generated by the measurement results to obtain a digital orthophoto map;

[0006] Select first images of different terrain regions in the digital orthophoto map, downsample the first images to obtain second images, construct first training data pairs based on the first images and the second images, and train a generative adversarial network based on the first training data pairs. Each of the first training data pairs includes a high-resolution image and a low-resolution image with the same region size and region coordinates;

[0007] Obtain a target image to be reconstructed, read the target image in blocks as at least two sub-blocks, generate reconstruction results of each of the sub-blocks based on the trained generative adversarial network, and splice the reconstruction results based on the edge gradient fusion method to obtain a reconstructed image corresponding to the target image.

[0008] Preferably, the obtaining aerial images collected by a UAV and performing aerial triangulation based on image control points set in the aerial images includes:

[0009] Obtain aerial images collected by the UAV based on oblique photography and display the aerial images;

[0010] Receive a setting instruction for the aerial image, determine the image control points in the aerial image, and perform aerial triangulation based on each of the image control points to obtain measurement data.

[0011] Preferably, the orthorectification of the aerial image using the digital elevation model generated according to the measurement results to obtain a digital orthophoto map includes:

[0012] Project the pixel coordinates in the aerial image into three-dimensional coordinates in three-dimensional space according to the measurement results, and determine the three-dimensional point cloud data in three-dimensional space. Construct a digital elevation model according to the three-dimensional point cloud data;

[0013] After optimizing each of the three-dimensional coordinates based on the ground elevation value of the digital elevation model, resample to obtain new pixel coordinates according to the optimized three-dimensional coordinates, and generate a preliminary orthophoto map corresponding to each of the aerial images according to the new pixel coordinates;

[0014] Stitch each of the preliminary orthophoto maps to obtain a digital orthophoto map.

[0015] Preferably, the construction of the first training data pair based on the first image and the second image includes:

[0016] After adding Gaussian noise to the second image, perform Gaussian blur processing on the second image to obtain a third image;

[0017] Cut the first image and the third image into at least two first sub-images and second sub-images respectively, and the image sizes of the first sub-images and the second sub-images are the same;

[0018] Construct the first sub-images and the second sub-images corresponding to the same image coordinates into the first training data pair, and integrate each of the first training data pairs to obtain a training data set.

[0019] Preferably, the integration of each of the first training data pairs to obtain a training data set includes:

[0020] Obtain the regional satellite images of each of the terrain regions, and adjust the regional satellite images based on geometric correction so that the geographical location of the regional satellite images is consistent with that of the first image;

[0021] Construct a second training data pair based on the first image and the regional satellite images;

[0022] Integrate each of the first training data pairs and the second training data pairs to obtain a training data set.

[0023] Preferably, the training of the adversarial network based on the first training data pair includes:

[0024] Obtain an initial generative adversarial network, where the generator of the initial generative adversarial network is a deep network with at least two dense residual learning blocks, and the discriminator of the initial generative adversarial network is a U-Net discriminator with skip connections;

[0025] Train the initial generative adversarial network based on the first training data to obtain a preliminary trained network, and optimize the preliminary trained network based on the second training data to obtain a trained generative adversarial network.

[0026] Preferably, the obtaining of the target image to be reconstructed, reading the target image in blocks into at least two sub-blocks, generating reconstruction results for each of the sub-blocks based on the trained generative adversarial network, and stitching the reconstruction results based on the edge gradient fusion method to obtain the reconstructed image corresponding to the target image, includes:

[0027] Obtain the target image to be reconstructed, read the target image in blocks into at least two sub-blocks, and there is an overlapping area of a preset number of pixels between adjacent sub-blocks;

[0028] Generate reconstruction results for each of the sub-blocks based on the trained generative adversarial network;

[0029] For the overlapping pixels corresponding to the overlapping area in the reconstruction results, use the weighted pixel sum of the edge distances of the overlapping pixels at the same position in each reconstruction result as the final pixel value at that position to obtain the reconstructed image corresponding to the target image.

[0030] According to a second aspect, there is provided a satellite image super-resolution reconstruction device assisted by UAV images, the device includes:

[0031] An acquisition module, configured to acquire aerial images collected by a UAV, perform aerial triangulation based on image control points set in the aerial images, and orthorectify the aerial images according to a digital elevation model generated by the measurement results to obtain a digital orthophoto map;

[0032] A training module, configured to select first images of different terrain regions in the digital orthophoto map, downsample the first images to obtain second images, construct a first training data pair based on the first images and the second images, and train a generative adversarial network based on the first training data pair, where each first training data pair includes a high-resolution image and a low-resolution image with the same region size and region coordinates;

[0033] A reconstruction module, configured to obtain the target image to be reconstructed, read the target image in blocks into at least two sub-blocks, generate reconstruction results for each of the sub-blocks based on the trained generative adversarial network, and stitch the reconstruction results based on the edge gradient fusion method to obtain the reconstructed image corresponding to the target image.

[0034] According to a third aspect, an electronic device is provided, including a processor and a memory;

[0035] The processor is connected to the memory;

[0036] The memory is used to store executable program code;

[0037] The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the steps of the method provided in the first aspect or any possible implementation manner of the first aspect.

[0038] According to a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer or a processor, the computer or the processor is enabled to execute the method provided in the first aspect or any possible implementation manner of the first aspect.

[0039] The method provided in the embodiments of this specification can obtain a digital orthophoto map after orthorectification processing of the aerial images collected by the unmanned aerial vehicle, and obtain a first image for characterizing a high-resolution image and a second image for characterizing a low-resolution image according to the digital orthophoto map, construct first training data to train a generative adversarial network, and finally realize super-resolution reconstruction of the target image according to the generative adversarial network. Through the above method, the downsampled low-resolution unmanned aerial vehicle images can be used as low-resolution satellite images, combined with high-resolution unmanned aerial vehicle images to construct training data to train a generative adversarial network, so that the trained generative adversarial network can perform super-resolution reconstruction on satellite images, and higher-resolution images can be obtained without upgrading or replacing hardware, saving costs and being easy to implement technically. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 is a schematic flowchart of a method for super-resolution reconstruction of satellite images assisted by unmanned aerial vehicle images in an embodiment of this specification.

[0042] Figure 2 is a schematic structural diagram of a device for super-resolution reconstruction of satellite images assisted by unmanned aerial vehicle images in an embodiment of this specification.

[0043] Figure 3 It is a schematic structural diagram of an electronic device in an embodiment of this specification. Detailed implementation manners

[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0045] In the following introduction, the terms "first" and "second" are only for the purpose of description and cannot be construed as indicating or implying relative importance. The following introduction provides multiple embodiments of the present application. Different embodiments can be replaced or combined. Therefore, the present application can also be considered to include all possible combinations of the same and / or different embodiments described. Thus, if one embodiment includes features A, B, and C, and another embodiment includes features B and D, then the present application should also be considered to include embodiments containing one or more of all other possible combinations of A, B, C, and D, even though such embodiments may not be explicitly described in the following content.

[0046] The following description provides examples and does not limit the scope, applicability, or examples set forth in the claims. Changes can be made to the functions and arrangements of the described elements without departing from the scope of the content of the present application. Each example can appropriately omit, substitute, or add various processes or components. For example, the described method can be executed in a different order from the described order, and various steps can be added, omitted, or combined. In addition, the features described in some examples can be combined into other examples.

[0047] See Figure 1 , Figure 1 It is a schematic flowchart of a satellite image super-resolution reconstruction method assisted by UAV images provided by an embodiment of the present application. In the embodiment of the present application, the method includes:

[0048] S101. Obtain aerial images collected by a UAV, perform aerial triangulation based on the image control points set in the aerial images, and orthorectify the aerial images according to the digital elevation model generated by the measurement results to obtain digital orthophoto maps.

[0049] The execution subject of the present application can be a cloud server.

[0050] In the embodiments of this specification, to construct a training set, first, the drone is controlled to perform aerial photography to obtain aerial images collected by the drone. Compared with satellites, the shooting distance of the drone from the ground is lower, so the obtained aerial images are clearer. These aerial images will be used as high-resolution images in the subsequent training set. To enable the aerial images to represent the situation of satellite images, before using the aerial images as the high-resolution images of the training set, it is also necessary to preprocess the aerial images and adjust them into orthophotos that match the satellite images. Specifically, after the drone has collected the aerial images, the staff can set image control points for the aerial images, that is, select some features on the ground that are easy to identify and not easy to change (such as road intersections, building corners, obvious terrain features, etc.) as image control points in the images and mark them in the aerial images. The aerial images obtained by the cloud server can be the images that have been marked, so the image control points set by the staff can be directly determined from the images. Through the image control points, aerial triangulation can be performed. Aerial Triangulation (AT) is a method to determine the geometric relationship between images through photogrammetry technology. It will find the same image control points in different images through feature matching algorithms and use the least squares method or other optimization methods to calculate the internal and external orientation elements of each image, that is, the measurement results, and then construct a three-dimensional space model of the corresponding positions of the images to obtain specific three-dimensional data of the corresponding positions. Then, a digital elevation model can be generated, and then the three-dimensional space model can be reversely optimized and adjusted through the digital elevation model. Finally, each pixel in the aerial image is mapped to an exact position on the ground through the three-dimensional space model to achieve orthorectification and obtain the digital orthophoto corresponding to the aerial image. Among them, after aerial triangulation, the aerial images can also be radiometrically corrected simultaneously to eliminate brightness changes caused by factors such as atmospheric scattering.

[0051] In an implementable manner, the obtaining of the aerial images collected by the drone and performing aerial triangulation based on the image control points set in the aerial images includes:

[0052] Obtain the aerial images collected by the drone based on oblique photography and display the aerial images;

[0053] Receive a setting instruction for the aerial images, determine the image control points in the aerial images, and perform aerial triangulation based on each of the image control points to obtain measurement data.

[0054] In the embodiments of this specification, the drone will first collect aerial images through oblique photography. Oblique photography means mounting multiple sensors on the same drone and collecting images from five different angles, namely one vertical angle and four oblique angles, simultaneously, in order to obtain richer high-resolution textures of the top surface and side views of buildings and improve the accuracy of subsequent processing results. After collecting the aerial images, they will be displayed on the terminal used by the staff. The staff can perform quality inspection on the images and then set image control points. The cloud server will receive the setting instructions generated by the staff's operations to determine the image control points set by the staff, and then perform aerial triangulation based on the image control points to obtain measurement data. Among them, after obtaining the measurement data, the measurement data can also be sent to the staff again, and the staff will perform quality inspection and screening on the measurement data.

[0055] In an implementable manner, orthorectifying the aerial image with the digital elevation model generated according to the measurement result to obtain a digital orthophoto map includes:

[0056] Projecting the pixel coordinates in the aerial image into three-dimensional coordinates in three-dimensional space according to the measurement result, and determining the three-dimensional point cloud data in three-dimensional space. A digital elevation model is constructed according to the three-dimensional point cloud data;

[0057] After optimizing each of the three-dimensional coordinates based on the ground elevation value of the digital elevation model, resampling according to the optimized three-dimensional coordinates to obtain new pixel coordinates, and generating a preliminary orthophoto map corresponding to each of the aerial images according to the new pixel coordinates;

[0058] Stitching each of the preliminary orthophoto maps to obtain a digital orthophoto map.

[0059] In the embodiments of this specification, the measurement results include the interior orientation elements (such as focal length, principal point coordinates) and exterior orientation elements (such as the coordinates of the image center and attitude angles) of each image calculated through the relationship between control points and overlapping images. By resolving the measurement results, the pixel coordinates in the aerial image can be projected into three-dimensional coordinates in a three-dimensional space to construct a three-dimensional space model, enabling each image to have an accurate positioning in the three-dimensional space. At the same time, after constructing the three-dimensional space model through aerial triangulation, a large amount of three-dimensional point cloud data will be generated. These point cloud data contain the height information of the ground surface. Therefore, a Digital Elevation Model (DEM) can be constructed based on this height information. The height coordinate parameters in the three-dimensional coordinates can be optimized and adjusted according to the ground elevation values in the digital elevation model to refine and adjust the three-dimensional model, making it more accurately reflect the actual terrain and landform. By resampling the optimized three-dimensional coordinates, new pixel coordinates corresponding to the aerial image can be obtained. By rewriting the pixel values of the new pixel coordinates into the image file, a preliminary orthophoto map can be obtained. The aerial images taken by the UAV are relatively small compared to satellite images. Therefore, it is necessary to splice the taken preliminary orthophoto maps, and quality inspection and local correction can also be carried out manually if necessary, and finally a digital orthophoto map is formed.

[0060] S102. Select a first image of different terrain regions in the digital orthophoto map, downsample the first image to obtain a second image, construct a first training data pair based on the first image and the second image, and train and generate an adversarial network based on the first training data pair.

[0061] Among them, each first training data pair includes a high-resolution image and a low-resolution image with the same region size and region coordinates.

[0062] In the embodiments of this specification, in order to construct a training sample set, first images of different terrain areas (such as urban areas, rural areas, mountains, water towns, etc.) of a certain size are selected from the digital orthophoto map as the high-resolution images of the UAV images. Then, the first images are downsampled to obtain second images as the low-resolution images of the UAV images. The downsampling can specifically adopt the bicubic downsampling method, and the downsampling ratio can be selected as 4 times downsampling. The low-resolution second images are similar to the satellite remote sensing images collected by satellites. Therefore, images of the same position and the same size in the first images and the second images can be matched together to construct the first training data pairs, and a generative adversarial network is trained based on the first training data pairs. In this way, the trained generative adversarial network can output the corresponding high-resolution image when receiving the low-resolution satellite remote sensing image. Among them, when the first images and the second images are large, the first images and the second images can also be cut first, and then the first training data pairs are constructed according to the cut images. In addition, since the second images have been downsampled, in order to make the sizes of the regions corresponding to the high-resolution images obtained from the first images and the low-resolution images obtained from the second images in each first training data pair the same, the sizes of the high-resolution images and the low-resolution images in each first training data pair need to be adjusted according to the downsampling ratio. For example, assuming the downsampling is 4 times, the low-resolution image can be an image of 256*256, and the high-resolution image can be an image of 64*64.

[0063] In one implementable manner, constructing the first training data pair based on the first image and the second image includes:

[0064] After adding Gaussian noise to the second image, perform Gaussian blur processing on the second image to obtain a third image;

[0065] Cut the first image and the third image into at least two first sub-images and second sub-images respectively, and the image sizes of the first sub-images and the second sub-images are the same;

[0066] Construct the first sub-images and the second sub-images corresponding to the same image coordinates into the first training data pairs, and integrate each of the first training data pairs to obtain a training data set.

[0067] In the embodiments of this specification, in order to make the quality of the low-resolution UAV images closer to that of the satellite images, the second images are also degraded. Specifically, first, Gaussian noise is added to the second image. Taking the addition of Gaussian noise with a mean of 0 and a standard deviation of 5 as an example, the calculation formula is as follows:

[0068]

[0069] Among them, is the original image, It is Gaussian noise.

[0070] Next, the image after adding noise is subjected to Gaussian blur processing using a 3*3 convolution kernel:

[0071]

[0072] where x and y are the offsets of the pixel positions, is the standard deviation of the Gaussian function.

[0073] Next, the first image and the third image are cut to obtain a first sub-image corresponding to the first image and a second sub-image corresponding to the third image, and a first training data pair is constructed based on the first sub-image and the second sub-image. The first training data pair can be understood as a high-resolution - low-resolution data pair of UAV images.

[0074] In an implementable manner, the integrating of each of the first training data pairs to obtain a training data set includes:

[0075] Obtain the regional satellite images of each of the terrain regions, and adjust the regional satellite images based on geometric correction so that the geographical location of the regional satellite images is consistent with that of the first image;

[0076] Construct a second training data pair based on the first image and the regional satellite images;

[0077] Integrate each of the first training data pairs and the second training data pairs to obtain a training data set.

[0078] In the embodiments of this specification, in order to further improve the accuracy of the trained model, in addition to the first training data pairs completely constructed based on UAV images, the actual regional satellite images of each terrain region are also obtained, and the positions of the regional satellite images are adjusted through geometric correction so that the regional satellite images can be basically consistent with the geographical location of the first image. Then, in the same way as the first training data pairs, the first image and the regional satellite images are cut, and a second training data pair representing UAV high-resolution - satellite low-resolution is constructed. Among them, since the time phases of UAV images and satellite images are not the same, it is necessary to perform manual screening on the regional satellite images to eliminate image pairs that do not match or have large differences, such as the presence of new buildings, building demolitions, the presence of dense vehicles, river boats, etc. The first training data pairs can be used for the initial training of the model, and the second training data pairs can be used for the optimization and fine-tuning of the model.

[0079] In an implementable manner, the training of the generative adversarial network based on the first training data pair includes:

[0080] Obtain an initial generative adversarial network, where the generator of the initial generative adversarial network is a deep network with at least two dense residual learning blocks, and the discriminator of the initial generative adversarial network is a U-Net discriminator with skip connections;

[0081] Train the initial generative adversarial network based on the first training data to obtain a preliminary trained network, and optimize the preliminary trained network based on the second training data to obtain a trained generative adversarial network.

[0082] In the embodiments of this specification, the initial generative adversarial network model adopted in this application can be Real-ESRGAN. Its generator is a deep network with several dense residual learning blocks (RRDB). All convolutional layers use a stride of 1 to maintain the resolution of the image in the generator. The concept of residuals can be used to prevent instability when training this deep network. The discriminator can use a U-Net discriminator with skip connections. The U-Net outputs the authenticity value of each pixel and can provide detailed feedback for each pixel to the generator. At the same time, in order to avoid the training instability brought by the U-Net network structure, frequency domain regularization can be adopted to stabilize the training dynamics, which also helps to mitigate the oversharpening situation caused during the GAN training process. During the training process, first train the network through the first training data to obtain a preliminary trained network, and then optimize the network model through the second training data to fine-tune it to obtain a trained generative adversarial network. Among them, the model uses the Adam optimizer, and the batch size is set to 12. The generator network is trained at a learning rate of 2x10e-4 for 1000k iterations, and fine-tuned at a learning rate of 1x10e-4 for 40k iterations. Exponential moving average (EMA) is used for more stable training. The L1 loss, perceptual loss, and GAN loss are used to train the model, and the weight ratios are 1, 1, and 0.1.

[0083] In addition, the peak signal-to-noise ratio (PSNR) can also be used as an evaluation index for the model. It is one of the standard indexes for evaluating the quality of reconstructed images. A higher PSNR indicates a higher image quality, which can be expressed as:

[0084]

[0085] Among them, MaxVal represents the maximum pixel value of the HR image, and MSE represents the mean square error.

[0086] S103. Obtain the target image to be reconstructed, read the target image in blocks as at least two sub-blocks, generate the reconstruction results of each sub-block based on the trained generative adversarial network, and splice the reconstruction results based on the edge gradient fusion method to obtain the reconstructed image corresponding to the target image.

[0087] In the embodiments of this specification, after the generative adversarial network is trained, the satellite image to be reconstructed, that is, the target image, will be obtained. Since the target image is large, it needs to be read in blocks first to obtain a number of smaller sub-blocks. For example, it can be read in blocks into sub-blocks of 500*500 size, and then the sub-blocks are respectively input into the trained generative adversarial network to obtain the reconstruction results of each sub-block. Since stitching lines are likely to occur at the image stitching locations, the reconstruction results will be stitched by means of edge gradient fusion to obtain the reconstructed image that finally completes the super-resolution reconstruction.

[0088] In an implementable manner, the obtaining of the target image to be reconstructed, reading the target image in blocks into at least two sub-blocks, generating the reconstruction results of each sub-block based on the trained generative adversarial network, and stitching the reconstruction results based on the edge gradient fusion method to obtain the reconstructed image corresponding to the target image includes:

[0089] Obtain the target image to be reconstructed, read the target image in blocks into at least two sub-blocks, and there is an overlapping area with a preset number of pixels between adjacent sub-blocks;

[0090] Generate the reconstruction results of each sub-block based on the trained generative adversarial network;

[0091] For the overlapping pixels corresponding to the overlapping area in the reconstruction results, the edge distance weighted pixel sum of the overlapping pixels at the same position in each reconstruction result is used as the final pixel value at that position to obtain the reconstructed image corresponding to the target image.

[0092] In the embodiments of this specification, in order to achieve edge gradient fusion, an overlapping area with a certain number of pixels is set between sub-blocks. For the pixels in the overlapping area, the final pixel value after stitching is the edge distance weighted pixel sum of the overlapping parts of each overlapping sub-block at that position. Specifically, taking the overlapping area with 150 pixels set as an example, for two sub-images A and B with one-side overlap at the edge of the large image, the pixels in their overlapping area can be expressed as:

[0093]

[0094] where ij represents the position of the pixel, 、 respectively represent the distances from the pixel at this point to the edges of sub-images A and B. The closer to the edge of the overlapping area, the less the contribution of the pixel value of the sub-image to the final pixel value at this point.

[0095] For the case where there is an overlapping area among four sub-images in the middle of the large image, the pixel value at this point is the weighted sum of the corresponding pixel points of the four sub-images according to the distance, and the pixels in its overlapping area can be expressed as:

[0096]

[0097] Among them, 、 、 、 respectively represent the distances from the pixel of this point to the edges of sub - images A, B, C, and D.

[0098] Next, in combination with the attached Figure 2 , the satellite image super - resolution reconstruction device assisted by UAV images provided by the embodiments of the present application will be introduced in detail. It should be noted that the satellite image super - resolution reconstruction device assisted by UAV images shown in the attached Figure 2 is used to execute the method of the embodiments of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the embodiments shown in the present application Figure 1 shown. Figure 1 shown.

[0099] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a satellite image super - resolution reconstruction device assisted by UAV images provided by the embodiments of the present application. As shown in Figure 2 , the device includes:

[0100] An acquisition module 201, configured to acquire aerial images collected by a UAV, perform aerial triangulation based on the image control points set in the aerial images, and orthorectify the aerial images according to the digital elevation model generated by the measurement results to obtain digital orthophoto maps;

[0101] A training module 202, configured to select first images of different terrain regions in the digital orthophoto map, downsample the first images to obtain second images, construct first training data pairs based on the first images and the second images, and train a generative adversarial network based on the first training data pairs. Each of the first training data pairs includes a high - resolution image and a low - resolution image with the same region size and region coordinates;

[0102] A reconstruction module 203, configured to acquire a target image to be reconstructed, read the target image in blocks into at least two sub - blocks, generate reconstruction results for each of the sub - blocks based on the trained generative adversarial network, and splice the reconstruction results based on the edge - gradient fusion method to obtain a reconstruction image corresponding to the target image.

[0103] In an implementable manner, the acquisition module 201 is specifically configured to:

[0104] Acquire aerial images collected by the UAV based on oblique photography and display the aerial images;

[0105] Receive a setting instruction for the aerial image, determine image control points in the aerial image, and perform aerial triangulation based on each of the image control points to obtain measurement data.

[0106] In an implementable manner, the obtaining module 201 is further specifically configured to:

[0107] Project the pixel coordinates in the aerial image into three-dimensional coordinates in a three-dimensional space according to the measurement result, determine three-dimensional point cloud data in the three-dimensional space, and construct a digital elevation model according to the three-dimensional point cloud data;

[0108] After optimizing each of the three-dimensional coordinates based on the ground elevation value of the digital elevation model, resample according to the optimized three-dimensional coordinates to obtain new pixel coordinates, and generate a preliminary orthophoto map corresponding to each of the aerial images according to the new pixel coordinates;

[0109] Stitch each of the preliminary orthophoto maps to obtain a digital orthophoto map.

[0110] In an implementable manner, the training module 202 is specifically configured to:

[0111] After adding Gaussian noise to the second image, perform Gaussian blur processing on the second image to obtain a third image;

[0112] Cut the first image and the third image into at least two first sub-images and second sub-images respectively, and the image sizes of the first sub-images and the second sub-images are the same;

[0113] Construct first training data pairs from the first sub-images and the second sub-images corresponding to the same image coordinates, and integrate each of the first training data pairs to obtain a training data set.

[0114] In an implementable manner, the training module 202 is further specifically configured to:

[0115] Obtain regional satellite images of each of the terrain regions, and adjust the regional satellite images based on geometric correction so that the geographical positions of the regional satellite images are consistent with those of the first image;

[0116] Construct second training data pairs based on the first image and the regional satellite images;

[0117] Integrate each of the first training data pairs and the second training data pairs to obtain a training data set.

[0118] In an implementable manner, the training module 202 is further specifically configured to:

[0119] Obtain an initial generative adversarial network, where the generator of the initial generative adversarial network is a deep network with at least two dense residual learning blocks, and the discriminator of the initial generative adversarial network is a U-Net discriminator with skip connections;

[0120] Train the initial generative adversarial network based on the first training data to obtain a preliminary trained network, and optimize the preliminary trained network based on the second training data to obtain a trained generative adversarial network.

[0121] In an implementable manner, the reconstruction module 203 is specifically configured to:

[0122] Obtain a target image to be reconstructed, read the target image in blocks as at least two sub-blocks, and there is an overlapping area with a preset number of pixels between adjacent sub-blocks;

[0123] Generate reconstruction results for each of the sub-blocks based on the trained generative adversarial network;

[0124] For the overlapping pixels corresponding to the overlapping area in the reconstruction results, use the weighted pixel sum of the edge distances of the overlapping pixels at the same position in each reconstruction result as the final pixel value at that position to obtain the reconstructed image corresponding to the target image.

[0125] Those skilled in the art can clearly understand that the technical solutions of the embodiments of the present application can be implemented by means of software and / or hardware. The "units" and "modules" in this specification refer to software and / or hardware that can independently complete or cooperate with other components to complete specific functions, where the hardware can be, for example, a Field-Programmable Gate Array (FPGA), an Integrated Circuit (IC), etc.

[0126] Each processing unit and / or module of the embodiments of the present application can be implemented by an analog circuit that implements the functions described in the embodiments of the present application, or can be implemented by software that executes the functions described in the embodiments of the present application.

[0127] See Figure 3 , which shows a schematic structural diagram of an electronic device involved in the embodiments of the present application. This electronic device can be used to implement Figure 1 The method in the shown embodiments. As Figure 3 shown, the electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.

[0128] Among them, the communication bus 302 is used to realize the connection and communication between these components.

[0129] Among them, the user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0130] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0131] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire electronic device 300 through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305, the processor 301 performs various functions of the electronic device 300 and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0132] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 305 may further be at least one storage device located far from the aforementioned processor 301. Such asFigure 3 As shown in Figure 3 , the memory 305, which is a computer storage medium, may include an operating system, a network communication module, a user interface module, and program instructions.

[0133] In Figure 3 In the electronic device 300 shown in Figure 3 , the user interface 303 is mainly used to provide an interface for the user to input data and obtain the data input by the user; and the processor 301 can be used to call the satellite image super-resolution reconstruction application program based on UAV images stored in the memory 305, and specifically perform the following operations:

[0134] Obtain the aerial images collected by the UAV, perform aerial triangulation based on the image control points set in the aerial images, and orthorectify the aerial images according to the digital elevation model generated by the measurement results to obtain digital orthophoto maps;

[0135] Select the first images of different terrain regions in the digital orthophoto maps, downsample the first images to obtain second images, construct the first training data pairs based on the first images and the second images, and train and generate a generative adversarial network based on the first training data pairs. Each of the first training data pairs includes a high-resolution image and a low-resolution image with the same region size and region coordinates;

[0136] Obtain the target image to be reconstructed, read the target image in blocks into at least two sub-blocks, generate the reconstruction results of the sub-blocks based on the trained generative adversarial network, and splice the reconstruction results based on the edge gradient fusion method to obtain the reconstructed image corresponding to the target image.

[0137] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above method are implemented. Among them, the computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0138] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0139] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0140] In the several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0141] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0143] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. And the aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0144] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.

[0145] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and practicing the present disclosure herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not described in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for super-resolution reconstruction of satellite images based on drone images, characterized in that: The method comprises: Acquire aerial images collected by a drone based on oblique photography, and display the aerial images; Receiving a setting instruction for the aerial image, determining image control points in the aerial image, performing aerial triangulation based on each of the image control points to obtain measurement data, and orthorectifying the aerial image according to a digital elevation model generated by the measurement results to obtain a digital orthophoto map; Selecting first images of different terrain areas in the digital orthophoto map, downsampling the first images to obtain second images, adding Gaussian noise to the second images, and performing Gaussian blur processing on the second images to obtain third images; Cutting the first image and the third image into at least two first sub-images and a second sub-image respectively, wherein the first sub-image and the second sub-image have the same image size; constructing a first sub-image and a second sub-image corresponding to the same image coordinates into a first training data pair, integrating each of the first training data pairs to obtain a training data set, and training a generative adversarial network based on the first training data pairs, wherein each of the first training data pairs includes a high-resolution image and a low-resolution image with the same region size and region coordinates; Obtain a target image to be reconstructed, divide the target image into at least two sub-blocks, generate reconstruction results of each sub-block based on a trained generative adversarial network, and splice the reconstruction results based on an edge gradient fusion method to obtain a reconstructed image corresponding to the target image.

2. The method according to claim 1, characterized in that The digital elevation model generated according to the measurement result is used to orthorectify the aerial image to obtain a digital orthophoto map, including: Projecting pixel coordinates in the aerial image into three-dimensional coordinates in a three-dimensional space according to the measurement results, determining three-dimensional point cloud data in the three-dimensional space, and constructing a digital elevation model according to the three-dimensional point cloud data; After optimizing each of the three-dimensional coordinates based on the ground elevation value of the digital elevation model, resampling the optimized three-dimensional coordinates to obtain new pixel coordinates, and generating a preliminary orthophoto map corresponding to each of the aerial images based on the new pixel coordinates; The preliminary orthophoto images are stitched together to obtain a digital orthophoto image.

3. The method according to claim 1, characterized in that The step of integrating the first training data pairs to obtain a training data set includes: Acquire a regional satellite image of each of the terrain areas, and adjust the regional satellite image based on geometric correction so that the geographical location of the regional satellite image is consistent with the first image; constructing a second training data pair based on the first image and a regional satellite image; The first training data pairs and the second training data pairs are integrated to obtain a training data set.

4. The method according to claim 3, characterized in that The training of a generative adversarial network based on the first training data comprises: Obtain an initial generative adversarial network, wherein the generator of the initial generative adversarial network is a deep network with at least two dense residual learning blocks, and the discriminator of the initial generative adversarial network is a U-Net discriminator with skip connections; The initial generative adversarial network is trained based on the first training data pair to obtain a preliminary training network, and the preliminary training network is optimized based on the second training data pair to obtain a trained generative adversarial network.

5. The method according to claim 1, characterized in that The step of obtaining a target image to be reconstructed, dividing the target image into at least two sub-blocks, generating reconstruction results of each sub-block based on a trained generative adversarial network, and splicing each reconstruction result based on an edge gradient fusion method to obtain a reconstructed image corresponding to the target image includes: Acquire a target image to be reconstructed, and read the target image into at least two sub-blocks, where there is an overlapping area of ​​a preset number of pixels between adjacent sub-blocks; Generate a reconstruction result of each sub-block based on a trained generative adversarial network; For overlapping pixels corresponding to overlapping areas in the reconstruction results, the weighted pixel sum of edge distances of overlapping pixels at the same position in each reconstruction result is taken as the final pixel value of the position to obtain a reconstructed image corresponding to the target image.

6. A satellite image super-resolution reconstruction device based on drone image assistance, characterized in that: The device comprises: The acquisition module is used to acquire the aerial image collected by the drone based on oblique photography and display the aerial image; receive the setting instruction for the aerial image, determine the image control points in the aerial image, and perform aerial triangulation based on each of the image control points to obtain measurement data, and perform orthorectification on the aerial image according to the digital elevation model generated by the measurement result to obtain a digital orthophoto map; A training module is used to select a first image of different terrain areas in the digital orthophoto map, obtain a second image after downsampling the first image, add Gaussian noise to the second image, and then perform Gaussian blur processing on the second image to obtain a third image; cut the first image and the third image into at least two first sub-images and second sub-images respectively, and the image sizes of the first sub-image and the second sub-image are the same; construct the first sub-image and the second sub-image corresponding to the same image coordinates into a first training data pair, integrate each of the first training data pairs to obtain a training data set, and train a generative adversarial network based on the first training data pairs, each of the first training data pairs includes a high-resolution image and a low-resolution image with the same area size and area coordinates; The reconstruction module is used to obtain the target image to be reconstructed, read the target image into at least two sub-blocks, generate reconstruction results of each sub-block based on the trained generative adversarial network, and splice the reconstruction results based on the edge gradient fusion method to obtain the reconstructed image corresponding to the target image.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, wherein the computer-readable storage medium has instructions stored therein, and when the instructions are executed on a computer or a processor, the computer or the processor executes the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and system for generating aeronautical digital orthoimage map

    CN111707238A

  • Remote sensing image super-resolution reconstruction method and device, equipment and storage medium

    CN113516591A