A super-resolution reconstruction method for UAV aerial images based on edge artifact removal

By integrating the Sparse Non-local attention mechanism and PatchGAN ideas into the super-resolution reconstruction algorithm of aerial images of drone and combining edge refinement perception strategies, the problems of edge blur and artifacts of drone aerial images were solved, and the reconstruction quality of images was significantly improved.

CN114283059BActive Publication Date: 2025-05-13YANCHENG POWER SUPPLY CO STATE GRID JIANGSU ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111506147.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-05-13
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

Existing drone aerial image super-resolution algorithms have shortcomings in edge processing and texture detail recovery, resulting in blurred edges and artifacts in reconstructed images.

Method used

The super-resolution reconstruction method based on edge artifact removal is adopted, and the Sparse Non-local attention mechanism and spatial pyramid pooling layer sparse representation is integrated into the ESRGAN generation network to enhance the global representation ability of the network; the PatchGAN idea is introduced into the discriminator, and the patch discriminator is used to avoid image blur and artifacts; the edge refinement perception strategy is added to the perception loss calculation to retain important edge information.

Benefits of technology

Effectively remove artifacts during drone aerial image reconstruction, improve the clarity of image edges and the recovery quality of texture details, and improve the effect of super-resolution reconstruction of drone aerial image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283059B_ABST
    Figure CN114283059B_ABST
Patent Text Reader

Abstract

The present invention provides a method for super-resolution reconstruction of unmanned aerial images based on edge artifact removal, and the method comprises the following steps: step (1): down-sampling a real high-resolution image to obtain a corresponding low-resolution image C1; step (2): passing C1 as input into a generation network to generate a high-resolution image C14; step (3): calculating the perceptual loss by combining C14 and the real high-resolution image, and updating the parameters of the generation network and the discriminant network according to the perceptual loss; step (4): judging whether the number of iterations is satisfied, and if not, returning to step (2), and if satisfied, outputting C14. The present invention provides a method for super-resolution reconstruction of unmanned aerial images based on edge artifact removal, which can improve the reconstruction effect of unmanned aerial images, remove artifacts generated during the reconstruction of unmanned aerial images, and better reconstruct the edges of unmanned aerial images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of super-resolution reconstruction of unmanned aerial vehicle aerial images, and in particular relates to a super-resolution reconstruction method of unmanned aerial vehicle aerial images based on edge artifact removal. Background Art

[0002] In the information age, images are the main carrier of information transmission. Drone aerial imaging technology is widely used in the fields of electricity, industrial manufacturing, etc. due to its strong penetration and other advantages. Power detection based on drone aerial imaging technology can quickly repair power equipment, thereby effectively reducing the time cost of equipment maintenance and improving the reliability of equipment operation. Engineers can easily visualize and quantify the thermal images of manufacturing equipment with the help of drone aerial imaging technology.

[0003] Due to the limitations of drone aerial imaging technology and the particularity of drone aerial imaging principles, drone aerial images have problems such as low resolution and blurred images, which affect their effective use. In recent years, with the increasing demand for high-quality drone aerial images in various fields, some drone aerial image super-resolution algorithms have emerged. Some algorithms have improved the accuracy and speed of drone aerial image super-resolution reconstruction by using faster and deeper convolutional neural networks, but the edge processing and texture details of the image cannot be well restored. Traditional network model structures, including super-resolution generative adversarial networks (SRGAN), reconstruct drone aerial images with blurred edges and artifacts, and the texture details cannot be guaranteed. How to improve the accuracy and speed of drone aerial image super-resolution reconstruction while ensuring the quality of drone aerial image reconstruction is the key to solving the problem of applying drone aerial image super-resolution reconstruction technology in real-time scenarios. Summary of the invention

[0004] In order to solve the above problems, the present invention provides a super-resolution reconstruction method for UAV aerial images based on edge artifact removal, which can improve the reconstruction effect of UAV aerial images, remove artifacts generated during the reconstruction of UAV aerial images, and better reconstruct the edges of UAV aerial images.

[0005] The present invention specifically provides a method for super-resolution reconstruction of unmanned aerial images based on edge artifact removal, and the method for super-resolution reconstruction of unmanned aerial images comprises the following steps:

[0006] Step (1): down-sample the real high-resolution image to obtain the corresponding low-resolution image C1;

[0007] Step (2): Pass C1 as input into the generative network to generate a high-resolution image C14;

[0008] First, C1 is passed into the generative network as input, and a layer of convolution preprocessing is performed. The features of the image are extracted through the first layer of convolution to generate the feature map C2;

[0009] Secondly, C2 is passed as input to the SNL-RRDB module for feature extraction and screening, and the output feature map C Nlast And the feature map R flast ;

[0010] Again, C Nlast and R flast Perform weighted operation to complete adaptive adjustment of channel features and obtain feature map C13 containing important features;

[0011] Finally, C13 is convolved once and then input into the upsampling module to enlarge the feature size. The enlarged result is convolved twice to output a high-resolution image C14.

[0012] Step (3): Calculate the perceptual loss by combining C14 and the real high-resolution image, and update the parameters of the generation network and the discriminant network according to the perceptual loss;

[0013] First, C14 and the real high-resolution image are input into the discriminative network to calculate the adversarial loss

[0014] Secondly, use C14 and the real high-resolution image to calculate the content loss L1;

[0015] Again, the edge loss L is calculated using the edge refinement-aware strategy E ;

[0016] Again, combined L1, L E Calculate the perceptual loss L G ;

[0017] Finally, according to the perceptual loss L G Update the parameters of the generating network and the discriminative network;

[0018] Step (4): Determine whether the number of iterations is met. If not, return to step (2). If it is met, output C14.

[0019] Compared with the prior art, the beneficial effects are:

[0020] (1) The Sparse Non-local Attention Mechanism is integrated into the dense residual block RRDB module of the ESRGAN generative network to form the SNL-RRBD block, which can establish a long-range dependency relationship between two pixels at any position in the image, integrate global image information into local features, and enhance the global representation of the network. At the same time, the spatial pyramid pooling layer sparse representation is adopted, which greatly reduces the computational overhead.

[0021] (2) The idea of ​​PatchGAN is introduced into the ESRGAN discriminator. The patch discriminator is used to change the output form of the discriminator into a matrix form, which can more effectively avoid image blur and suppress artifacts.

[0022] (3) When calculating the perception loss, the edge loss guided by the edge refinement perception strategy is added to retain the important edge information of the UAV aerial images and better reconstruct the edges of the UAV aerial images. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flow chart of a method for super-resolution reconstruction of UAV aerial images based on edge artifact removal according to the present invention;

[0024] Figure 2 A network structure diagram of a generation network of a method for super-resolution reconstruction of unmanned aerial images based on edge artifact removal according to the present invention;

[0025] Figure 3 It is a structural diagram of the SNL module of the generation network of the method for super-resolution reconstruction of unmanned aerial images based on edge artifact removal of the present invention;

[0026] Figure 4 This is a structural diagram of the RRDB module of the generation network of the method for super-resolution reconstruction of UAV aerial images based on edge artifact removal of the present invention. DETAILED DESCRIPTION

[0027] The following is a detailed description of a specific implementation of a method for super-resolution reconstruction of drone aerial images based on edge artifact removal according to the present invention in conjunction with the accompanying drawings.

[0028] like Figure 1 As shown, the super-resolution reconstruction method of unmanned aerial image of the present invention comprises the following steps:

[0029] Step (1): down-sample the real high-resolution image to obtain the corresponding low-resolution image C1;

[0030] Step (2): C1 is passed as input to the generative network to generate a high-resolution image C14. The network structure of the generative network is referenced Figure 2 ;

[0031] First, C1 is passed into the generative network as input, and a layer of convolution preprocessing is performed. The features of the image are extracted through the first layer of convolution to generate the feature map C2;

[0032] Secondly, C2 is passed as input to the SNL-RRDB module for feature extraction and screening, and the output feature map C Nlast And the feature map R flast :

[0033] SNL-RRDB is divided into SNL module and RRDB module.

[0034] like Figure 3 The processing in the SNL module is as follows:

[0035] Input feature map C2, size is H×W, channel is C0, C2 is transformed into three feature maps through three 1×1 convolutional layers, denoted as C3, C4, C5, size is H×W×C, where H, W, C are the height, width, and number of channels of the image;

[0036] Input C3 and C4 into the spatial pyramid pooling layer (SPP) for feature extraction to obtain feature maps C6 and C7 with a size of C×S; flatten and transpose functions are used to flatten and transpose C5 to obtain feature map C8 with a size of N×C, where N=W×H;

[0037] Perform matrix multiplication on C7 and C8 to obtain the similarity matrix M of C7 and C8:

[0038] M N×S =Q N×C ×K C×S ,

[0039] Where Q N×C Indicates C8, K C×S Indicates C7;

[0040] Pass M through the Softmax activation function to obtain the normalized similarity matrix M′ with a size of N×S:

[0041] M′=softmax(M),

[0042] Reshape C6 using the reshape function to obtain feature map C9, with a size of S×C;

[0043] Perform matrix multiplication on C8 and C9 to output feature map C10, with size N×C:

[0044] O N×C =M′ N×S ×VS×C ,

[0045] Among them, N×C Represents the feature map C10, V S×C Represents feature map C9;

[0046] C10 is transposed and reshaped by the transpose function and the reshape function. After a layer of convolution, the extracted feature map C11 is output with a size of H×W×C0.

[0047] like Figure 4 The processing in the RRDB module is as follows:

[0048] The input feature map C2 is convolved once, and then the features are enhanced through the ReLU activation function to obtain the enhanced feature map R:

[0049] R = δ(Conv(R pre )),

[0050] Where R pre Represents the input feature map C2, δ is the ReLU activation function;

[0051] The enhanced feature map R is then subjected to a second convolution operation, and then enhanced using the ReLU activation function to obtain the feature map R i , i represents the number of times the convolution process has been performed, and the processing function is as follows:

[0052]

[0053] Where R i It represents the feature map output by the ReLU activation function after the i-th convolution process. δ is the ReLU activation function. After the last convolution layer of RRDB is processed, the feature map R is output by the ReLU activation function. last ;

[0054] R last After performing a convolution and multiplying it by a balance factor β, the output feature map R is f :

[0055] R f =Conv(R last )×β,

[0056] The feature C11 output by the SNL module and the feature map R output by the RRDB module f Perform weighted operations to obtain feature graph C12, and pass C12 as input to the next SNL-RRDB module for processing until the last SNL-RRDB module finishes processing and outputs feature C Nlast And the feature map Rflast ;

[0057] Again, C Nlast and R flast Perform concat operation to complete the adaptive adjustment of channel features and obtain feature map C13 containing important features;

[0058] Finally, C13 is convolved once and then input into the upsampling module to enlarge the feature size. The enlarged result is convolved twice to output a high-resolution image C14.

[0059] Step (3): Calculate the perceptual loss by combining C14 and the real high-resolution image, and update the parameters of the generation network and the discriminant network according to the perceptual loss;

[0060] First, C14 and the real high-resolution image are input into the discriminative network to calculate the adversarial loss

[0061] The discriminator of the PatchGAN model divides C14 into several regions (patches), performs convolution processing on each region, determines the probability of the patch block in the receptive field being true or false, outputs a probability value, and finally outputs a probability matrix, which is weighted averaged and output to obtain

[0062] Secondly, use C14 and the real high-resolution image to calculate the content loss L1:

[0063] L1=||G(x)-X||1,

[0064] Where G(x) refers to the generated high-resolution image C14, and X refers to the real high-resolution image;

[0065] Again, the edge loss L is calculated using the edge refinement-aware strategy E :

[0066] (1) Use the convolution operator to calculate the edge feature map. The weight of the convolution operator is defined as the following matrix M:

[0067]

[0068] C14 and the real high-resolution image are respectively matrix multiplied with the convolution operator to obtain the edge feature maps P1 and P2 of the image. For each value B(i, j) of the feature maps P1 and P2, it is as follows:

[0069]

[0070] Where (i, j) represents the position of the pixel in the image, d ijRepresents the Euclidean distance from the edge of the image to the nearest pixel on the edge of pixel (i, j);

[0071] (2) After taking the union of the two feature maps, an upsampling operation is performed to obtain the cross feature map P 12 :

[0072] P 12 =f up (B G ∪B P ),

[0073] Among them B G Indicates P2, B p represents P1;

[0074] (3) According to the cross feature map P 12 Calculate L E :

[0075]

[0076] Where H, W, C are the height, width, and number of channels of the image, P nij is the value of the pixel corresponding to the cross feature map, δ is the balance coefficient, l nij is the binary cross entropy of the corresponding position;

[0077] Again, combined L1, L E Calculate the perceptual loss L G :

[0078]

[0079] Among them, λ, η, μ are coefficients that balance different loss terms. is the adversarial loss of the discriminator, L1 is the content loss, and L E It is the marginal loss;

[0080] Finally, according to the perceptual loss L G Update the parameters of the generating network and the discriminative network;

[0081] Step (4): Determine whether the number of iterations is met. If not, return to step (2). If it is met, output C14.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit the present invention. A person skilled in the art should understand that the specific implementation of the present invention can be modified or replaced by equivalents, but these modifications or changes are within the scope of protection of the pending claims.

Claims

1. A super-resolution reconstruction method for drone aerial images based on edge artifact removal, characterized in that: The method for super-resolution reconstruction of unmanned aerial images comprises the following steps: Step (1): down-sample the real high-resolution image to obtain the corresponding low-resolution image C1; Step (2): Pass C1 as input into the generative network to generate a high-resolution image C14; Step (3): Calculate the perceptual loss by combining C14 and the real high-resolution image, and update the parameters of the generation network and the discrimination network according to the perceptual loss; Step (4): Determine whether the number of iterations is met. If not, return to step (2). If it is met, output C14. In step (3), the perceptual loss is calculated by combining C14 and the real high-resolution image, and the parameters of the generative network and the discriminative network are updated according to the perceptual loss, including the following steps: First, C14 and the real high-resolution image are input into the discriminative network to calculate the adversarial loss Secondly, use C14 and the real high-resolution image to calculate the content loss L1; Again, the edge loss L is calculated using the edge refinement-aware strategy E , the specific algorithm is: (1) Use the convolution operator to calculate the edge feature map. The weight of the convolution operator is defined as the following matrix M: C14 and the real high-resolution image are respectively matrix multiplied with the convolution operator to obtain the edge feature maps P1 and P2 of the image. For each value B(i, j) of the feature maps P1 and P2, it is as follows: Where (i, j) represents the position of the pixel in the image, d ij Represents the Euclidean distance from the edge of the image to the nearest pixel on the edge of pixel (i, j); (2) After taking the union of the two feature maps, an upsampling operation is performed to obtain a cross feature map P 12 : P 12 =f up (B G ∪B P ), where B G Indicates P2, B p represents P1; (3) According to the cross-feature graph P 12 Calculate L E : Where H, W, C are the height, width, and number of channels of the image, P nij is the value of the pixel corresponding to the cross feature map, δ is the balance coefficient, l nij is the binary cross entropy of the corresponding position; Again, combined L1, L E Calculate the perceptual loss L G ; Finally, according to the perceptual loss L G Update the parameters of the generator and discriminator networks.

2. According to claim 1, a method for super-resolution reconstruction of drone aerial images based on edge artifact removal is characterized in that: In step (2), C1 is passed as input to the generative network, and the generation of the high-resolution image C14 includes the following steps: First, C1 is passed into the generative network as input, and a layer of convolution preprocessing is performed. The features of the image are extracted through the first layer of convolution to generate the feature map C2; Secondly, C2 is passed as input to the SNL-RRDB module for feature extraction and screening, and the output feature map C Nlast And the feature map R flast ; Again, C Nlast and R flast Perform concat operation to complete the adaptive adjustment of channel features and obtain feature map C13 containing important features; Finally, C13 is convolved once and then input into the upsampling module to enlarge the feature size. The enlarged result is convolved twice to output a high-resolution image C14.

3. The method for super-resolution reconstruction of drone aerial images based on edge artifact removal according to claim 2 is characterized in that: C2 is passed as input to the SNL-RRDB module for feature extraction and screening, and the feature graph C is output. Nlast And the characteristic map R flast The specific method is: The SNL-RRDB module is divided into an SNL module and an RRDB module. The processing in the SNL module is as follows: Input feature map C2, size is H×W, channel is C0, C2 is transformed into three feature maps through three 1×1 convolutional layers, denoted as C3, C4, C5, size is H×W×C, where H, W, C are the height, width, and number of channels of the image; Input C3 and C4 into the spatial pyramid pooling layer (SPP) for feature extraction to obtain feature maps C6 and C7 with a size of C×S; flatten and transpose functions are used to flatten and transpose C5 to obtain feature map C8 with a size of N×C, where N=W×H; Perform matrix multiplication on C7 and C8 to obtain the similarity matrix M of C7 and C8: M N×S =Q N×C ×K C×S , Where Q N×C Indicates C8, K C×S Indicates C7; Pass M through the Softmax activation function to obtain the normalized similarity matrix M′ with a size of N×S: M′=softmax(M), Reshape C6 using the reshape function to obtain feature map C9, with a size of S×C; Perform matrix multiplication on C8 and C9 to output feature map C10, with size N×C: O N×C =M′ N×S ×V S×C , Among them, N×C Represents the feature map C10, V S×C Represents feature map C9; C10 is transposed and reshaped by the transpose function and the reshape function. After a layer of convolution, the extracted feature map C11 is output with a size of H×W×C0. The processing in the RRDB module is as follows: The input feature map C2 is convolved once, and then the features are enhanced through the ReLU activation function to obtain the enhanced feature map R: R=δ(Conv(R pre )), Where R pre Represents the input feature map C2, δ is the ReLU activation function; The enhanced feature map R is then subjected to a second convolution operation, and then enhanced using the ReLU activation function to obtain the feature map R i , i represents the number of times the convolution process has been performed, and the processing function is as follows: Where R i It represents the feature map output by the ReLU activation function after the i-th convolution process. δ is the ReLU activation function. After the last convolution layer of RRDB is processed, the feature map R is output by the ReLU activation function. last ; R last After performing a convolution and multiplying it by a balance factor β, the feature map R is output. f : R f =Conv(R last )×β, The feature C11 output by the SNL module and the feature map R output by the RRDB module f The weighted operation is performed to obtain the feature graph C12, and C12 is passed as input to the next SNL-RRDB module for processing, until the last SNL-RRDB module finishes processing and outputs the feature graph C12. Nlast And the characteristic map R flast .

4. The method for super-resolution reconstruction of drone aerial images based on edge artifact removal according to claim 1 is characterized in that: Input C14 and the real high-resolution image into the discriminative network to calculate the adversarial loss The algorithm is as follows: the PatchGAN model discriminator divides C14 into several regions, performs convolution processing on each region, determines the probability of the patch block in the receptive field being true or false, outputs a probability value, and finally outputs a probability matrix, performs weighted average output on the probability matrix, and obtains 5. The method for super-resolution reconstruction of drone aerial images based on edge artifact removal according to claim 1 is characterized in that: The algorithm for calculating the content loss L1 using C14 and the real high-resolution image is: L1 = ||G(x)-X||1, where G(x) refers to the generated high-resolution image C14, and X refers to the real high-resolution image.

6. The method for super-resolution reconstruction of drone aerial images based on edge artifact removal according to claim 1 is characterized in that: Combination L1, L E Calculate the perceptual loss L G The algorithm is: Among them, λ, η, μ are coefficients that balance different loss terms, is the adversarial loss of the discriminator, L1 is the content loss, L E is the edge loss.

Citation Information

Patent Citations

  • Super-resolution reconstruction method based on conditional generative adversarial network

    CN109978762A

  • KR20200084434A