A UAV view distortion correction method based on high-order polynomial and deep network fusion

Through the method of fusion of higher-order polynomials and deep networks, the distortion parameters are extracted in combination with geometric encoder and spatial encoder, and correction flow is generated and fusion is solved, and the image distortion problem of drone is achieved with high precision and robust image correction effect.

CN119417734BActive Publication Date: 2025-05-06SHAANXI FUTURE ECONOMIC IND HOLDINGS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510028276.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Drone images are prone to introduce radial or tangential distortion under factors such as flight altitude, viewing angle changes and lens design, which affects the accuracy and reliability of subsequent visual analysis tasks.

Method used

Using the method of fusion of higher-order polynomials and deep networks, the radial and tangential distortion parameters are extracted respectively through geometric encoder and spatial encoder, and the correction flow is generated in combination with the higher-order polynomial model, and the final correction flow is fused by a confidence-based exponential attenuation function to obtain the final correction flow applied to the distorted image.

Benefits of technology

Without relying on special hardware such as calibration boards, the image distortion is stably restored, which improves the accuracy and robustness of image processing, especially in complex scenarios, and shows stronger adaptability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119417734B_ABST
    Figure CN119417734B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing, and in particular to a method for correcting the distortion of unmanned aerial vehicle views by fusing a high-order polynomial and a deep network, the method comprising the following steps: respectively using a geometric encoder and a spatial encoder to parse the distorted image, and analyzing the radial distortion parameters and the tangential distortion parameters; respectively inputting the radial distortion parameters and the tangential distortion parameters output by the two encoders into a high-order polynomial model to generate a specific correction flow, and calculating the confidence corresponding to each pixel of the correction flow through a fully connected layer; fusing the correction flows output by the two encoders through an exponential decay function based on the confidence to obtain a final correction flow, and applying the final correction flow to the distorted image to obtain a corrected image. The present invention can stably restore the distortion of images with radial or tangential distortion without relying on special hardware such as a calibration plate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a method for correcting unmanned aerial vehicle (UAV) view distortion by fusing a high-order polynomial with a deep network. Background Art

[0002] As an important tool for data collection, drones are widely used in scenarios such as traffic flow monitoring, disaster assessment, and urban planning. However, due to the flight altitude and viewing angle changes of drones, as well as the wide-angle or telephoto design of the lens, radial or tangential distortion is easily introduced, especially in large areas or complex scenes. The distortion problem is more significant. This image distortion will affect the accuracy and reliability of subsequent visual analysis tasks, such as target recognition and trajectory tracking. The present invention is committed to proposing a correction method for restoring distorted images caused by the lens, generating a confidence-based correction flow, effectively solving the image distortion problem, and improving the processing accuracy and robustness of traffic monitoring data. Summary of the invention

[0003] The present application provides a drone view distortion correction method that integrates high-order polynomials and deep networks, which can stably restore the distortion of images with radial or tangential distortion without relying on special hardware such as calibration plates.

[0004] To solve the above problems, this application provides the following solutions:

[0005] R1. Use geometric encoders respectively and spatial encoder Distorted image Analyze and find out the radial distortion parameters and and tangential distortion parameters and ,in and It is a geometric encoder The extracted radial distortion parameters and tangential distortion parameters are and It is a geometric encoder The extracted radial distortion parameters and tangential distortion parameters;

[0006] R2. Input the radial distortion parameters and tangential distortion parameters output by the two encoders into the high-order polynomial model to generate a specific correction flow and And the confidence of the correction flow corresponding to each pixel is calculated through the fully connected layer and ;

[0007] R3. The correction streams output by the two encoders are fused through an exponential decay function based on confidence to obtain the final correction stream , and the correction flow Applied to distorted images The corrected image is obtained .

[0008] Wherein, the geometric encoder The processing process is divided into two stages: edge feature extraction and geometric feature encoding;

[0009] The geometric encoder Edge feature extraction refers to using the Canny edge detector to extract the edge features of the input image. Processing to generate edge map ;

[0010] The geometric encoder The geometric feature encoding refers to the use of ResNet50 to encode the edge graph Encode and generate radial distortion parameters and tangential distortion parameters ;

[0011] The spatial encoder The processing process is realized by encoding and decoding the input image The encoding of

[0012] The spatial encoder The encoding process refers to using MobileNet to encode the input image Extract The spatial encoding feature map of the layer ,in For passing Multi-scale feature maps generated by layer convolutional neural networks, feature maps With different resolutions and feature dimensions ;

[0013] The spatial encoder The decoding process refers to gradually restoring the spatial resolution of the feature map by upsampling layer by layer through transposed convolution, introducing skip connections, and combining with the spatial encoder The low-level detail features of the encoding process and the high-level semantic features of the decoder are combined through the fully connected layer. Spatial decoding feature map of the layer Calculate the radial distortion parameters and tangential distortion parameters ;

[0014] The spatial encoder The decoding process of is to gradually restore the feature map by upsampling layer by layer through transposed convolution. The process is described as:

[0015]

[0016] Where ConvTranspose represents the transposed convolution operation, For the Spatial decoding feature map of the layer After The layer transposed convolutional neural network and passed Multi-scale feature maps generated by layer convolutional neural network The spliced ​​generated The spatial decoded feature map of the layer;

[0017] The correction flow represents a mapping from the image coordinate system to the correction coordinate system, where is the pixel coordinate in the distorted image. After correction, the pixel will be mapped to the new position ,Right now:

[0018] , ,

[0019] in, is the corrected pixel coordinate, and Represents the pixel width and pixel height of the image respectively.

[0020] The radial distortion parameters and tangential distortion parameters output by the input encoder are used to generate a specific correction flow using a high-order polynomial model. The process is as follows: Each pixel coordinate on , according to the radial distortion parameter Calculate the radial distortion factor , and then using the radial distortion factor , tangential distortion parameters and pixel coordinates , calculate the new position of the pixel in the corrected coordinate system ;

[0021] The radial distortion factor The calculation method is as follows:

[0022]

[0023] in, Indicates that it contains A set of radial distortion parameters, Represents pixels To the center of the image The square distance of

[0024] The pixel coordinates New position in the corrected coordinate system The calculation method is:

[0025]

[0026]

[0027] The confidence corresponding to each pixel of the correction flow is calculated through the fully connected layer The specific steps are:

[0028]

[0029]

[0030] in, and To learn the parameters, The Sigmoid activation function ensures that the confidence level is between 0 and 1.

[0031] The confidence-based exponential decay function is used to fuse the final correction flow The process is as follows:

[0032]

[0033]

[0034] in, Reason and The confidence level of the calculated horizontal axis, Reason and The confidence level of the calculated ordinate, yes and The calculated correction flow of the abscissa, yes and The calculated correction flow of the ordinate, Reason and The confidence level of the calculated horizontal axis, Reason and The confidence level of the calculated ordinate, yes and The calculated correction flow of the abscissa, yes and The calculated correction flow of the ordinate, is the attenuation coefficient, which is used to control the degree of confidence attenuation.

[0035] The will correct flow Applied to distorted images The corrected image is obtained Refers to the correction flow New position mapped to the corrected coordinate system .

[0036] Compared with the prior art, the above technical solution provides a method for correcting drone view distortion by integrating high-order polynomials and deep networks, which has the following beneficial effects:

[0037] 1. The geometric encoder and spatial encoder are used to extract the geometric features and spatial features of the input image respectively, so that the method can adapt to different scenes and various types of surveillance image distortion, especially in complex scenes with significant depth changes such as intersections and high-altitude surveillance, showing stronger robustness and adaptability;

[0038] 2. The high-order polynomial model is introduced as prior knowledge, which effectively ensures the stability and accuracy of the distortion correction process. It can handle radial distortion and tangential distortion at the same time and meet the requirements of high-precision correction.

[0039] 3. A confidence-based exponential decay function is introduced to dynamically weighted fuse the correction streams generated by the geometric encoder and the spatial encoder, thereby significantly improving the accuracy and stability of the correction streams and ensuring the accuracy and consistency of the final correction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0041] Figure 1 The present invention is a flowchart of a method for correcting drone view distortion by integrating high-order polynomials and deep networks in an embodiment of the present invention.

[0042] Figure 2 The distorted image in the present invention Holistic model framework for radial and tangential distortion correction. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0044] An embodiment of the present application provides a method for correcting drone view distortion by fusing high-order polynomials and deep networks. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0045] like Figure 1 As shown, an embodiment of a method for correcting drone view distortion by integrating a high-order polynomial and a deep network in an embodiment of the present application, the specific steps are as follows:

[0046] R1. The embodiments of the present invention use geometric encoders respectively and spatial encoder Distorted image Analyze and find out the radial distortion parameters and and tangential distortion parameters and ,in and It is a geometric encoder The extracted radial distortion parameters and tangential distortion parameters are and It is a geometric encoder The extracted radial distortion parameters and tangential distortion parameters;

[0047] R2. In the embodiment of the present invention, the radial distortion parameters and tangential distortion parameters output by the two encoders are respectively input into the high-order polynomial model to generate a specific correction flow and And the confidence of the correction flow corresponding to each pixel is calculated through the fully connected layer and ;

[0048] R3. In the embodiment of the present invention, the correction streams output by the two encoders are merged through an exponential decay function based on confidence to obtain the final correction stream , and the correction flow Applied to distorted images The corrected image is obtained .

[0049] Among them, Figure 2 As shown, the geometric encoder described in the embodiment of the present invention The processing process is divided into two stages: edge feature extraction and geometric feature encoding;

[0050] The geometric encoder Edge feature extraction refers to using the Canny edge detector to extract the edge features of the input image. Processing to generate edge map ;

[0051] The geometric encoder The geometric feature encoding refers to the use of ResNet50 to encode the edge graph Encode and generate radial distortion parameters and tangential distortion parameters ;

[0052] The spatial encoder The processing process is realized by encoding and decoding the input image The encoding of

[0053] The spatial encoder The encoding process refers to using MobileNet to encode the input image Extract The spatial encoding feature map of the layer ,in For passing Multi-scale feature maps generated by layer convolutional neural networks, feature maps With different resolutions and feature dimensions ;

[0054] The spatial encoder The decoding process refers to gradually restoring the spatial resolution of the feature map by upsampling layer by layer through transposed convolution, introducing skip connections, and combining with the spatial encoder The low-level detail features of the encoding process and the high-level semantic features of the decoder are combined through the fully connected layer. Spatial decoding feature map of the layer Calculate the radial distortion parameters and tangential distortion parameters ;

[0055] The spatial encoder The decoding process of is to gradually restore the feature map by upsampling layer by layer through transposed convolution. The process is described as:

[0056]

[0057] Where ConvTranspose represents the transposed convolution operation, For the Spatial decoding feature map of the layer After The layer transposed convolutional neural network and passed Multi-scale feature maps generated by layer convolutional neural network The spliced ​​generated The spatial decoded feature map of the layer;

[0058] The correction flow represents a mapping from the image coordinate system to the correction coordinate system, where is the pixel coordinate in the distorted image. After correction, the pixel will be mapped to the new position ,Right now:

[0059] , ,

[0060] in, is the corrected pixel coordinate, and Represents the pixel width and pixel height of the image respectively.

[0061] The radial distortion parameters and tangential distortion parameters output by the input encoder are used to generate a specific correction flow using a high-order polynomial model. The process is as follows: Each pixel coordinate on , according to the radial distortion parameter Calculate the radial distortion factor , and then using the radial distortion factor , tangential distortion parameters and pixel coordinates , calculate the new position of the pixel in the corrected coordinate system ;

[0062] The radial distortion factor The calculation method is as follows:

[0063]

[0064] in, Indicates that it contains A set of radial distortion parameters, Represents pixels To the center of the image The square distance of

[0065] The pixel coordinates New position in the corrected coordinate system The calculation method is:

[0066]

[0067]

[0068] The confidence corresponding to each pixel of the correction flow is calculated through the fully connected layer The specific steps are:

[0069]

[0070]

[0071] in, and To learn the parameters, The Sigmoid activation function ensures that the confidence level is between 0 and 1.

[0072] The confidence-based exponential decay function is used to fuse the final correction flow The process is as follows:

[0073]

[0074]

[0075] in, Reason and The confidence level of the calculated horizontal axis, Reason and The confidence level of the calculated ordinate, yes and The calculated correction flow of the abscissa, yes and The calculated correction flow of the ordinate, Reason and The confidence level of the calculated horizontal axis, Reason and The confidence level of the calculated ordinate, yes and The calculated correction flow of the abscissa, yes and The calculated correction flow of the ordinate, is the attenuation coefficient, which is used to control the degree of confidence attenuation.

[0076] The will correct flow Applied to distorted images The corrected image is obtained Refers to the correction flow New position mapped to the corrected coordinate system .

[0077] The embodiments of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments. For ordinary technicians in this technical field, after knowing the contents recorded in the present invention, they can make several equivalent changes and substitutions without departing from the principle of the present invention. These equivalent changes and substitutions should also be regarded as belonging to the protection scope of the present invention.

Claims

1. A drone view distortion correction method based on high-order polynomial and deep network fusion, characterized in that: The method for correcting the distortion of drone view by integrating high-order polynomial and deep network includes: Use the geometric encoder D g and spatial encoder D s Analyze the distorted image I and find out the radial distortion parameter K g and K s and tangential distortion parameters and Where K g and is the geometric encoder D g The extracted radial distortion parameters and tangential distortion parameters, K s and is the spatial encoder D s The extracted radial distortion parameters and tangential distortion parameters, the radial distortion parameters are used to describe the outward or inward deformation of the image center, and the tangential distortion parameters describe the tangential displacement caused by the tilt of the lens. The geometric encoder D g It is an encoder module that extracts edge features and encodes geometric features of the input content. The geometric encoder D g The input is the distorted image I, and the output is the radial distortion parameter K g and tangential distortion parameters Spatial Encoder D s It is an encoder module that encodes the input content and then decodes it. The spatial encoder D s The input is the distorted image I, and the output is the radial distortion parameter K s and tangential distortion parameters Spatial Encoder D s The encoding process refers to using MobileNet to extract n layers of spatial encoding feature maps from the input image I in is a multi-scale feature map generated by the i-layer convolutional neural network, and the feature map i has different resolutions H i ×W i and feature dimension C i , the spatial encoder D s The decoding process refers to gradually restoring the spatial resolution of the feature map by upsampling layer by layer through transposed convolution, introducing skip connections, and combining with the spatial encoder D s The low-level detail features of the encoding process and the high-level semantic features of the decoder are combined, and the spatial decoding feature map of the nth layer is decoded through the fully connected layer Calculate the radial distortion parameter K s and tangential distortion parameters The radial distortion parameters and tangential distortion parameters output by the two encoders are respectively input into the high-order polynomial model to generate a specific correction flow f g and f s And calculate the confidence C corresponding to each pixel of the correction flow through the fully connected layer g and C s ; The correction streams output by the two encoders are fused through an exponential decay function based on confidence to obtain the final correction stream f, and the correction stream f is applied to the distorted image I to obtain the corrected image I′.

2. The method for correcting drone view distortion by integrating high-order polynomials and deep networks according to claim 1, characterized in that: in, The geometric encoder D g The edge feature extraction refers to processing the input image I using the Canny edge detector to generate an edge map I e ; The geometric encoder D g The geometric feature encoding refers to the use of ResNet50 to encode the edge graph I e Encode and generate radial distortion parameter K g and tangential distortion parameters The spatial encoder D s The decoding process of is to gradually restore the feature map by upsampling layer by layer through transposed convolution. The process is described as: Where ConvTranspose represents the transposed convolution operation. is the spatial decoding feature map of the i-th layer The multi-scale feature map generated by the i+1th layer of transposed convolutional neural network and the i-th layer of convolutional neural network The spatial decoding feature map of the i+1th layer generated by splicing; The correction flow f represents a mapping from the image coordinate system to the correction coordinate system, where (u, v) ∈ W × H is the pixel coordinate in the distorted image. After correction, the pixel will be mapped to a new position (u′, v′), that is: Where (u′, v′) is the corrected pixel coordinate, W and H represent the pixel width and pixel height of the image respectively; The radial distortion parameters and tangential distortion parameters output by the input encoder are input to a high-order polynomial model to generate a specific correction flow, and the process is as follows: for each pixel coordinate (u, v)∈W×H on the input image I, a radial distortion factor R is calculated according to the radial distortion parameter K, and then the new position (u′, v′) of the pixel in the correction coordinate system is calculated using the radial distortion factor R, the tangential distortion parameter (p1, p2) and the pixel coordinate (u, v); The radial distortion factor R is calculated as follows: R=1+k1·r 2 +k2·r 4 +...+k m ·r 2m Where K = {k1, k2, …, k m } represents a set of m radial distortion parameters, r represents the square distance from the pixel (u, v) to the image center (W / 2, H / 2); The new position (u′, v′) of the pixel coordinate (u, v) in the correction coordinate system is calculated as follows: u′=(u-W / 2)·R+2·p1·(u-W / 2)·(v-H / 2)+P2·(r 2 +2·(u-W / 2) 2 ) v′=(v-H / 2)·R+p1·(r 2 +2·(v-H / 2) 2 )+2·p2·(u-W / 2)·(v-H / 2) The confidence C corresponding to each pixel of the correction flow is calculated through the fully connected layer. u ,c v ) The specific steps are: c u =σ(W f ·((u′-u)+I)+b f ) c v =σ(W f ·((v′-v)+I)+b f ) Among them, W f and b f is the learning parameter, σ is the Sigmoid activation function, ensuring that the confidence is between 0 and 1; The confidence-based exponential decay function is used to fuse the final correction flow f = (f u ,f v )The process is as follows: in, By K g and The confidence level of the calculated horizontal axis, By K g and The confidence level of the calculated ordinate, It's K g and The calculated correction flow of the abscissa, It's K g and The calculated correction flow of the ordinate, By K s and The confidence level of the calculated horizontal axis, By K s and The confidence level of the calculated ordinate, It's K s and The calculated correction flow of the abscissa, It's K s and The calculated correction flow of the ordinate, α is the attenuation coefficient, which is used to control the attenuation degree of the confidence; Applying the correction flow f to the distorted image I to obtain the corrected image I′ refers to mapping the correction flow f to a new position (u′, v′) of the correction coordinate system.

Citation Information

Patent Citations

  • Method for constructing correction model in unmanned aerial vehicle image radiation correction

    CN110110730A

  • Method for realizing on-orbit star map geometric distortion correction based on two-dimensional Legendre neural network and star sensor

    CN115689915A