Real-time positioning method for drone inspection images

By calculating the matching parameters of the drone video frame and DOM image, the real-time and accuracy problems of drone video image positioning are solved, and the stable positioning of interest points under the conditions of large changes in lighting and land objects is achieved, achieving m-level accuracy and second-level speed.

CN115950435BActive Publication Date: 2025-07-22YELLOW RIVER ENG CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310137098.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2025-07-22
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

The prior art cannot achieve real-time, accurate and robust positioning of drone video images, especially when lighting, landform details and resolution changes greatly, it is impossible to accurately locate the location of the target points of interest.

Method used

By obtaining the POS data of the DOM image in the inspection area and the UAV video frame, the center latitude, azimuth, length and width of the reference DOM image corresponding to the video frame are calculated, and the optimal correction parameters of the video frame are solved based on the gradient intensity, and the longitude and latitude coordinates of the points of interest are finally solved.

Benefits of technology

Under the large changes in lighting, land detail and resolution, a steady matching between video frames and DOM images is achieved, and the longitude and latitude coordinates of ground interest points are solved in real time with m-level accuracy. The solution time of a single frame is in seconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115950435B_ABST
    Figure CN115950435B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time positioning method for UAV inspection images. By reading the video frames captured during the UAV inspection process, the POS data of the video frames, and the DOM images of the inspection area in real time, the central longitude and latitude, azimuth angle, length, and width of the reference DOM image corresponding to the video frame are calculated, and the reference DOM image is extracted. The video frame is robustly and roughly matched with the reference DOM image, and the optimal correction parameters of the video frame are solved based on the gradient intensities of the video frame and the roughly matched DOM image. The longitude and latitude coordinates of the interest points are solved according to the pixel coordinates of the interest points in the video frame and the optimal correction parameters. The advantages of the present invention are that it proposes methods for rough matching of video frames and precise correction of video frames. Even when there are significant differences in lighting, ground object details, and resolution between the UAV video frames and the DOM of the inspection area, it can still robustly achieve the rough matching between the two, and through precise correction, the longitude and latitude coordinates of the ground interest points can be solved in real time with an accuracy of the order of meters, and the solution time for a single frame is at the second level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of UAV monitoring, and particularly to a real-time positioning method for UAV inspection images. Background Art

[0002] Inspection is a routine task in many industries. In the early stage, manual visual inspection was mainly used. With the development of technology, UAVs have become new inspection tools due to their significantly higher efficiency than human labor. A common method is to equip UAVs with imaging devices to capture video images of inspection targets in real time, and determine the inspection time, the location of the inspection target, and the specific situation of the inspection target through the video images. Common methods for accurately determining the target location include manual verification method, photogrammetry method, and computer vision technology.

[0003] Among them, in the manual verification method, the inspection video of the UAV is browsed manually, suspicious target locations are marked, and then accurate location information is obtained through on-site measurement. This method cannot meet the real-time requirements of UAV inspection, and there are also significant safety hazards in manual on-site verification operations under special circumstances (such as heavy rain, landslides, and debris flows), and it has gradually become unable to meet the requirements of inspection work.

[0004] Photogrammetry mainly includes forward intersection method of stereo image pairs and aerial triangulation method. The forward intersection method of stereo image pairs can real-time intersect the three-dimensional coordinates of ground points using the POS information of UAV video frames, but the accuracy is extremely poor and cannot meet the accuracy requirements of inspection image positioning. The aerial triangulation method includes strip block adjustment method and bundle block adjustment method, etc. It can accurately calculate the three-dimensional coordinates of ground points using the highly overlapping images of multiple strips, but it requires post-processing of relevant data and cannot meet the real-time requirements of inspection positioning.

[0005] Computer vision technology mainly includes stereo vision methods, image matching methods, and optical flow tracking methods. Among them, the stereo vision method uses a binocular camera to accurately calculate the three-dimensional coordinates of image points by using the precise three-dimensional pose of the binocular camera and the parallax of corresponding image points in two images. However, existing binocular camera calibration methods, including the DLT (Direct Linear Transform) method, PnP (Perspective-n-Point) method, PST (Perspective Similar Triangle) method, etc., are only applicable to binocular measurement systems where the relative positions of the two cameras remain unchanged and the distance is short, and are not applicable to the three-dimensional pose calibration of drones during flight, resulting in the inability to accurately calculate the three-dimensional position information of inspection images. The image matching method can use the inspection image to match the existing DOM, and calculate the position information of the inspection image based on the coordinates of the DOM. However, existing image matching technologies, including SIFT, SURF, ORB, etc., cannot achieve accurate image matching at all when the ground object details change greatly, which also leads to the inability to accurately calculate the position information of the inspection image. The optical flow tracking methods, including the Farneback polynomial method and the LK optical flow method, all have strong preset conditions, namely small displacement, unchanged brightness, and consistent regional motion. They can only be used for the relative positioning of inspection video images and cannot be used for the absolute positioning of inspection images.

[0006] That is to say, at present, there is no method that can calculate the position of drone video images in real time, accurately and robustly. Summary of the Invention

[0007] The purpose of the present invention is to provide a real-time positioning method for drone inspection images, which can calculate the position information of the target of interest in real time, accurately and robustly through drone video images.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions:

[0009] The real-time positioning method for drone inspection images of the present invention includes the following steps:

[0010] S1, obtaining the DOM image of the inspection area;

[0011] S2, reading the video frames and the POS data of the video frames captured during the drone inspection in real time;

[0012] S3, calculating the central longitude and latitude, azimuth, length, and width of the reference DOM image corresponding to the video frame according to the POS data of the video frame and the DOM image of the inspection area, and extracting the reference DOM image;

[0013] S4, performing a robust rough match between the video frame and the reference DOM image to obtain a roughly matched DOM image;

[0014] S5. Calculate the optimal correction parameters of the video frame based on the gradient intensity of the video frame and the coarsely matched DOM image;

[0015] S6. Calculate the longitude and latitude coordinates of the interest points based on the pixel coordinates of the interest points in the video frame and the optimal correction parameters.

[0016] Further, the central longitude and latitude of the reference DOM image are determined by the geodetic longitude and latitude in the POS data of the video frame; the azimuth angle is determined by the course deviation angle in the POS data of the video frame; the length The calculation formula in pixels is: , where , represents the flight height in the POS data of the video frame, represents the physical length of the image acquisition device, represents the focal length of the image acquisition device; the width The calculation formula in pixels is: , where , represents the physical width of the image acquisition device; represents the ground resolution of the inspection area DOM image.

[0017] Further, step S4 specifically includes the following steps:

[0018] S4.1. Scale the extracted reference DOM image according to the ground resolution of the inspection area DOM image and the ground resolution of the video frame to obtain the first reference DOM image;

[0019] S4.2. Calculate the best coarse matching position of the RGB between the video frame and the first reference DOM image;

[0020] S4.3. Calculate the best coarse matching position of the gradient intensity between the video frame and the first reference DOM image;

[0021] S4.4. Calculate the robust best coarse matching position of the video frame and the first reference DOM image according to the best coarse matching position of the RGB and the best coarse matching position of the gradient intensity;

[0022] S4.5. Extract the coarsely matched DOM image according to the robust best coarse matching position, the length and width of the video frame.

[0023] Further, step S5 specifically includes the following steps:

[0024] S5.1. Construct an energy difference equation based on the gradient intensity of the video frame and the coarsely matched DOM image;

[0025] S5.2; Construct a parameter correction equation based on the energy difference equation to correct the scaling ratio, displacement, rotation, and distortion between the video frame and the coarsely matched DOM image;

[0026] S5.3; Construct an optimization criterion based on the presumption that there should be a minimum energy difference between the corresponding image points of the video frame and the coarsely matched DOM image;

[0027] S5.4; Solve the optimal correction parameters of the video frame based on the optimization criterion.

[0028] The advantage of the present invention lies in proposing a method for coarse matching of video frames and fine correction of video frames. Even when there are significant differences in illumination, ground object details, and resolution between the UAV video frames and the DOM of the inspection area, it can still robustly achieve the coarse matching of the two, and through precise correction, the longitude and latitude coordinates of the ground interest points can be calculated in real time with an accuracy of the order of meters, and the calculation time for a single frame is at the second level. Description of the Drawings

[0029] Figure 1 is the flowchart of the real-time positioning method for the UAV inspection images of the present invention.

[0030] Figure 2 is the schematic diagram for extracting the reference DOM image in the method of the present invention.

[0031] Figure 3 is the schematic diagram for the coarse matching between the video frame and the reference DOM image in the method of the present invention.

[0032] Figure 4 is an example of the DOM image of the UAV inspection area in the method of the present invention.

[0033] Figure 5 is an example of a certain video frame captured in real time by the UAV in the method of the present invention.

[0034] Figure 6 is Figure 5 the reference DOM image corresponding to the video frame in

[0035] Figure 7 is Figure 5 the coarsely matched DOM image corresponding to the video frame in

[0036] Figure 8 is Figure 7 the gradient intensity image of the coarsely matched DOM image in

[0037] Figure 9 is Figure 7 the gradient intensity image of the video frame in

[0038] Figure 10 is for Figure 5 the example after the correction of the video frame image in

[0039] Figure 11 Yes Figure 9 Effect display of superposition of video frame and DOM image of inspection area. Specific implementation manners

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] When using an unmanned aerial vehicle (UAV) for inspection, if the target can be detected in real time from the video images captured by the UAV in real time and the precise position of the target can be determined, it has extremely high value for timely discovering and eliminating potential hazards. However, there is currently no method for real-time, precise and robust calculation of the target position information from the UAV video images.

[0042] As Figure 1 shown, the real-time positioning method for the UAV inspection images of the present invention proposes a strategy based on the overall matching of the video frame and the DOM image of the inspection area, and can calculate the position of the inspection video frame in real time with an accuracy of m level, specifically including the following steps:

[0043] S1. Obtain the DOM image of the inspection area;

[0044] The DOM image (Digital Orthophoto Map) is a set of digital orthophoto images generated by digital differential rectification and mosaicking of images and cropping according to a certain map sheet range. It is an image with both map geometric accuracy and image features and can be used as the background control information for map analysis.

[0045] S2. Read the video frame and the POS data of the video frame captured during the UAV inspection in real time;

[0046] The POS data (position and orientation system) is a high-precision position and attitude measurement system combined with IMU / DGPS. It uses the GPS receivers installed on the UAV and the GPS receivers on one or more base stations on the ground to synchronously and continuously observe the GPS satellite signals, and adopts differential GPS positioning (DGPS) technology to achieve precise positioning. At the same time, an inertial measurement unit (IMU) is used to sense the acceleration of the aircraft or other carriers to obtain information such as the speed and attitude of the carrier.

[0047] S3. Calculate the central longitude and latitude, azimuth, length, and width of the reference DOM image corresponding to the video frame based on the POS data of the video frame and the DOM image of the inspection area, and extract the reference DOM image;

[0048] As Figure 2 shown, Region 1 is the complete DOM image covering the entire inspection area, Region 2 is the drone during inspection, and the video frame captured by the drone in real time, i.e., the image, is Region 3. The rough-matched DOM image area of this video frame on the DOM is as shown in Region 4 in Figure 2 . Region 5 in Figure 2 is the extracted reference DOM image.

[0049] The methods for determining the central longitude and latitude, azimuth, length, and width of the reference DOM image are as follows:

[0050] The central longitude and latitude are determined by the geodetic longitude and latitude in the POS data of the video frame ; the azimuth is determined by the flight deviation angle in the POS data of the video frame ; the length is calculated in pixels using formula (1) as:

[0051] Formula (1)

[0052] where , represents the flight altitude in the POS data of the video frame, represents the physical length of the image acquisition device, represents the focal length of the image acquisition device; represents the ground resolution of the DOM image of the inspection area.

[0053] The width is calculated in pixels using formula (2) as:

[0054] Formula (2)

[0055] where , represents the flight altitude in the POS data of the video frame, represents the physical width of the image acquisition device, represents the focal length of the image acquisition device; represents the ground resolution of the DOM image of the inspection area. The physical width and physical length of the image acquisition device are actually the length and width of the captured image.

[0056] According to the determined central longitude and latitude , flight deviation angle , length and width , it is possible to extract the reference DOM image shown in Region 5 from the inspection area DOM image shown in Region 1 of Figure 2 .

[0057] S4. Robustly and roughly match the video frame with the reference DOM image to obtain a roughly matched DOM image;

[0058] Since the acquisition time and acquisition means of the video frames captured by the drone and the inspection area DOM image are inconsistent, there are significant differences between the two in terms of illumination, ground object details, and resolution. As a result, existing image feature matching methods, such as those based on SIFT (Scale-invariant feature transform), SURF (Speeded Up Robust Features), and ORB (Oriented FAST and Rotated BRIEF), cannot be applied to the automatic matching of drone video frames and the inspection area DOM at all. This application can ensure a robust rough match between the two in the case of significant differences in illumination, ground object details, and resolution between the drone video frames and the inspection area DOM.

[0059] Specifically, it includes the following steps:

[0060] S4.1. Scale the extracted reference DOM image according to the ground resolution of the inspection area DOM image and the ground resolution of the video frame to obtain a first reference DOM image; The calculation formula (3) is:

[0061] Formula (3)

[0062] Where, represents the scaled first reference DOM image; represents the reference DOM image; , represents the ground resolution of the inspection area DOM image, represents the ground resolution of the video frame.

[0063] S4.2. Calculate the RGB best rough matching position of the video frame and the first reference DOM image ; is the pixel coordinate of the upper left corner of the RGB best rough matching area, as shown in Figure 3 .

[0064] Determine the RGB best rough matching position of the video frame and the scaled reference DOM according to formula (4), that is

[0065]

[0066] Formula (4)

[0067] wherein

[0068] represents the gray value of the R channel of the video frame,

[0069] represents the gray value of the R channel of the scaled reference DOM;

[0070] represents the gray value of the G channel of the video frame,

[0071] represents the gray value of the G channel of the scaled reference DOM;

[0072] represents the gray value of the B channel of the video frame,

[0073] represents the gray value of the B channel of the scaled reference DOM,

[0074] represents the total number of pixels of a single video frame.

[0075] S4.3, calculate the best rough matching position of the gradient intensity between the video frame and the first reference DOM image ; wherein is the pixel coordinate of the upper left corner of the best rough matching area of the gradient intensity, as Figure 2 shown.

[0076] That is, determine the best rough matching position of the gradient intensity between the video frame and the scaled first reference DOM image according to formula (5).

[0077]

[0078] Formula (5)

[0079] wherein

[0080]

[0081] represents the first-order gradient of the video frame in the x direction, represents the first-order gradient of the video frame in the y direction; represents the first-order gradient of the scaled first reference DOM image in the x direction, represents the first-order gradient of the scaled first reference DOM image in the y direction.

[0082] S4.4, calculate the robust best rough matching position of the video frame and the first reference DOM image , is the pixel coordinate of the upper left corner of the robust optimal rough matching region, as Figure 3 shown. That is, calculate the robust optimal rough matching position of the video frame and the first reference DOM image after scaling according to formula (6):

[0083]

[0084] Formula (6)

[0085] where

[0086] ,

[0087] is the RGB optimal rough matching position, is the gradient intensity optimal rough matching position, represents the length of the video frame (in pixels), represents the width of the video frame (in pixels); , represents the total number of pixels of a single video frame.

[0088] S4.5, as Figure 3 shown, according to the robust optimal rough matching position , the length and width of the video frame, the rough matching DOM image region 4 that best matches the video frame can be extracted from the reference DOM image region 5 as shown in Figure 3 .

[0089] As Figures 4 - 7 shown, it is an effect example of using the method of this application to determine the rough matching DOM image region that best matches the video frame. Among them Figure 4 is the DOM image of the inspection area, Figure 5 is the video frame image taken by the drone in real time; Figure 6 is the effect diagram of extracting the reference DOM image after step S3. Figure 7 is the effect of the rough matching DOM image region that best matches the video frame obtained after step S4 calculation, scaling the reference DOM image, performing RGB optimal rough matching, gradient intensity optimal rough matching, and robust optimal rough matching. As Figures 4 - 7 can be seen, in the case of large differences in illumination, ground object details, and resolution, this application can still robustly achieve the rough matching of the video frame and the reference DOM, and the matching accuracy is about 10m level.

[0090] S5, solve the optimal correction parameters of the video frame based on the gradient intensity of the video frame and the rough matching DOM image.

[0091] This application uses the coarsely matched DOM as a reference, selects the points in the video frame where the gradient intensity is greater than a certain threshold as feature points, constructs a system of polynomial parameter equations, and corrects the video frame through a robust optimization process, and can calculate the longitude and latitude coordinates of the ground interest points with an accuracy of m levels. Specifically, it includes the following steps:

[0092] S5.1, construct an energy difference equation based on the gradient intensities of the video frame and the coarsely matched DOM image; that is, construct an energy difference equation according to formula (7).

[0093]

[0094] Formula (7)

[0095] Among them, represents the gradient intensity of the video frame at ; represents the gradient intensity of the coarsely matched DOM image at ; represents the shooting time of the video frame, represents the generation time of the DOM image of the inspection area; represents the displacement amount of the image point on the video frame to the homologous image point on the coarsely matched DOM image.

[0096] S5.2, construct a parameter correction equation based on the energy difference equation to correct the scaling ratio, displacement, rotation, and distortion between the video frame and the coarsely matched DOM image; that is, construct a polynomial parameter correction equation formula (8) based on formula (7)

[0097] Formula (8)

[0098] Among them,

[0099] ;

[0100] ;

[0101] ;

[0102] Among them, (x, y) are the pixel coordinates of the interest point on the video frame;

[0103] are the parameters to be solved; in essence, they are the coefficients of two third-order polynomials with respect to the pixel coordinates x and y of the interest point on the video frame. X is a column vector composed of - .

[0104] S5.3, construct an optimization criterion formula (9) according to formula (8) and based on the presumption that there should be a minimum energy difference between the homologous image points of the video frame and the coarsely matched DOM image;

[0105]

[0106] Formula (9)

[0107] S5.4. Solve for the optimal correction parameters of the video frame based on the optimization criterion. The process is as follows:

[0108] Select feature points in the video frame that satisfy Formula (10), that is

[0109] Formula (10)

[0110] Where represents the gradient intensity of the video frame at , represents the mean of the gradient intensity of the video frame, represents the standard deviation of the gradient intensity of the video frame.

[0111] Use the feature points to construct the coefficient matrix and the constant matrix ;

[0112] Let , represents the identity matrix;

[0113] Solve for the parameter , represents the estimate of;

[0114] Solve for the energy difference ;

[0115] Re - weight , ;

[0116] Repeat steps -[[]]END , if the difference between the results of two consecutive iterations is less than the preset threshold, then stop the iteration and obtain the optimal robust estimate .

[0117] S6. Solve for the longitude and latitude coordinates of the interest points based on the pixel coordinates of the interest points in the video frame and the optimal correction parameters. The calculation formula (11) is

[0118]

[0119] Formula (11)

[0120] Among them,

[0121] , represents the conversion parameter from pixel coordinates to longitude and latitude coordinates in the DOM image of the inspection area, which is provided by the producer of the DOM image of the inspection area. represents the pixel coordinates of the interest point of the video frame; represents the longitude and latitude coordinates of the interest point of the video frame; represents the optimal robust estimation of the calibration parameter. Using the method of this application, the pixel coordinates of the interest point can be solved in real time during the UAV shooting process, and the solution time for a single frame is at the second level.

[0122] As Figures 8 - 11 shown, it is a schematic diagram of the effect of gradient intensity calibration in step S5 using the method of the present invention. Among them Figure 8 is the gradient amplitude image of the reference DOM; Figure 9 is the gradient amplitude image of the video frame; Figure 10 is the video frame image after precise calibration, and each pixel point of it already has longitude and latitude coordinate information. In the Arcgis software, with the coordinate information as the reference, Figure 10 is superimposed on Figure 4 the DOM image of the inspection area shown in Figure 11 the gray area shown. After measurement, Figure 10 and Figure 4 the matching accuracy of the DOM image of the inspection area shown can reach the m level.

Claims

1. A real-time positioning method for drone inspection images, characterized in that: It includes the following steps: S1. Obtain the DOM image of the inspection area; S2. Read the video frames and the POS data of the video frames captured during the UAV inspection in real time; S3. According to the POS data of the video frames and the DOM image of the inspection area, calculate the central longitude and latitude, azimuth, length and width of the reference DOM image corresponding to the video frames, and extract the reference DOM image; The longitude and latitude of the center of the reference DOM image are determined by the geodetic longitude and latitude in the POS data of the video frame; the azimuth angle is determined by the course deviation angle in the POS data of the video frame; the length The calculation formula in pixels is: , where , represents the flight altitude in the POS data of the video frame, represents the physical length of the image acquisition device, represents the focal length of the image acquisition device; The said width The calculation formula in pixels is as follows: , where , represents the physical width of the image acquisition device; represents the ground resolution of the DOM image in the inspection area; S4. Perform a robust rough matching between the video frames and the reference DOM image to obtain a roughly matched DOM image; S5. Solve the optimal correction parameters of the video frames based on the gradient intensities of the video frames and the roughly matched DOM image; S6. Solve the longitude and latitude coordinates of the interest points according to the pixel coordinates of the interest points in the video frames and the optimal correction parameters.

2. The real-time positioning method of the drone inspection image according to claim 1, wherein: Step S4 specifically includes the following steps: S4.

1. Scale the extracted reference DOM image according to the ground resolution of the DOM image of the inspection area and the ground resolution of the video frames to obtain a first reference DOM image; S4.

2. Calculate the best rough matching position of the RGB between the video frames and the first reference DOM image; S4.

3. Calculate the best rough matching position of the gradient intensity between the video frames and the first reference DOM image; S4.

4. Calculate the robust best rough matching position of the video frames and the first reference DOM image according to the best rough matching position of the RGB and the best rough matching position of the gradient intensity; S4.

5. Extract the roughly matched DOM image according to the robust best rough matching position, the length and width of the video frames.

3. The real-time positioning method of the drone inspection image according to claim 1, characterized in that: Step S5 specifically includes the following steps: S5.

1. Construct an energy difference equation based on the gradient intensities of the video frames and the roughly matched DOM image; S5.

2. Construct a parameter correction equation based on the energy difference equation to correct the scaling ratio, displacement, rotation, and distortion between the video frames and the roughly matched DOM image; S5.

3. Construct an optimization criterion based on the presumption that the minimum energy difference should exist between the homologous image points of the video frames and the roughly matched DOM image; S5.

4. Solve the optimal correction parameters of the video frames based on the optimization criterion.