A vision-based method, device and system for obtaining elevation information

By acquiring and analyzing the tilted environment image sequences by drones, a three-dimensional environmental map was established and plane fitting was performed, which solved the accuracy and efficiency of elevation data acquisition under dense vegetation, and achieved high accuracy and low computational volume of elevation information.

CN118482684BActive Publication Date: 2025-06-24STATE GRID JIANGSU ELECTRIC POWER CO LTD TAIZHOU POWER SUPPLY BRANCH +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410581064.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-11
Publication Date
2025-06-24
Estimated Expiration
2044-05-11

AI Technical Summary

Technical Problem

In simulated construction sites, especially in the lever frame setting process, it is difficult for the prior art to accurately obtain ground elevation data under dense vegetation, and the data processing efficiency is low.

Method used

The drone-mounted camera collects dynamic tilt environment image sequences, performs coverage area recognition and pixel filling, analyzes image motion information, establishes a three-dimensional environment map, and solves elevation information through plane fitting.

Benefits of technology

Improve the accuracy of elevation information, reduce the amount of data calculation, avoid environmental climate impact, and improve modeling accuracy by continuously optimizing camera posture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118482684B_ABST
    Figure CN118482684B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and system for obtaining elevation information based on vision. The method includes: collecting a dynamic sequence of tilted environment images of a region to be detected by a camera carried by a drone; performing coverage area recognition on the sequence of tilted environment images, and removing and pixel-filling the recognized coverage areas to obtain a processed sequence of environment images; analyzing based on the processed sequence of environment images to obtain image motion information; establishing a three-dimensional environment map according to the image motion information; performing plane fitting based on the three-dimensional environment map to establish a plane equation, and solving the plane equation to obtain the elevation information of the region to be detected; this method can effectively improve the accuracy of obtaining elevation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of engineering technology, and particularly to a method, device and system for obtaining elevation information based on vision. Background Art

[0002] In a simulated construction site, especially in the link of erecting a spanning frame, complex natural environment challenges are often faced. For example, the construction area may be located in a densely vegetated area such as dense trees or paddy fields, and is affected by signal interference from buildings and vegetation in the urban environment or dense forests. These factors will affect the oblique photogrammetry, making it impossible to accurately obtain the absolute elevation information of the bottom of the spanning frame, and high-performance computing resources are required to process the data.

[0003] Remote sensing technology is difficult to accurately obtain the ground elevation data under thick vegetation, and the mapping accuracy cannot be effectively guaranteed. While the ground vehicle-mounted lidar scanning technology can effectively penetrate vegetation to obtain real ground information, but due to the extremely large number of laser points and image points generated by vehicle-mounted laser scanning, up to hundreds of millions of points, it has high requirements for data processing software and hardware devices, and the data processing efficiency is low.

[0004] Patent (CN113723568A) discloses a method for obtaining the elevation of feature points of a remote sensing image based on multi-sensors and sea level, belonging to the fields of autonomous navigation and remote sensing images. This method controls the drone to cruise, uses the SLAM method to establish a SLAM map, obtains the fitting equation of the drone cruising plane in the SLAM coordinate system according to the coordinates at each moment during the cruising process, obtains the equation of the sea level in the SLAM coordinate system according to the altitude information of the drone SLAM initialization collected by the barometer, performs feature matching on the remote sensing image and the airborne camera image, obtains the elevation information of the feature points corresponding to the real-world three-dimensional points, and finally adds the elevation information to the remote sensing image. The equation of the sea level in the SLAM coordinate system can be obtained, and the elevation information of the feature points in the remote sensing map can be obtained, providing more help for the users of the remote sensing map. In this method, the barometer is used to obtain the altitude information of the drone. However, the accuracy of the barometer is greatly affected by the environment, so the accuracy of the elevation information obtained by this method is not high. Summary of the Invention

[0005] The present invention provides a method, device and system for obtaining elevation information based on vision, which can effectively improve the accuracy of the obtained elevation information.

[0006] A method for obtaining elevation information based on vision includes:

[0007] Collecting a dynamic sequence of oblique environment images of a to-be-detected area by a camera carried by a drone;

[0008] Perform coverage area recognition on the sequence of inclined environment images, and perform rejection and pixel filling on the recognized coverage areas to obtain a processed sequence of environment images;

[0009] Analyze based on the processed sequence of environment images to obtain image motion information;

[0010] Establish a three-dimensional environment map according to the image motion information;

[0011] Perform plane fitting based on the three-dimensional environment map, establish a plane equation, and solve the plane equation to obtain the elevation information of the area to be detected.

[0012] Further, recognize the coverage area in the sequence of inclined environment images based on a convolutional neural network.

[0013] Further, the image motion information includes the motion speed of each pixel in the sequence of environment images;

[0014] Analyze based on the processed sequence of environment images to obtain image motion information, including:

[0015] Based on the optical flow equation, obtain the relationship between the motion speed of each pixel and the pixel gray value in the sequence of environment images;

[0016] Select a neighborhood window for each pixel, convert the optical flow equation into a system of linear equations within the neighborhood window, and solve the system of linear equations based on the least squares method to obtain the motion speed of each pixel in the sequence of environment images.

[0017] Further, establish a three-dimensional environment map according to the image motion information, including:

[0018] In the sequence of environment images, select feature points;

[0019] According to the motion speed of the pixels of the feature points in the sequence of environment images, obtain the pixel coordinates of the feature points in the sequence of environment images;

[0020] Establish an epipolar constraint between feature points in two adjacent frames of environment images;

[0021] Form a cost function with all the epipolar constraints of the feature points, and optimize the cost function using the least squares method to obtain the optimal fundamental matrix of each epipolar constraint;

[0022] According to the optimal fundamental matrix, map the feature points of two adjacent frames of environment images to each other's epipolar lines to obtain feature point pairs;

[0023] According to the pixel coordinates of the feature point pairs, use triangulation to calculate the three-dimensional point coordinates corresponding to the feature point pairs;

[0024] Estimate the camera pose based on the projection of the three-dimensional point coordinates in the corresponding environmental image, and optimize the projection error;

[0025] Construct a three-dimensional environmental map based on the optimized three-dimensional point coordinates of multiple feature points obtained by calculation.

[0026] Further, estimating the camera pose based on the projection of the three-dimensional point coordinates in the corresponding environmental image and optimizing the projection error includes:

[0027] Establish a camera model based on the three-dimensional point coordinates and the pixel coordinates of the corresponding feature points;

[0028] Establish an error model for the pixel coordinates of the feature points and the three-dimensional point coordinates based on the camera model;

[0029] Optimize the error model, find the optimal camera pose parameters that minimize the result of the error model, and adjust the pose of the camera according to the calculated optimal camera pose parameters.

[0030] Further, perform plane fitting based on the three-dimensional environmental map, establish a plane equation, and solve the plane equation to obtain the elevation information of the area to be detected, including:

[0031] Construct a design matrix and an observation vector for each three-dimensional point;

[0032] Establish a system of linear equations based on the design matrix and the observation vector;

[0033] Solve the system of linear equations using the least squares method to obtain the plane parameters;

[0034] Construct a plane equation based on the plane parameters, and obtain the elevation information of the area to be detected according to the reference plane elevation and the plane equation.

[0035] Further, the elevation information is calculated by the following formula:

[0036]

[0037] Where Z is the elevation of the area to be detected, Z0 represents the reference plane elevation, a, b, c, d are plane parameters, and x, y are the point coordinates on the plane of the area to be detected.

[0038] Further, the method further includes:

[0039] Collect the true elevation data of the ground control points;

[0040] Align the real elevation data with the elevation information spatially, calculate the root mean square error between the real elevation data and the elevation information, and obtain the confidence level of the elevation information based on the root mean square error.

[0041] An apparatus for obtaining elevation information based on vision as described above, comprising:

[0042] An image acquisition module, configured to collect a dynamic sequence of tilted environment images of the area to be detected through a camera carried by a drone;

[0043] An image processing module, configured to identify the covered areas in the sequence of tilted environment images, and perform rejection and pixel filling on the identified covered areas to obtain a processed sequence of environment images;

[0044] An image analysis module, configured to analyze based on the processed sequence of environment images to obtain image motion information;

[0045] A map building module, configured to build a three-dimensional environment map according to the image motion information;

[0046] A fitting module, configured to perform plane fitting based on the three-dimensional environment map, establish a plane equation, and solve the plane equation to obtain the elevation information of the area to be detected.

[0047] Further, the image processing module identifies the covered areas in the sequence of tilted environment images based on a convolutional neural network.

[0048] Further, the image motion information includes the motion speed of each pixel in the sequence of environment images;

[0049] The image analysis module analyzes based on the processed sequence of environment images to obtain image motion information, including:

[0050] Based on the optical flow equation, obtain the relationship between the motion speed of each pixel in the sequence of environment images and the pixel gray value;

[0051] For each pixel, select a neighborhood window, convert the optical flow equation into a linear equation system within the neighborhood window, and solve the linear equation system based on the least squares method to obtain the motion speed of each pixel in the sequence of environment images.

[0052] Further, the map building module builds a three-dimensional environment map according to the image motion information, including:

[0053] In the sequence of environment images, select feature points;

[0054] According to the motion speed of the pixels of the feature points in the sequence of environment images, obtain the pixel coordinates of the feature points in the sequence of environment images.

[0055] Establish the epipolar constraint of feature points in two adjacent frames of environmental images;

[0056] Form a cost function with the epipolar constraints of all feature points, and optimize the cost function using the least squares method to obtain the optimal fundamental matrix for each epipolar constraint;

[0057] According to the optimal fundamental matrix, map the feature points of two adjacent frames of environmental images to each other's epipolar lines to obtain feature point pairs;

[0058] According to the pixel coordinates of the feature point pairs, use triangulation to calculate the three-dimensional point coordinates corresponding to the feature point pairs;

[0059] Estimate the camera pose based on the projection of the three-dimensional point coordinates in the corresponding environmental image, and optimize the projection error;

[0060] Construct a three-dimensional environmental map based on the optimized three-dimensional point coordinates of multiple calculated feature points.

[0061] Further, the map building module estimates the camera pose based on the projection of the three-dimensional point coordinates in the corresponding environmental image and optimizes the projection error, including:

[0062] Establish a camera model based on the three-dimensional point coordinates and the pixel coordinates of their corresponding feature points;

[0063] Based on the camera model, establish an error model between the pixel coordinates of the feature points and the three-dimensional point coordinates;

[0064] Optimize the error model, find the optimal camera pose parameters that minimize the result of the error model, and adjust the camera pose according to the calculated optimal camera pose parameters.

[0065] Further, the fitting module performs plane fitting based on the three-dimensional environmental map, establishes a plane equation, and solves the plane equation to obtain the elevation information of the area to be detected, including:

[0066] Construct a design matrix and an observation vector for each three-dimensional point;

[0067] Establish a system of linear equations based on the design matrix and the observation vector;

[0068] Solve the system of linear equations using the least squares method to obtain the plane parameters;

[0069] Construct a plane equation based on the plane parameters, and obtain the elevation information of the area to be detected according to the elevation of the reference plane and the plane equation.

[0070] Further, the elevation information is calculated by the following formula:

[0071]

[0072] Among them, Z is the elevation of the area to be detected, Z0 represents the elevation of the reference plane, a, b, c, and d are plane parameters, and x and y are the coordinates of points on the plane of the area to be detected.

[0073] Furthermore, the method further includes:

[0074] Collect the true elevation data of the ground control points;

[0075] Align the true elevation data and the elevation information spatially, calculate the root mean square error between the true elevation data and the elevation information, and obtain the confidence level of the elevation information according to the root mean square error.

[0076] A vision-based elevation information acquisition system includes a processor, a storage device, and a drone. A camera is mounted on the drone. The storage device stores multiple instructions, and the processor is configured to read the multiple instructions and execute the above method.

[0077] The vision-based elevation information acquisition method, device, and system provided by the present invention at least include the following beneficial effects:

[0078] (1) Construct a three-dimensional environmental map based on a dynamic sequence of tilted environment images. Compared with laser point clouds, it can effectively reduce the amount of data calculation and is not affected by environmental climate, resulting in relatively high accuracy of the obtained elevation information;

[0079] (2) Eliminate the covered areas in the obtained sequence of tilted environment images, reduce the impact of the covered areas on modeling, and thereby improve the accuracy of the elevation information;

[0080] (3) Continuously optimize the camera pose during the modeling process to improve the accuracy of the modeling, and thereby improve the accuracy of the elevation information. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 It is a flowchart of an embodiment of the vision-based elevation information acquisition method provided by the present invention.

[0082] Figure 2 It is a flowchart of an embodiment of analyzing the sequence of environmental images to obtain image motion information in the vision-based elevation information acquisition method provided by the present invention.

[0083] Figure 3 It is for establishing a three-dimensional environmental map in the vision-based elevation information acquisition method provided by the present invention Figure 1 a flowchart of an embodiment.

[0084] Figure 4 Flow chart of an embodiment for optimizing projection error in the vision-based elevation information acquisition method provided by the present invention.

[0085] Figure 5 Flow chart of an embodiment for plane fitting in the vision-based elevation information acquisition method provided by the present invention.

[0086] Figure 6 Schematic structural diagram of an embodiment of the vision-based elevation information acquisition device provided by the present invention. Detailed implementation manners

[0087] To better understand the above technical solution, the following will describe the above technical solution in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0088] Referring to Figure 1 , in some embodiments, a vision-based elevation information acquisition method is provided, including:

[0089] S1. Collect a dynamic sequence of inclined environment images of the area to be detected through a camera carried by a drone;

[0090] S2. Identify the covered areas in the sequence of inclined environment images, and remove and fill pixels in the identified covered areas to obtain a processed sequence of environment images;

[0091] S3. Analyze based on the processed sequence of environment images to obtain image motion information;

[0092] S4. Establish a three-dimensional environment map according to the image motion information;

[0093] S5. Perform plane fitting based on the three-dimensional environment map, establish a plane equation, and solve the plane equation to obtain the elevation information of the area to be detected.

[0094] Specifically, in step S1, the camera carried by the drone continuously and synchronously collects images from different perspectives of vertical and four inclinations to obtain a dynamic sequence of inclined environment images.

[0095] Further, in step S2, the covered areas in the sequence of inclined environment images are identified based on a convolutional neural network, where the covered areas may be densely vegetated areas.

[0096] In a vegetation-dense environment, vegetation may occlude the collected image data, resulting in deviations in map modeling and positioning. To solve this problem, vegetation occlusion processing techniques such as vegetation occlusion culling or partial occlusion processing can be adopted to improve the accuracy of map modeling and positioning. Specifically, CNN can be used here. There are relatively mature techniques in the existing technology that can perform vegetation occlusion culling and partial occlusion processing. For the pixels in the culled part, pixel filling is performed to ensure the continuity of the image.

[0097] Further, in step S3, the image motion information includes the motion speed of each pixel in the environmental image sequence.

[0098] Reference Figure 2 , based on the processed environmental image sequence for analysis, image motion information is obtained, including:

[0099] S31. Based on the optical flow equation, obtain the relationship between the motion speed of each pixel and the pixel gray value in the environmental image sequence.

[0100] S32. Select a neighborhood window for each pixel, convert the optical flow equation into a linear equation system within the neighborhood window, and solve the linear equation system based on the least squares method to obtain the motion speed of each pixel in the environmental image sequence.

[0101] Specifically, in step S31, assuming that between consecutive frames of the environmental image sequence, the pixel gray value in the image remains unchanged, the interval between two adjacent frames of the image is short, and the motion of adjacent pixels is consistent in space, the mathematical expression of the brightness constancy constraint can be obtained:

[0102] I(x,y,t) = I(x + Δx,y + Δy,t + Δt); (1)

[0103] where I(x,y,t) represents the gray value of the image at pixel (x, y) at time t, and I(x + Δx,y + Δy,t + Δt) represents the gray value of pixel (x, y) moving to (x + Δx,y + Δy) at time t + Δt.

[0104] Take the partial derivatives of both sides of the brightness constancy constraint with respect to space and time, and use the chain rule for differentiation. Assuming that I(x,y,t) is a continuously differentiable function, then there is:

[0105]

[0106] Perform Taylor expansion of (x + Δx,y + Δy,t + Δt) at (x,y,t), and get:

[0107]

[0108] Substitute the Taylor expansion result into the derivative formula of the brightness constancy constraint, and we get:

[0109]

[0110] and are the motion speeds of the pixel on the x-axis and y-axis, which are denoted as u and v respectively. After organizing the obtained equations, we can get the optical flow equation:

[0111]

[0112] where (u, v) is the motion speed of the pixel (x, y) on the x-axis and y-axis in the image, is the gradient of the image gray value in the x and y directions, is the change rate of the gray value with time.

[0113] The (u, v) obtained by solving equation (5) is the motion speed of the pixel.

[0114] Further, in step S32, a neighborhood window is selected for each pixel, and the optical flow equation is converted into a linear equation system within the neighborhood window. The linear equation system is solved based on the least squares method to obtain the motion speed of each pixel in the environmental image sequence.

[0115] Specifically, denote and as Ix, Iy, and It respectively. If n pixels within the neighborhood window all satisfy the above equations, then we have:

[0116]

[0117] It can be converted into the following form of a linear equation system:

[0118]

[0119] Solve the equation system (7) using the least squares method, then we have:

[0120]

[0121] Thus, the velocity estimation of the pixel can be obtained to achieve feature tracking.

[0122] Further, referring to Figure 3 , in step S4, according to the image motion information, a three-dimensional environmental map is established, including:

[0123] S41. Select feature points in the environmental image sequence;

[0124] S42. Obtain the pixel coordinates of the feature points in the environmental image sequence according to the motion speed of the pixels of the feature points in the environmental image sequence;

[0125] S43. Establish the epipolar constraint of the feature points in two adjacent frames of environmental images;

[0126] S44. Combine the epipolar constraints of all the feature points into a cost function, and optimize the cost function by using the least squares method to obtain the optimal fundamental matrix of each epipolar constraint;

[0127] S45. According to the optimal fundamental matrix, map the feature points of two adjacent frames of environmental images to each other's epipolar lines to obtain feature point pairs;

[0128] S46. According to the pixel coordinates of the feature point pairs, use triangulation to calculate the three-dimensional point coordinates corresponding to the feature point pairs;

[0129] S47. Estimate the camera pose according to the projection of the three-dimensional point coordinates in the corresponding environmental image, and optimize the projection error;

[0130] S48. Construct a three-dimensional environmental map according to the optimized three-dimensional point coordinates of multiple feature points obtained by calculation.

[0131] Specifically, in step S41, the feature points can be prominent features in the environment to be detected, such as the corners, edges, windows of buildings, etc. They have unique geometric features, which are helpful for subsequent positioning and mapping.

[0132] Furthermore, in step S42, according to the motion speed of the pixels obtained by the above method, the pixel coordinates of the feature points in the environmental image sequence can be obtained.

[0133] Furthermore, in step S43, the principle of multi-view geometry is used for three-dimensional reconstruction, and the epipolar constraint is used to describe the geometric relationship between two images to recover the camera motion and the scene.

[0134] The epipolar constraint refers to the positional relationship between the feature points in these views and the camera when observing the same scene point in two views. For the corresponding pixel point pairs (x, x') in two images, their epipolar constraint is expressed as:

[0135] x' T Fx = 0; (9)

[0136] where F is the fundamental matrix, and x, x' respectively represent the corresponding feature points in the two views.

[0137] Furthermore, in step S44, combine the epipolar constraint equations of all feature point pairs into a cost function, and use the least squares method to minimize this cost function:

[0138]

[0139] Among them, F is the fundamental matrix to be estimated, and ∑ i represents the summation over all feature point pairs. By optimizing the above cost function, the optimal fundamental matrix F can be obtained, where i is the number of the feature point.

[0140] Among them, the least squares method is used to optimize the cost function to obtain the optimal fundamental matrix for each epipolar constraint, including:

[0141] Define the minimized residual r:

[0142] r = F * X - c; (11)

[0143] Among them, the fundamental matrix F to be estimated is an N*M matrix, X is an M*1 unknown vector, and c is an N*1 known vector, which can be a feature point;

[0144] By finding the least squares solution x LS , such that the norm ||r|| 2 of r is minimized, the least squares solution x LS can be obtained through the following formula:

[0145] x LS = (F T F) -1 F T c; (12)

[0146] Furthermore, in step S45, after obtaining the optimal fundamental matrix F, the feature points in the image are mapped to the corresponding epipolar lines between each other. That is, for the feature point x i in the first image, the corresponding point x' i in the second image must lie on its corresponding epipolar line.

[0147] Furthermore, in step S46, for each pair of known feature point pairs (x i , x' i ), their corresponding 3D point coordinates can be calculated by the method of triangulation.

[0148] Specifically, in two adjacent views, the optical center point of the camera and the feature point form a straight line, and the intersection of the two straight lines formed by the two feature points and the optical center point of the camera respectively can be regarded as the actual spatial point. By determining the two straight lines and the connection line of the two feature points to form a triangle, the 3D coordinates of the corresponding spatial point can be calculated according to the pixel coordinates of the two feature points, that is, the 3D point coordinates corresponding to the feature point.

[0149] Furthermore, referring to Figure 4, in step S47, the obtained feature points corresponding to the three-dimensional points can be regarded as the projections of the three dimensions under different camera poses in two views. However, there will be a certain deviation between the feature points and the actual projection points. Therefore, the camera pose is adjusted to minimize the distance between the feature points and the actual projection points, specifically including:

[0150] S47a. Establish a camera model based on the three-dimensional point coordinates and the pixel coordinates of the corresponding feature points;

[0151] S47b. Establish an error model regarding the pixel coordinates of the feature points and the three-dimensional point coordinates according to the camera model;

[0152] S47c. Optimize the error model, find the optimal camera pose parameters that minimize the result of the error model, and adjust the camera pose according to the calculated optimal camera pose parameters.

[0153] Specifically, in step S47a, assuming n three-dimensional points, the coordinates of the i-th three-dimensional point are: P i = [X i , Y i , Z i T , and the pixel coordinates of its projection, that is, the pixel coordinates of the feature point, are p(p xi , p yi ). The camera internal parameter is K, the pose parameter is T, and the camera model is:

[0154]

[0155] where s i represents the depth of the feature point in the camera coordinate system.

[0156] Furthermore, in step S47b, according to this camera model, an error model regarding the pixel coordinates of the feature points and the three-dimensional point coordinates is established. In formula (13), the left side represents the projection position of the feature point, and the right side represents the observed position of the feature point. Written in matrix form as:

[0157] s i p = KTP i ; (14)

[0158] Due to the camera pose and the observation point noise, there is an error between the left and right sides of the above equation. Therefore, the errors are summed to construct a least squares problem to obtain the error model e:

[0159]

[0160] Further, in step S47c, the above error model is optimized, that is, the optimal pose parameters T of the camera are found to minimize the result of the error model. Specifically, the iterative formula can be expressed as:

[0161] f(x + Δx) ≈ f(x) + J(x) T Δx; (16)

[0162] where J(x) T is the derivative of f(x) with respect to x and is a Jacobian matrix.

[0163] The current goal is to find the increment Δx such that ||f(x + Δx)|| 2 is minimized. To solve for Δx, a linear least squares problem needs to be solved:

[0164]

[0165] According to the extreme value condition, taking the derivative of equation (14) with respect to Δx and setting it to 0, we get:

[0166] J(x)f(x) + J(x)J(x) T Δx = 0; (18)

[0167] The following system of equations can be obtained:

[0168] J(x)J(x) T Δx = -J(x)f(x); (19)

[0169] Equation (19) is a linear system of equations about Δx and is the increment equation.

[0170] The steps to solve the increment equation are as follows:

[0171] a. Given the initial value x0;

[0172] b. For the k-th iteration, find the current Jacobian matrix J(x k ) and the error f(x k );

[0173] c. Solve the increment equation;

[0174] d. If Δx k is less than the preset value, stop the iteration; otherwise, let x k+1 = x k + Δx k , and return to step b.

[0175] After calculating the optimal pose parameters of the camera, adjust the pose of the camera according to the optimal pose parameters of the camera.

[0176] Further, in step S48, by continuously adding new feature points and camera poses, a three-dimensional environmental map can be gradually constructed.

[0177] Further, referring to Figure 5 , perform plane fitting based on the three-dimensional environmental map, establish a plane equation, and solve the plane equation to obtain the elevation information of the area to be detected, including:

[0178] S51. Construct a design matrix and an observation vector for each three-dimensional point;

[0179] S52. Establish a system of linear equations according to the design matrix and the observation vector;

[0180] S53. Solve the system of linear equations by the least squares method to obtain plane parameters;

[0181] S54. Construct a plane equation according to the plane parameters, and obtain the elevation information of the area to be detected according to the plane equation.

[0182] Specifically, in step S51, for each three-dimensional point (X i , Y i , Z i ), construct a design matrix Q. Each row of the design matrix corresponds to the two-dimensional coordinates of a three-dimensional point, in the form of (X i , Y i , 1). At the same time, construct an observation vector k, and each element is the z coordinate of this point.

[0183] Further, in step S52, the system of linear equations established according to the design matrix and the observation vector is Qx = k, where x = [a, b, c, d] T are the plane parameters to be solved.

[0184] In steps S53 and S54, the parameters a, b, c, d obtained by solving by the least squares method can be used to construct a plane equation ax + by + cz + d = 0, where a and b are the normal vectors of the plane, c is the intercept of the plane, and d is the distance from the plane to the origin. According to the plane equation obtained above, the elevation information of the ground can be estimated.

[0185] The elevation information is calculated by the following formula:

[0186]

[0187] where Z is the elevation of the area to be detected, Z0 represents the elevation of the reference plane, a, b, c, d are plane parameters, and x, y are the coordinates of points on the plane of the area to be detected.

[0188] Further, in some embodiments, the elevation information obtained by the above method may be compared with the actually collected elevation data to evaluate the confidence level of the elevation information obtained by the above method. Therefore, the method provided in the embodiments further includes:

[0189] S6. Collect the actual elevation data of the ground control points;

[0190] S7. Align the actual elevation data and the elevation information spatially, calculate the root mean square error between the actual elevation data and the elevation information, and obtain the confidence level of the elevation information according to the root mean square error.

[0191] Specifically, the calculation formula for the root mean square error RMSE is:

[0192]

[0193] where H i is the i-th elevation data obtained by the above method, H' i is the i-th elevation data obtained by actual measurement, and n is the number of samples.

[0194] The value of the root mean square error is the confidence level. The credible threshold here can be set to 2. That is to say, when the calculated root mean square error is less than or equal to 2, that is, the error between the two is within 2%, it means that the obtained elevation information is credible. If the root mean square error is greater than 2, it is not credible.

[0195] Reference Figure 6 , in some embodiments, a vision-based elevation information acquisition device applied to the above method is provided, including:

[0196] An image acquisition module 201, configured to collect a dynamic tilt environment image sequence of the area to be detected through a camera carried by a drone;

[0197] An image processing module 202, configured to perform coverage area recognition on the tilt environment image sequence, and perform rejection and pixel filling on the recognized coverage area to obtain a processed environment image sequence;

[0198] An image analysis module 203, configured to analyze based on the processed environment image sequence to obtain image motion information;

[0199] A map building module 204, configured to build a three-dimensional environment map according to the image motion information;

[0200] A fitting module 205, configured to perform plane fitting based on the three-dimensional environment map, establish a plane equation, and solve the plane equation to obtain the elevation information of the area to be detected.

[0201] Further, the image processing module 202 identifies the covered area in the sequence of slanted environment images based on a convolutional neural network.

[0202] Further, the image motion information includes the motion speed of each pixel in the sequence of environment images;

[0203] The image analysis module 203 analyzes based on the processed sequence of environment images to obtain image motion information, including:

[0204] Based on the optical flow equation, obtain the relationship between the motion speed of each pixel in the sequence of environment images and the pixel gray value;

[0205] For each pixel, select a neighborhood window, convert the optical flow equation into a system of linear equations within the neighborhood window, and solve the system of linear equations based on the least squares method to obtain the motion speed of each pixel in the sequence of environment images.

[0206] Specifically, assuming that between consecutive frames of the sequence of environment images, the pixel gray value in the image remains unchanged, the interval between two adjacent frames of the image is short, and the motion of adjacent pixels in space is consistent, the mathematical expression of the brightness constancy constraint can be obtained:

[0207] I(x,y,t) = I(x + Δx,y + Δy,t + Δt); (1)

[0208] where I(x,y,t) represents the gray value of the image at the pixel (x,y) at time t, and I(x + Δx,y + Δy,t + Δt) represents the gray value of the pixel (x,y) at the time t + Δt when it moves to (x + Δx,y + Δy).

[0209] Take the partial derivatives of both sides of the brightness constancy constraint with respect to space and time, and use the chain rule for differentiation. Assuming that I(x,y,t) is a continuously differentiable function, then there is:

[0210]

[0211] Perform a Taylor expansion of (x + Δx,y + Δy,t + Δt) at (x,y,t) to obtain:

[0212]

[0213] Substitute the Taylor expansion result into the derivative formula of the brightness constancy constraint to obtain:

[0214]

[0215] and Let \(u\) and \(v\) be the velocities of the pixel in the \(x\)-axis and \(y\)-axis directions respectively. After organizing the obtained equations, the optical flow equation can be derived:

[0216]

[0217] where \((u, v)\) is the velocity of the pixel \((x, y)\) in the \(x\)-axis and \(y\)-axis directions on the image, \(\nabla I\) is the gradient of the image gray value in the \(x\) and \(y\) directions, \(\frac{\partial I}{\partial t}\) is the rate of change of the gray value with respect to time.

[0218] The \((u, v)\) obtained by solving Equation (5) is the velocity of the pixel.

[0219] Furthermore, for each pixel, a neighborhood window is selected. The optical flow equation is converted into a system of linear equations within the neighborhood window, and the system of linear equations is solved based on the least squares method to obtain the motion velocity of each pixel in the environmental image sequence.

[0220] Specifically, let and be denoted as \(I_x\), \(I_y\), and \(I_t\) respectively. If \(n\) pixels within the neighborhood window all satisfy the above equation, then we have:

[0221]

[0222] It can be converted into the following form of a system of linear equations:

[0223]

[0224] Using the least squares method to solve the system of equations (7), we get:

[0225]

[0226] From this, the velocity estimation of the pixel can be obtained to achieve feature tracking.

[0227] Furthermore, the map building module 204 establishes a three-dimensional environmental map based on the image motion information, including:

[0228] Selecting feature points in the environmental image sequence;

[0229] Obtaining the pixel coordinates of the feature points in the environmental image sequence according to the motion velocities of the pixels of the feature points in the environmental image sequence;

[0230] Establishing the epipolar constraint between the feature points in two adjacent frames of environmental images;

[0231] Form a cost function with the epipolar constraints of all feature points, and use the least squares method to optimize the cost function to obtain the optimal fundamental matrix for each epipolar constraint;

[0232] According to the optimal fundamental matrix, map the feature points of two adjacent frames of environmental images to each other's epipolar lines to obtain feature point pairs;

[0233] According to the pixel coordinates of the feature point pairs, use triangulation to calculate the three-dimensional point coordinates corresponding to the feature point pairs;

[0234] Estimate the camera pose based on the projection of the three-dimensional point coordinates in the corresponding environmental image, and optimize the projection error;

[0235] Construct a three-dimensional environmental map based on the optimized three-dimensional point coordinates of multiple feature points obtained by calculation.

[0236] Specifically, the feature points can be prominent features in the environment to be detected, such as the corners, edges, windows of buildings, etc. They have unique geometric features, which are helpful for subsequent positioning and mapping.

[0237] Furthermore, based on the pixel motion speed calculated by the above method, the pixel coordinates of the feature points in the environmental image sequence can be obtained.

[0238] Furthermore, use the principle of multi-view geometry for three-dimensional reconstruction, and use epipolar constraints to describe the geometric relationship between two images to recover the camera motion and the scene.

[0239] Epipolar constraint refers to the positional relationship between the feature points in these views and the camera when observing the same scene point in two views. For the corresponding pixel point pairs (x, x') in two images, their epipolar constraint is expressed as:

[0240] x' T Fx = 0; (9)

[0241] where F is the fundamental matrix, and x, x' represent the corresponding feature points in the two views respectively.

[0242] Furthermore, in step S44, combine the epipolar constraint equations of all feature point pairs into a cost function, and use the least squares method to minimize this cost function:

[0243]

[0244] where F is the fundamental matrix to be estimated, and ∑ i denotes the sum over all feature point pairs. Optimizing the above cost function can obtain the optimal fundamental matrix F, and i is the number of the feature point.

[0245] Among them, the least squares method is used to optimize the cost function to obtain the optimal fundamental matrix for each epipolar constraint, including:

[0246] Define the minimized residual r:

[0247] r = F * X - c; (11)

[0248] Among them, the fundamental matrix F to be estimated is an N * M matrix, X is an M * 1 unknown vector, and c is an N * 1 known vector, which can be feature points;

[0249] By finding the least squares solution x LS , such that the norm ||r|| of r 2 is minimized, and the least squares solution x LS can be obtained through the following formula:

[0250] x LS = (F T F) -1 F T c; (12)

[0251] Furthermore, after obtaining the optimal fundamental matrix F, the feature points in the images are corresponded to the epipolar lines between each other. That is, for the feature point x i in the first image, the corresponding point x' i in the second image must be located on its corresponding epipolar line.

[0252] Furthermore, for each pair of known feature point pairs (x i , x' i ), the corresponding 3D point coordinates can be calculated by the method of triangulation.

[0253] Specifically, in two adjacent views, the optical center point of the camera and the feature points form a straight line, and the intersection point of the two straight lines formed by the two feature points and the optical center point of the camera respectively can be regarded as the actual spatial point. By determining the two straight lines and the connection line of the two feature points to form a triangle, the 3D coordinates of the corresponding spatial point can be calculated according to the pixel coordinates of the two feature points, that is, the 3D point coordinates corresponding to the feature points.

[0254] Furthermore, the feature points corresponding to the obtained 3D points can be regarded as the projections of the 3D under different camera poses in the two views. However, there will be a certain deviation between the feature points and the actual projection points. Therefore, by adjusting the camera pose to minimize the distance between the feature points and the actual projection points, specifically including:

[0255] Establish a camera model according to the 3D point coordinates and the pixel coordinates of the corresponding feature points;

[0256] According to the camera model, an error model regarding the pixel coordinates of feature points and the three-dimensional point coordinates is established;

[0257] Optimize the error model, find the optimal pose parameters of the camera that minimize the result of the error model, and adjust the pose of the camera according to the calculated optimal pose parameters of the camera.

[0258] Specifically, assume n three-dimensional points, and the coordinates of the i-th three-dimensional point are:

[0259] P i =[X i ,Y i ,Z i T , and the pixel coordinates of its projection, that is, the pixel coordinates of the feature point, are p(p xi ,p yi ). The internal parameters of the camera are K, the pose parameters are T, and the camera model is:

[0260]

[0261] where s i represents the depth of the feature point in the camera coordinate system.

[0262] Furthermore, according to this camera model, an error model regarding the pixel coordinates of feature points and the three-dimensional point coordinates is established. In Equation (13), the left side represents the projection position of the feature point, and the right side represents the observed position of the feature point. Written in matrix form as:

[0263] s i p = KTP i ; (14)

[0264] Due to the camera pose and the noise of the observed points, there is an error between the left and right sides of the above equation. Therefore, sum the errors to construct a least squares problem and obtain the error model e:

[0265]

[0266] Furthermore, in step S47c, optimize the above error model, that is, find the optimal pose parameters T of the camera to minimize the result of this error model. Specifically, the iterative formula can be expressed as:

[0267] f(x + Δx) ≈ f(x) + J(x) T Δx; (16)

[0268] where J(x) T is the derivative of f(x) with respect to x and is a Jacobian matrix.

[0269] The current goal is to find the increment Δx such that ||f(x + Δx)||​2 Reaches the minimum. To solve for Δx, a linear least squares problem needs to be solved:

[0270]

[0271] According to the extreme value condition, take the derivative of Equation (14) with respect to Δx and set it to 0, obtaining:

[0272] J(x)f(x)+J(x)J(x) T Δx = 0; (18)

[0273] The following system of equations can be obtained:

[0274] J(x)J(x) T Δx = -J(x)f(x); (19)

[0275] Equation (17) is a linear system of equations with respect to Δx and is the increment equation.

[0276] The steps to solve the increment equation are as follows:

[0277] a. Given the initial value x0;

[0278] b. For the k-th iteration, find the current Jacobian matrix J(x k ) and the error f(x k );

[0279] c. Solve the increment equation;

[0280] d. If Δx k is less than the preset value, stop the iteration; otherwise, let x k+1 = x k +Δx k , and return to step b.

[0281] After calculating the optimal pose parameters of the camera, adjust the pose of the camera according to the optimal pose parameters of the camera.

[0282] Furthermore, by continuously adding new feature points and camera poses, a three-dimensional environmental map can be gradually constructed.

[0283] Furthermore, the fitting module 205 performs plane fitting based on the three-dimensional environmental map, establishes a plane equation, and solves the plane equation to obtain the elevation information of the area to be detected, including:

[0284] Construct a design matrix and an observation vector for each three-dimensional point;

[0285] Establish a system of linear equations according to the design matrix and the observation vector;

[0286] Solve the system of linear equations using the least squares method to obtain the plane parameters;

[0287] Construct a plane equation according to the plane parameters, and obtain the elevation information of the area to be detected according to the plane equation.

[0288] Specifically, for each three-dimensional point (X i , Y i , Z i ), construct a design matrix Q. Each row of the design matrix corresponds to the two-dimensional coordinates of a three-dimensional point, in the form of (X i , Y i , 1). At the same time, construct an observation vector k, and each element is the z coordinate of this point.

[0289] Furthermore, the linear equation system established according to the design matrix and the observation vector is Qx = k, where x = [a, b, c, d] T is the plane parameter to be solved.

[0290] The parameters a, b, c, d obtained by using the least squares method can be used to construct the plane equation ax + by + cz + d = 0, where a and b are the normal vectors of the plane, c is the intercept of the plane, and d is the distance from the plane to the origin. According to the plane equation obtained above, the elevation information of the ground can be estimated.

[0291] The elevation information is calculated by the following formula:

[0292]

[0293] where Z is the elevation of the area to be detected, Z0 represents the elevation of the reference plane, a, b, c, d are plane parameters, and x, y are the coordinates of points on the plane of the area to be detected.

[0294] Furthermore, the device further includes an evaluation module for:

[0295] Collect the true elevation data of the ground control points;

[0296] Align the true elevation data with the elevation information in space, calculate the root mean square error between the true elevation data and the elevation information, and obtain the confidence level of the elevation information according to the root mean square error.

[0297] Specifically, the calculation formula of the root mean square error RMSE is:

[0298]

[0299] where H i is the i-th elevation data obtained by the above method, H′ i is the i-th elevation data obtained by actual measurement, and n is the number of samples.

[0300] The value of the root mean square error is the confidence level. The credible threshold here can be set to 2. That is to say, when the calculated root mean square error is less than or equal to 2, that is, there is an error within 2% between the two, it means that the obtained elevation information is credible. If the root mean square error is greater than 2, it is not credible.

[0301] In some embodiments, a vision-based elevation information acquisition system is further provided, including a processor, a storage device, and a drone. A camera is mounted on the drone. The storage device stores multiple instructions, and the processor is configured to read the multiple instructions and execute the above method.

[0302] The vision-based elevation information acquisition method, device, and system provided by the above embodiments have at least the following beneficial effects:

[0303] (1) Constructing a three-dimensional environmental map based on a dynamic sequence of inclined environment images can effectively reduce the amount of data calculation compared to laser point clouds and is not affected by environmental climate, resulting in relatively high accuracy of the obtained elevation information;

[0304] (2) Removing vegetation from the obtained sequence of inclined environment images to reduce the impact of vegetation on modeling, thereby improving the accuracy of elevation information;

[0305] (3) Continuously optimizing the camera pose during the modeling process to improve the accuracy of modeling, and thus improve the accuracy of elevation information.

[0306] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for obtaining elevation information based on vision, characterized in that: include: The camera carried by the drone is used to collect a dynamic sequence of oblique environmental images of the area to be inspected; Identifying the coverage area of ​​the inclined environment image sequence, and removing and filling pixels of the identified coverage area to obtain a processed environment image sequence; Analyze the processed environment image sequence to obtain image motion information; Establishing a three-dimensional environment map according to the image motion information; Performing plane fitting based on the three-dimensional environment map, establishing a plane equation, and solving the plane equation to obtain elevation information of the area to be detected; The image motion information includes the motion speed of each pixel in the environment image sequence; Based on the analysis of the processed environmental image sequence, image motion information is obtained, including: Based on the optical flow equation, the relationship between the motion speed and the pixel gray value of each pixel in the environment image sequence is obtained; Selecting a neighborhood window for each pixel, converting the optical flow equation into a linear equation group within the neighborhood window, solving the linear equation group based on the least squares method, and obtaining the movement speed of each pixel in the environment image sequence; According to the image motion information, a three-dimensional environment map is established, including: In the environment image sequence, selecting feature points; Obtaining pixel coordinates of the feature point in the environment image sequence according to the movement speed of the pixel of the feature point in the environment image sequence; Establish epipolar constraints of feature points in two adjacent frames of environmental images; The epipolar constraints of all feature points are combined into a cost function, and the cost function is optimized using the least squares method to obtain the optimal basic matrix of each epipolar constraint; According to the optimal basic matrix, feature points of two adjacent frames of environment images are matched to each other's epipolar lines to obtain feature point pairs; Calculating the three-dimensional point coordinates corresponding to the feature point pair using triangulation according to the pixel coordinates of the feature point pair; Estimate the camera posture according to the projection of the three-dimensional point coordinates in the corresponding environment image, and optimize the projection error; A three-dimensional environment map is constructed based on the calculated coordinates of multiple three-dimensional points.

2. The method according to claim 1, characterized in that The coverage area in the tilted environment image sequence is identified based on a convolutional neural network.

3. The method according to claim 1, characterized in that The camera posture is estimated according to the projection of the three-dimensional point coordinates in the corresponding environment image, and the projection error is optimized, including: A camera model is established according to the three-dimensional point coordinates and the pixel coordinates of the corresponding feature points; According to the camera model, an error model about the pixel coordinates of the feature points and the three-dimensional point coordinates is established; The error model is optimized to find the optimal camera posture parameters that minimize the error model result, and the camera posture is adjusted according to the calculated optimal camera posture parameters.

4. The method according to claim 1, characterized in that: Performing plane fitting based on the three-dimensional environment map, establishing a plane equation, and solving the plane equation to obtain elevation information of the area to be detected includes: Construct the design matrix and observation vector for each 3D point; Establishing a linear equation system according to the design matrix and the observation vector; Solving the linear equations by the least square method to obtain the plane parameters; A plane equation is constructed according to the plane parameters, and elevation information of the area to be detected is obtained according to the reference plane elevation and the plane equation.

5. The method according to claim 4, characterized in that The elevation information is calculated using the following formula: Among them, Z is the elevation of the area to be detected, Z0 represents the elevation of the reference plane, a, b, c, d are plane parameters, and x and y are the coordinates of the point on the plane of the area to be detected.

6. The method according to claim 1, characterized in that The method further comprises: Collect true elevation data of ground control points; The true elevation data and the elevation information are spatially aligned, and a root mean square error between the true elevation data and the elevation information is calculated, and the confidence of the elevation information is obtained according to the root mean square error.

7. A visual-based elevation information acquisition device applied to any one of claims 1 to 6, characterized in that: include: An image acquisition module is used to collect a dynamic tilted environment image sequence of the area to be detected through a camera carried by the drone; An image processing module is used to identify the coverage area of ​​the inclined environment image sequence, and to remove and fill pixels of the identified coverage area to obtain a processed environment image sequence; An image analysis module, used for analyzing the processed environment image sequence to obtain image motion information; A map building module, used to build a three-dimensional environment map according to the image motion information; The fitting module is used to perform plane fitting based on the three-dimensional environment map, establish a plane equation, and solve the plane equation to obtain the elevation information of the area to be detected.

8. A vision-based elevation information acquisition system, characterized in that: It comprises a processor, a storage device and a drone, wherein the drone is equipped with a camera, the storage device stores a plurality of instructions, and the processor is used to read the plurality of instructions and execute any method according to claims 1-6.

Citation Information

Patent Citations

  • Remote sensing image feature point elevation acquisition method based on multiple sensors and sea level

    CN113723568A

  • Underwater sonar image matching method based on Gaussian distribution clustering

    CN113313172A

  • High-precision positioning method based on oblique photography of unmanned aerial vehicle

    CN113340277A