Three-Dimensional Map Reconstruction Method for Large-Scale Scenes Incorporating Multi-Modal Data Sensing
By registering and pose correction of camera and radar point clouds, a more accurate three-dimensional map model is generated, which solves the problem that the camera and radar have difficulty in combining their respective advantages in three-dimensional map reconstruction, and achieves higher-precision three-dimensional map reconstruction.
Patent Information
- Application Number
- CN202510571975.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In the prior art, the three-dimensional point cloud generated by the camera has a wide range but low accuracy, while the three-dimensional point cloud generated by the radar has a high accuracy but small range. How to make full use of the advantages of the two to obtain a more accurate three-dimensional map model is a difficult problem.
The point cloud generated by the camera and radar are registered through the registration algorithm, the optimal rotation matrix and translation vector are obtained, the camera's pose is corrected, and the camera's point cloud is corrected using the corrected pose, and a three-dimensional map model is generated by combining the surface reconstruction algorithm.
Improve the correction accuracy of the camera's long-distance point cloud, obtain a more accurate three-dimensional map model, and make full use of the respective advantages of the camera and radar.
Smart Images

Figure CN120088423B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and specifically relates to a method for reconstructing a three-dimensional map of a large-scale scene space that integrates multi-modal data perception. Background Art
[0002] In the field of low-altitude economy, through three-dimensional map reconstruction, accurate navigation information can be provided for low-altitude aircraft, including airspace routes, takeoff and landing point positions, etc., thereby improving flight safety and efficiency. At the same time, the three-dimensional map can also be used for flight supervision to monitor the position and status of the aircraft in real time to ensure that flight activities comply with regulations. In urban planning, the three-dimensional map can provide detailed spatial information to help planners better understand the urban terrain, building layout, etc., so as to make more reasonable planning decisions. In addition, the three-dimensional map can also be used for the simulation and prediction of urban construction to evaluate the impact of different construction plans on the urban environment and traffic. Even when natural disasters or emergencies occur, the three-dimensional map can quickly provide information such as the terrain and building distribution in the disaster area to provide decision-making support for emergency rescue. At the same time, by comparing the three-dimensional maps before and after the disaster, the losses and impacts caused by the disaster can be evaluated to provide a scientific basis for post-disaster reconstruction.
[0003] Current three-dimensional reconstruction data comes from video data of cameras (RGB and IR) and point cloud data based on radar. Specifically, when performing three-dimensional reconstruction based on camera image data, its image data can be reconstructed into three-dimensional point cloud data through three-dimensional coordinate point recovery operations, and then a surface reconstruction operation can be performed on the three-dimensional point cloud data according to a preset surface reconstruction algorithm to determine the corresponding three-dimensional model. Or directly convert the point cloud data of the radar into a three-dimensional model based on a surface reconstruction algorithm or the like. In the field of low-altitude economy, the shooting range of cameras is often far, reaching hundreds of meters or even several kilometers, while the point cloud generated by the radar is limited by cost within 200m. Usually, the range of the point cloud generated by the camera is wider than that generated by the radar, but the accuracy of the point cloud generated by the radar is higher than that of the point cloud generated by the camera. How to make full use of the advantages of both to obtain a more accurate three-dimensional map model is a technical problem that technicians in this field have always had to solve. Summary of the Invention
[0004] Aiming at the problems in the prior art, this application provides a method for reconstructing a three-dimensional map of a large-scale scene space that integrates multi-modal data perception to improve the accuracy of three-dimensional map modeling in the field of low-altitude economy by utilizing the breadth advantage of camera image data and the accuracy advantage of radar point cloud data.
[0005] To solve at least one of the above problems, this application provides the following technical solutions:
[0006] In a first aspect, the present application provides a method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception, including:
[0007] Obtain a sequence of video frames captured by a camera, and calculate a first point cloud based on adjacent frame videos; and, obtain a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and have the same orientation;
[0008] Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose;
[0009] Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud;
[0010] Convert the first corrected point clouds in multiple corrected camera coordinate systems to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
[0011] Further, the step of calculating the first point cloud based on adjacent frame videos includes:
[0012] Match feature points of the same target in adjacent frame videos to obtain feature point pairs of each same target;
[0013] Construct a fundamental matrix based on the feature point pairs, and calculate an essential matrix according to the fundamental matrix and the internal parameters of the camera;
[0014] Decompose the essential matrix to obtain a first rotation matrix and a first translation vector between two adjacent frame cameras;
[0015] Determine the three-dimensional coordinates of each feature point through triangulation according to the internal parameters of the camera and the first rotation matrix and the first translation vector; the three-dimensional coordinates of the feature points of multiple targets form the first point cloud.
[0016] Further, the step that the three-dimensional coordinates of the feature points of multiple targets form the first point cloud includes:
[0017] Take the center point of the camera rectangular screen range as the origin, the X% of the rectangle length as the major axis, and the X% of the rectangle width as the minor axis to draw an ellipse, and select the feature points located inside the intersection area of the rectangle and the ellipse for feature point matching; the X% is 5% - 15%.
[0018] Further, the step of registering the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector includes:
[0019] For each point in the first point cloud, find the point in the second point cloud that is closest to it in distance through the nearest neighbor search algorithm to form a matching point pair;
[0020] Based on the principle of rigid transformation, calculate a second rotation matrix and a second translation vector that can align all the matching point pairs;
[0021] Use the second rotation matrix and the second translation vector to transform the points in the first point cloud to make them closer to the point in the first point cloud that is closest to it in distance;
[0022] Repeat the above steps until the number of iterations reaches the upper limit, or the distance change between the first point cloud and the second point cloud is less than a preset value. The obtained second rotation matrix and second translation vector are the optimal rotation matrix and the optimal translation vector.
[0023] Further, the step of calculating a second rotation matrix and a second translation vector that can align all the matching point pairs based on the principle of rigid transformation includes:
[0024] In each iteration, randomly select some matching point pairs and calculate the transformation matrix for aligning these matching point pairs. The transformation matrix includes the second rotation matrix and the second translation vector;
[0025] Apply the transformation matrix to the first point cloud and calculate the distance between the transformed first point cloud and other matching point pairs in the second point cloud;
[0026] Regard the matching point pairs whose distances meet the set threshold as inliers and retain them, and regard other non-conforming matching points as outliers and remove them.
[0027] Further, the step of randomly selecting some matching point pairs and calculating the transformation matrix for aligning these matching point pairs in each iteration includes:
[0028] Take the camera parameters, the points in the first point cloud, and the points in the second point cloud as variable nodes, and construct a factor graph with the reprojection error between the corresponding point coordinates in the first point cloud and the second point cloud as the observation edge; the camera parameters include internal parameters and external parameters; the external parameters include the second rotation matrix and the second translation vector;
[0029] Perform graph optimization on the factor graph to obtain the transformation matrix; the factor graph uses sparse linear algebra and a solver for graph optimization.
[0030] Further, the method for obtaining the initial pose is as follows: Based on the known camera pose, the initial pose of the camera for capturing each frame of the video is calculated by accumulating relative poses frame by frame.
[0031] In a second aspect, the present application provides a large-scale scene space three-dimensional map reconstruction device that fuses multi-modal data perception, including:
[0032] A point cloud generation module, configured to obtain a sequence of video frames captured by a camera, and calculate a first point cloud according to adjacent video frames; and obtain a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and have the same orientation;
[0033] A pose correction module, configured to register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose;
[0034] A point cloud correction module, configured to convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud;
[0035] A map construction module, configured to convert multiple first corrected point clouds in the corrected camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
[0036] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for reconstructing a large-scale scene space three-dimensional map that fuses multi-modal data perception are implemented.
[0037] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for reconstructing a large-scale scene space three-dimensional map that fuses multi-modal data perception are implemented.
[0038] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method for reconstructing a large-scale scene space three-dimensional map that fuses multi-modal data perception are implemented.
[0039] As can be seen from the above technical solutions, the present application provides a method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception. A first point cloud is generated from the images captured by a camera, and a second point cloud is generated by a radar. By taking advantage of the more accurate radar data, the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera is obtained, and the pose of the camera is corrected by this deviation value. Then, the three-dimensional points of the first point cloud are transformed to the corrected camera coordinate system through coordinate transformation to obtain a first corrected point cloud. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a schematic flowchart of the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception in the embodiments of the present application;
[0042] Figure 2 It is a regional diagram of the point selection range of the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception in the embodiments of the present application;
[0043] Figure 3 It is a structural diagram of the device for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception in the embodiments of the present application;
[0044] Figure 4 It is a schematic structural diagram of the electronic device in the embodiments of the present application.
[0045] REFERENCE SIGNS:
[0046] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0048] In the technical solutions of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0049] Considering the problems existing in the prior art, this application provides a method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception. A first point cloud is generated from the images captured by a camera, and a second point cloud is generated by a radar. By taking advantage of the more accurate radar data, the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera is obtained, and the pose of the camera is corrected by this deviation value. Then, the three-dimensional points of the first point cloud are transformed to the corrected camera coordinate system through coordinate transformation to obtain a first corrected point cloud. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.
[0050] To obtain a more accurate three-dimensional map modeling in a larger range, this application makes full use of the respective advantages of the camera and the radar, and provides an embodiment of a method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception. Refer to Figure 1 , the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception specifically includes the following content:
[0051] Step S101: Obtain the video frame sequence captured by the camera, and calculate the first point cloud based on adjacent frame videos; and obtain the second point cloud generated by the radar; the camera and the radar are arranged adjacent to each other and have the same orientation.
[0052] Optionally, in this embodiment, the camera and the radar are arranged on the belly of a low-altitude aircraft, and the two are installed very close to each other in the physical space, or even share the same installation position or area. Therefore, the second point cloud can be used to correct the first point cloud, and the obtained deviation value can be used to reversely correct the absolute pose of the camera. The point cloud data of the camera is obtained again using the corrected pose, and finally, the first corrected point cloud data with higher accuracy than the first point cloud data is obtained.
[0053] Optionally, in this embodiment, the steps of calculating the first point cloud based on adjacent frame videos include: using image processing algorithms such as SIFT, SURF, and ORB to extract feature points in adjacent frame videos, and then using the nearest neighbor algorithm to perform feature point matching on the same targets in adjacent frame videos to obtain feature point pairs of each same target; constructing a fundamental matrix based on the feature point pairs, and calculating an essential matrix according to the fundamental matrix and the internal parameters of the camera (focal length, principal point coordinates, distortion coefficients, etc.); by decomposing the essential matrix, obtaining a first rotation matrix and a first translation vector between two adjacent frame cameras, that is, the relative pose of two adjacent frame cameras; finally, according to the internal parameters of the camera and the first rotation matrix and the first translation vector, determining the three-dimensional coordinates of each feature point through triangulation; the three-dimensional coordinates of the feature points of multiple targets constitute the first point cloud.
[0054] Specifically, the definition of the fundamental matrix is:
[0055] F = K -T EK -1
[0056] where F is the fundamental matrix, E is the essential matrix, and K is the internal parameter matrix of the camera.
[0057] The relationship between the essential matrix E and the relative pose of the camera (rotation matrix R and translation vector t) is:
[0058] E = [t] × R
[0059] where [t] × is the skew-symmetric matrix of the translation vector t.
[0060] The calculation formula for the three-dimensional coordinates of the feature points is:
[0061] P = K -1 p1 × (RK -1 p2 + t)
[0062] where p1 and p2 are feature points in the camera coordinate systems of two adjacent frames, K is the internal parameter matrix of the camera, (R, t) is the relative pose of the camera, and × is the cross product of vectors.
[0063] Optionally, in this embodiment, considering that the usage scenario of the present invention is a large spatial scale, small errors at the camera pose end will be amplified at the far end of the picture. Although distortion correction of the picture can restore the picture, there is still a distortion problem at the edge of the restored picture. Therefore, to eliminate the influence of lens distortion on the final result, in this embodiment, when selecting feature points, feature points in the middle area of the picture are selected, edge feature points are ignored, and at the same time, a part of the edge picture is discarded, and only most of the middle picture is used to minimize the influence of lens distortion on calculating the camera pose.
[0064] The selection point range area mask is as follows Figure 2 shown. Taking the center point of the camera's rectangular screen range as the origin, with X% of the rectangle's length as the major axis and X% of the rectangle's width as the minor axis, an ellipse is drawn. Feature points located within the intersection area of the rectangle and the ellipse are selected for feature point matching, and the remaining feature points are discarded. The rectangle is the original screen range. Exemplarily, if the camera is 16:9, the resolution is 1920 × 1080. X is determined according to the specific distortion situation, and the range is 5% - 15%. For example, 5%, 10%, 15%, etc. are selected.
[0065] Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain the optimal rotation matrix and the optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose.
[0066] Optionally, in this embodiment, the ICP (Iterative Closest Point) algorithm is used to register the first point cloud and the second point cloud. That is, taking the second point cloud of the radar as the target point cloud and the first point cloud of the camera as the source point cloud, minimizing the error function E between the first point cloud and the second point cloud:
[0067]
[0068] where P i and Q i are the matching points in the first point cloud and the second point cloud respectively, and R and t are the rotation matrix and translation vector of the camera respectively.
[0069] That is, during the point cloud registration process, the ICP algorithm iteratively finds the closest point pairs between the two point clouds (the first point cloud and the second point cloud) and calculates the optimal rigid transformation (including the rotation matrix R and the translation vector t) so that the source point cloud (the first point cloud) can be aligned in the coordinate system of the target point cloud (the second point cloud). If each point in the first point cloud is transformed according to the rotation matrix R and the translation vector t, a corrected point cloud can be obtained. This corrected point cloud is closer to the second point cloud generated by the radar in space, which can improve the accuracy of the first point cloud. Then, since the rotation matrix R and the translation vector t are the transformation matrices for transforming the first point cloud to the second point cloud, then to correct the pose of the camera in the reverse direction, the inverse matrix of R (or the transpose matrix, because R is an orthogonal matrix and its inverse matrix is equal to the transpose matrix) and -t (i.e., the opposite of the translation vector) are used to update the pose of the camera so that it can be aligned with the corrected point cloud. Finally, using the corrected pose to obtain the point cloud data of the camera again, the first corrected point cloud data with a higher accuracy than the first point cloud data is finally obtained.
[0070] Specifically, in this embodiment, the registration step includes: for each point in the first point cloud, finding the point in the second point cloud that is closest to it in distance through the nearest neighbor search algorithm to form a matching point pair; based on the principle of rigid transformation, calculating the second rotation matrix and the second translation vector that can align all the matching point pairs; using the second rotation matrix and the second translation vector to transform the points in the first point cloud to make them closer to the point in the first point cloud that is closest to it in distance; repeating the above steps until the number of iterations reaches the upper limit, or the distance change between the first point cloud and the second point cloud is less than the preset value, and the obtained second rotation matrix and second translation vector are the optimal rotation matrix and the optimal translation vector.
[0071] Optionally, in this embodiment, considering that the distortion effect in the central part of the camera's picture is the smallest, and the closer the radar is to the central area, the higher its accuracy, when performing ICP matching, a weight ratio can be increased for these central areas. The specific weight increase scheme is as follows: in order to make the weight of the middle area greater than that of the edge area, an appropriate σ value can be set so that the value of the Gaussian function is larger in the central area and smaller in the edge area, and these weights are used to adjust the point cloud data so that the points in the central area have a greater influence during the registration process. The specific formula is as follows:
[0072]
[0073] where d represents the distance from the point to a certain reference point (the center point of the picture in this example), and σ is the standard deviation, which is used to control the width of the Gaussian distribution.
[0074] In point cloud registration, the closer the point is to the central area (i.e., the smaller the d value), the greater its weight; the farther the point is from the central area, the smaller its weight. If the standard deviation σ is larger, the Gaussian distribution is wider, indicating a greater degree of dispersion of the data points; if the standard deviation σ is smaller, the Gaussian distribution is narrower, indicating that the data points are more concentrated.
[0075] Specifically in this example, first check the accuracy of the radar data. If the measurement results show that the deviation between the middle data error and the edge data error is greater, that is, the central data in the radar data is more accurate, then a smaller value should be selected for d. Exemplarily, in this example, the d value can be selected as 1.
[0076] Optionally, in this embodiment, considering that the point cloud registration is multi-point matching, the quality of the matching points is crucial for the registration result. If the matching point pairs are inaccurate, the registration result will be poor. Therefore, this embodiment introduces the RANSAC algorithm (RANdom Sample Consensus) to optimize the registration. That is, using the RANSAC algorithm, by presetting a reasonable threshold to distinguish normal points (inliers) and abnormal points (outliers), the correct transformation parameters (optimal rotation matrix and optimal translation vector) are screened out from the matching points to reduce the influence of outliers on the registration and improve the robustness and accuracy of the registration.
[0077] Specifically, in this embodiment, the step of calculating the second rotation matrix and the second translation vector that can align all matching point pairs based on the principle of rigid transformation includes: in each iteration, randomly select a part of the matching point pairs, preferably the minimum number of matching point pairs, such as 3 matching point pairs, and calculate the transformation matrix that aligns these matching point pairs. This transformation matrix includes the second rotation matrix and the second translation vector; apply the transformation matrix to the first point cloud and calculate the distance between the transformed first point cloud and other matching point pairs in the second point cloud to evaluate the quality of the transformation; regard the matching point pairs whose distances meet the set threshold as inliers and retain them, and regard other non-conforming matching points as outliers and remove them. The RANSAC algorithm will continue to iterate and update the optimal transformation matrix until a transformation matrix that can maximize the matching of inliers is found.
[0078] Optionally, since in the RANSAC algorithm, it is usually necessary to find the optimal model in a large amount of data. Therefore, to improve the operation speed, this embodiment uses a factor graph to process these constraints in a more structured and efficient way. That is, in the RANSAC algorithm of this embodiment, a model is estimated by randomly selecting samples from the matching points, and it is checked which point pairs conform to the model (i.e., inliers). By introducing the factor graph into the RANSAC algorithm, this process can be accelerated by modeling constraints and optimization problems, especially in the case of applying to large-scale data.
[0079] A factor graph is a graphical model that consists of variable nodes and factor nodes, and the factor nodes represent the constraints between variables. Usually, it is used to express the relationship between different variables in an optimization problem. Therefore, in the RANSAC algorithm of this embodiment, the factor graph can structurally represent the transformation (such as rotation and translation) between point clouds and find the optimal transformation through optimization, so that this embodiment can efficiently find the optimal solution without directly brute-forcing through all possible matching points and improve the search speed of the optimal solution.
[0080] Specifically, in this embodiment, the general framework of introducing a factor graph into the RANSAC algorithm is as follows: Camera parameters (including internal parameters such as focal length and principal point coordinates, and external parameters such as the camera pose described by the second rotation matrix and the second translation vector), points in the first point cloud, and points in the second point cloud are used as variable nodes, and the BA (Bundle Adjustment) algorithm is added as a constraint to the factor graph to further optimize the camera pose and the three-dimensional coordinates of feature points for refining the geometric model obtained from the initial SIFT / SURF / ORB and other feature matches.
[0081] In this embodiment, the reprojection error between the corresponding point coordinates in the first point cloud and the second point cloud is used as an observation edge, or a factor node, to construct a factor graph. Each matching point pair serves as a factor node, connecting two variable nodes (e.g., the transformation matrix corresponding to the first point cloud and the second point cloud, or the transformation parameters). The task of the factor node is to measure the error of the point cloud matching under the current transformation. By minimizing the reprojection error, the positions of the three-dimensional points in the scene and the internal and external parameters of the camera are jointly optimized to obtain the most accurate geometric model and camera pose.
[0082] More specifically, for a known three-dimensional point P and camera parameters, the projected coordinate u of this point on the image can be calculated. If the position of the feature point observed by the radar is u′, the reprojection error is defined as the difference between the two: e = u′ - u. The goal of BA is to find a set of camera parameters and three-dimensional point coordinates that minimize the sum of the reprojection errors of all feature points on all images.
[0083] Therefore, in each iteration of the RANSAC algorithm in this embodiment, the estimated transformation parameters are used as variable nodes in the factor graph, and these transformation parameters are optimized through optimization algorithms such as the Gauss-Newton method and the Levenberg-Marquardt algorithm. The goal of the optimization is to minimize the error of the constraints in all factor nodes to obtain the optimal transformation matrix. Then, based on the current optimized transformation, the error of each matching point pair is calculated to determine whether the matching point is an inlier. If the error is less than a certain threshold, the matching point is considered an inlier, and the factor graph is updated. That is, after each iteration, the transformation parameters obtained by optimizing the factor graph are compared with the current optimal transformation, and the optimal transformation (the transformation with the most inliers) is updated. As a preferred solution, this embodiment uses sparse linear algebra and an efficient solver for graph optimization.
[0084] In the traditional RANSAC algorithm, for each fitting, all three-dimensional points need to be traversed to calculate the inlier or outlier status of each three-dimensional point. Factor graphs can help identify and group three-dimensional points with similar constraints, thereby accelerating the processing of these three-dimensional points. By decomposing these tasks and executing them in parallel, factor graphs can significantly improve the computational efficiency of the RANSAC algorithm. Moreover, through factor graphs, the RANSAC algorithm can more efficiently represent and utilize the relationship between inliers and outliers. This structured representation helps reduce redundant calculations when selecting inliers and outliers. Moreover, factor graphs can be iteratively optimized so that each round of calculations can converge to a reasonable solution faster. Compared with the random selection and repeated trials of the traditional RANSAC algorithm, factor graphs quickly determine which three-dimensional points belong to inliers through a more sophisticated optimization process, thereby accelerating the convergence of the overall algorithm.
[0085] At the same time, the factor graph can not only select inliers and outliers in each RANSAC iteration, but also optimize the quality of the solution through a global optimization method. The factor graph can simultaneously process the constraint relationship of multiple sets of data, and correct the deviation through an optimization algorithm to reduce the influence of outliers on the final model, thereby improving the accuracy of the selection of inliers at the global level. In this way, the system does not need to repeatedly calculate the inliers and outliers of all points every time, while improving the calculation speed. Moreover, in traditional methods, the selection of inliers and outliers usually depends on local distance metrics, which may lead to large errors. In this embodiment, the factor graph helps to more accurately screen out inliers in three-dimensional points by utilizing more complex global constraints, such as pose constraints. This precise screening further reduces unnecessary calculations, especially when the amount of data is large. Therefore, since the factor graph can take into account the global constraints between all points, it is more accurate in the selection of inliers than the traditional RANSAC method.
[0086] Therefore, in this embodiment, through the factor graph, the RANSAC algorithm can use global constraints and optimization methods when selecting the interior points and exterior points of three-dimensional points to improve the accuracy of interior point selection and reduce redundant calculations, thereby significantly speeding up the calculation speed, improving robustness and accelerating the convergence process to adapt to the processing of large-scale point cloud data.
[0087] Optionally, in this embodiment, the method for obtaining the initial pose is: based on the known camera pose, the initial pose of the camera that shoots each frame of video is calculated by accumulating the relative pose frame by frame. That is, the pose of the camera when shooting the video (i.e., the rotation matrix R and translation vector t of each frame of the camera) is calculated by analyzing the relative motion between adjacent frames of the camera. That is, assuming that the pose of the camera in the first frame T0=[R0|t0] is known and is usually set to the unit matrix (i.e., R0=I and t0=0), the goal is to calculate the pose T of the camera in each subsequent frame i =[Ri |t i , where R i is the rotation matrix and t i is the translation vector, such as the first rotation matrix and the first translation vector in the above steps.
[0088] Specifically, through the LK optical flow (Lucas-Kanade optical flow) algorithm or the dense optical flow method, the motion of the camera is estimated using the motion of the pixel points in the image. The rotation matrix R and the translation vector t obtained from the above methods represent the relative pose of the camera. In practical applications, if the initial pose of the camera (e.g., T0) is known, then the formula for calculating the absolute pose of each frame by accumulating the relative pose frame by frame is:
[0089] T i = T i-1 ·T rel
[0090] where, T i-1 is the camera pose of the previous frame, and T rel =[R rel |t rel is the relative pose from the previous frame to the current one. Each time a new camera pose is calculated, the absolute pose of the camera in the current frame is updated.
[0091] Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain the first corrected point cloud.
[0092] When calculating the three-dimensional point cloud, the triangulation method uses the relative position (baseline) between two frames of images and the parallax between corresponding points to calculate the position of each point in three-dimensional space. This process depends on the internal parameters (focal length, principal point coordinates, etc.) and external parameters (the pose of the camera, such as rotation and translation) of the camera to proceed. That is, the position of each three-dimensional point is reconstructed through the pose of the camera and the corresponding points in the image. The pose of the camera defines the conversion from the world coordinate system to the camera coordinate system. Therefore, if a three-dimensional point cloud has been obtained through the triangulation method, then the coordinates of these points are based on the camera coordinate system. If the camera pose is corrected, the three-dimensional point cloud needs to be converted from the old camera coordinate system to the new camera coordinate system.
[0093] Specifically, assume the original three-dimensional point cloud P i =(X i , Y i , Z i) is calculated relative to the initial camera pose \(T_0 = [R_0|t_0]\) (where \(R_0\) is the rotation matrix and \(t_0\) is the translation vector). Now, the radar point cloud corrects the camera pose to obtain the corrected pose \(T_0'=[R_0'|t_0']\). The conversion steps for transforming all 3D point clouds from the original coordinate system to the new coordinate system are as follows:
[0094] 1. Convert the 3D points from the camera coordinate system to the world coordinate system:
[0095] Use the camera pose \(T_0 = [R_0|t_0]\) to transform each 3D point \(P\) i from the camera coordinate system to the world coordinate system:
[0096] \(P\) i world \(=R_0\cdot P\) i \(+t_0\)
[0097] 2. Convert the 3D points from the world coordinate system to the new camera coordinate system:
[0098] Use the new camera pose \(T_0'=[R_0'|t_0']\) to transform the point cloud to the new camera coordinate system:
[0099] \(P\) i camera \('=R_0'\) T \(\cdot(P\) i world \(-t_0')\)
[0100] where \(R_0'\) T is the transpose matrix of \(R_0'\) (representing the rotation from the world coordinate system to the camera coordinate system).
[0101] 3. Combine the transformations:
[0102] Combining these two steps, the 3D point cloud \(P\) in the initial camera coordinate system \(T_0\) i can be calculated to obtain the 3D point cloud in the corrected camera coordinate system \(T_0'\):
[0103] \(P\) i camera \('=R_0'\) T \(\cdot(R_0\cdot P\) i \(+t_0 - t_0')\)
[0104] Here, \(R_0\cdot P\) i \(+t_0\) is the result of converting the 3D points from the camera coordinate system to the world coordinate system. Subtracting \(t_0'\) and multiplying by \(R_0'\) T completes the conversion of the 3D points from the world coordinate system to the corrected camera coordinate system.
[0105] When using a camera and a radar to detect the low-altitude economic field, the point clouds formed by the camera and the radar only partially overlap within the fan-shaped radiation range near the main point of the device. If the point clouds are directly corrected using a transformation matrix after point cloud registration, as the distance increases, the correction error of the camera point cloud will become larger and larger, making it difficult to supplement the data range of the radar point cloud. Therefore, compared with traditional methods, the method of first correcting the camera pose using the radar point cloud and then correcting the camera point cloud using the corrected camera pose can greatly improve the correction accuracy of the camera points at a long distance, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.
[0106] Step S104: Convert the first corrected point clouds in multiple corrected camera coordinate systems into the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
[0107] Since multiple consecutive three-dimensional point clouds calculated by the triangulation method are respectively located in the three-dimensional coordinate systems of multiple camera frames, after calculating multiple consecutive three-dimensional point clouds using the triangulation method, it is necessary to align the three-dimensional coordinate systems of multiple camera frames into the same world coordinate system through point cloud fusion and merge multiple point clouds into a unified three-dimensional model.
[0108] The specific fusion method is as follows:
[0109] Coordinate transformation: Perform coordinate transformation on each new point cloud (for example, the point cloud calculated in frame k), and convert it from the camera coordinate system of this frame to the world coordinate system. Referring to the above calculation process of the initial camera pose, obtain the rotation matrix R k and translation vector t k of the camera k relative to the reference frame (usually frame 0), then the formula for converting the point P k from the camera k coordinate system to the world coordinate system P0 of the reference frame is:
[0110] P0 = R k ·P k + t k
[0111] where R k is the coordinate of a certain three-dimensional point in the point cloud in the camera k coordinate system, and P0 is the coordinate of this point in the world coordinate system of the reference frame.
[0112] Point cloud registration: Use point cloud registration algorithms such as the ICP algorithm or the NDT algorithm to align the point clouds, and fuse the point clouds through stitching to obtain a complete and fused three-dimensional point cloud. Preferably, post-processing methods such as voxel grid filtering can also be used to reduce the density and noise of the point cloud.
[0113] Meshing: Perform meshing operations on the fused point cloud to obtain a three-dimensional grid model; map the surface information in the image onto the three-dimensional grid to form a three-dimensional map model with texture.
[0114] As can be seen from the above description, the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception provided by the embodiments of the present application generates a first point cloud from the images captured by the camera, generates a second point cloud from the radar, and utilizes the more accurate advantage of the radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera by this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation, a first corrected point cloud is obtained. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.
[0115] In order to obtain more accurate three-dimensional map modeling in a larger range, the present application makes full use of the respective advantages of the camera and the radar, and provides an embodiment of a large-scale scene space three-dimensional map reconstruction device for implementing all or part of the content of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception. See Figure 3 The large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception specifically includes the following content:
[0116] Point cloud generation module 10, configured to obtain a sequence of video frames captured by the camera, and calculate a first point cloud based on adjacent frame videos; and obtain a second point cloud generated by the radar; the camera and the radar are arranged adjacent to each other and have the same orientation;
[0117] Pose correction module 20, configured to register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose;
[0118] The point cloud correction module 30 is configured to convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose, so as to obtain a first corrected point cloud;
[0119] The map construction module 40 is configured to convert the first corrected point clouds in multiple corrected camera coordinate systems to the world coordinate system for fusion, so as to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
[0120] As can be seen from the above description, the large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception provided by the embodiments of the present application generates a first point cloud from the images captured by the camera, generates a second point cloud from the radar, and takes advantage of the more accurate radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera by this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation, a first corrected point cloud is obtained. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.
[0121] From a hardware perspective, in order to obtain more accurate three-dimensional map modeling in a larger range, the present application makes full use of the respective advantages of the camera and the radar, and provides an embodiment of an electronic device for implementing all or part of the content in the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception. The electronic device specifically includes the following:
[0122] A processor, a memory, a communication interface, and a bus; wherein, the processor, the memory, and the communication interface complete communication with each other through the bus; the communication interface is used to implement information transmission between the large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception and related devices such as a core business system, a user terminal, and a related database. The logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception and the embodiments of the large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception, and the content thereof is incorporated herein, and the repeated parts will not be described again.
[0123] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0124] In practical applications, part of the method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception can be executed on the electronic device side as described above, or all operations can be completed in the client device. The specific selection can be based on the processing capability of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor.
[0125] The client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.
[0126] Figure 4 FIG. 9 is a schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 4 As shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is worth noting that Figure 4 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0127] In one embodiment, the function of the large scene space three-dimensional map reconstruction method integrating multimodal data perception can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:
[0128] Step S101: obtaining a video frame sequence captured by a camera, and calculating a first point cloud according to adjacent frames of video; and obtaining a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and face the same direction;
[0129] Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose;
[0130] Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud;
[0131] Step S104: Convert the first corrected point clouds in multiple corrected camera coordinate systems to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
[0132] As can be seen from the above description, the electronic device provided in the embodiment of the present application generates a first point cloud from an image captured by a camera, generates a second point cloud from a radar, and utilizes the more accurate advantage of radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera by this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation to obtain a first corrected point cloud. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and further obtaining a more accurate three-dimensional map model.
[0133] In another embodiment, the large-scale scene space three-dimensional map reconstruction device that fuses multi-modal data perception can be separately configured from the central processing unit 9100. For example, the large-scale scene space three-dimensional map reconstruction device that fuses multi-modal data perception can be configured as a chip connected to the central processing unit 9100, and the functions of the large-scale scene space three-dimensional map reconstruction method that fuses multi-modal data perception can be realized through the control of the central processing unit.
[0134] As Figure 4 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include all the components shown in Figure 4 ; in addition, the electronic device 9600 may further include components not shown in Figure 4 , and reference can be made to the prior art.
[0135] As Figure 4 shown, the central processing unit 9100, sometimes also referred to as a controller or operation control, may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives inputs and controls the operation of various components of the electronic device 9600.
[0136] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. It can store the above-mentioned failure-related information, and can also store programs for executing relevant information. And the central processing unit 9100 can execute the programs stored in the memory 9140 to implement information storage or processing, etc.
[0137] The input unit 9120 provides inputs to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.
[0138] The memory 9140 can be a solid-state memory. For example, it can be a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be a memory that stores information even when powered off, can be selectively erased and has more data. Examples of such a memory are sometimes referred to as EPROMs, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 can include an application / function storage unit 9142, which is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.
[0139] The memory 9140 can also include a data storage unit 9143, which is used to store data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 can include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0140] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which can be the same as in the case of a conventional mobile communication terminal.
[0141] Based on different communication technologies, in the same electronic device, multiple communication modules 9110 can be provided, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, so as to implement normal telecommunication functions. The audio processor 9130 may include any suitable buffers, decoders, amplifiers, etc. Additionally, the audio processor 9130 is also coupled to a central processor 9100, so that it is possible to record sound on the local device through the microphone 9132, and it is possible to play the sound stored on the local device through the speaker 9131.
[0142] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception with the execution subject being a server or a client in the above embodiments. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, all steps of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception with the execution subject being a server or a client in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:
[0143] Step S101: Obtain a sequence of video frames captured by a camera, and calculate a first point cloud based on adjacent frame videos; and, obtain a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and have the same orientation;
[0144] Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose;
[0145] Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud;
[0146] Step S104: Convert multiple first corrected point clouds in the corrected camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform a surface reconstruction operation on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
[0147] As can be seen from the above description, the computer-readable storage medium provided by the embodiments of the present application generates a first point cloud from an image captured by a camera, generates a second point cloud from a radar, and takes advantage of the more accurate radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera by the deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation, a first corrected point cloud is obtained. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.
[0148] Embodiments of the present application also provide a computer program product that can implement all steps of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception with the execution subject being a server or a client in the above embodiments. When the computer program / instructions are executed by a processor, the steps of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception are implemented. For example, the computer program / instructions implement the following steps:
[0149] Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose;
[0150] Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud;
[0151] Step S104: Convert the first corrected point clouds in multiple corrected camera coordinate systems to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
[0152] As can be seen from the above description, the computer program product provided by the embodiments of the present application generates a first point cloud from an image captured by a camera, generates a second point cloud from a radar, and takes advantage of the more accurate radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera with this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation to obtain a first corrected point cloud. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.
[0153] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0154] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more of these flows Figure 1 or a combination of multiple flows and / or blocks
[0155] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one or more of these flows Figure 1 or a combination of multiple flows and / or blocks
[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks. Figure 1 The steps for implementing the functions specified in one block or a plurality of blocks.
[0157] Specific embodiments are used in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception, characterized in that, The method includes: Obtaining a sequence of video frames captured by a camera, and calculating a first point cloud based on adjacent frame videos; and, obtaining a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and have the same orientation; Registering the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculating the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correcting the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; Converting the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then converting the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud; Converting multiple first corrected point clouds in the corrected camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; performing surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
2. The method for reconstructing a three-dimensional map of a large-scale scene integrating multi-modal data perception according to claim 1, wherein The step of calculating the first point cloud based on adjacent frame videos includes: Performing feature point matching on the same target in adjacent frame videos to obtain feature point pairs of each same target; Constructing a fundamental matrix based on the feature point pairs, and calculating an essential matrix according to the fundamental matrix and the internal parameters of the camera; Decomposing the essential matrix to obtain a first rotation matrix and a first translation vector between two adjacent frame cameras; Determining the three-dimensional coordinates of each feature point through triangulation according to the internal parameters of the camera and the first rotation matrix and the first translation vector; the three-dimensional coordinates of the feature points of multiple targets form the first point cloud.
3. The method for reconstructing a three-dimensional map of a large-scale scene that fuses multi-modal data perception according to claim 2, wherein The step that the three-dimensional coordinates of the feature points of multiple targets form the first point cloud includes: Taking the center point of the camera rectangular screen range as the origin, taking X% of the rectangle length as the major axis, and taking X% of the rectangle width as the minor axis to draw an ellipse, and selecting the feature points located inside the intersection area of the rectangle and the ellipse for feature point matching; the X% is 5%-15%.
4. The method for reconstructing a three-dimensional map of a large-scale scene integrating multi-modal data perception according to claim 1, characterized in that, The step of registering the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector includes: For each point in the first point cloud, finding the point in the second point cloud that is closest to it through the nearest neighbor search algorithm to form a matching point pair; Calculating a second rotation matrix and a second translation vector that can align all the matching point pairs based on the principle of rigid transformation; Using the second rotation matrix and the second translation vector to transform the points in the first point cloud to make them closer to the point in the first point cloud that is closest to it; Repeating the above steps until the iteration times reach the upper limit, or when the distance change between the first point cloud and the second point cloud is less than a preset value, the obtained second rotation matrix and second translation vector are the optimal rotation matrix and the optimal translation vector.
5. The method for reconstructing a three-dimensional map of a large-scale scene integrating multi-modal data perception according to claim 4, wherein The step of calculating a second rotation matrix and a second translation vector that can align all the matching point pairs based on the principle of rigid transformation includes: In each iteration, a part of the matching point pairs are randomly selected, and a transformation matrix for aligning these matching point pairs is calculated, where the transformation matrix includes the second rotation matrix and the second translation vector; Apply the transformation matrix to the first point cloud, and calculate the distances between the transformed first point cloud and other matching point pairs in the second point cloud; The matching point pairs whose distances meet the set threshold are regarded as inliers and retained, while other non-conforming matching points are regarded as outliers and removed.
6. The method for reconstructing a three-dimensional map of a large-scale scene integrating multi-modal data perception according to claim 5, characterized in that The step of randomly selecting a part of the matching point pairs in each iteration and calculating the transformation matrix for aligning these matching point pairs includes: Taking the camera parameters, the points of the first point cloud, and the points in the second point cloud as variable nodes, and constructing a factor graph with the reprojection error between the corresponding point coordinates of the first point cloud and the second point cloud as the observation edges; the camera parameters include internal parameters and external parameters; the external parameters include the second rotation matrix and the second translation vector; Perform graph optimization on the factor graph to obtain the transformation matrix; the factor graph uses sparse linear algebra and a solver for graph optimization.
7. The method for reconstructing a three-dimensional map of a large-scale scene integrating multi-modal data perception according to claim 1, wherein The method for obtaining the initial pose is: based on the known camera poses, calculate the initial poses of the cameras for shooting each video frame by accumulating the relative poses frame by frame.
8. A three-dimensional map reconstruction device for large-scale scene space that integrates multi-modal data perception, characterized in that, The device includes: A point cloud generation module, configured to obtain a sequence of video frames captured by a camera, and calculate a first point cloud based on adjacent video frames; and obtain a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and have the same orientation; A pose correction module, configured to register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; A point cloud correction module, configured to convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud; A map construction module, configured to convert multiple first corrected point clouds in the corrected camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for reconstructing a three-dimensional map of a large-scale scene with multi-modal data perception according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for reconstructing a three-dimensional map of a large-scale scene with multi-modal data perception according to any one of claims 1 to 7.
Citation Information
Patent Citations
Three-dimensional scene reconstruction method and device for reconstruction and extension site
CN115100367A
Method, device and system for generating three-dimensional model point cloud of object to be modeled
CN115830217A