Multi-modal data perception fused large-scene space three-dimensional map reconstruction method

By using registration correction technology of camera and radar data in the low-altitude economy, the problem of insufficient accuracy of three-dimensional map modeling is solved, and more accurate three-dimensional map model generation is achieved.

CN120088423AActive Publication Date: 2025-06-03UNIVERSAL UBIQUITOUS TECH CO LTD

Patent Information

Application Number
CN202510571975.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-03
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In the low-altitude economy field, how to make full use of the breadth advantages of camera image data and the accuracy advantages of radar point cloud data to improve the accuracy of three-dimensional map modeling.

Method used

By acquiring point cloud data generated by the camera and radar, using the registration algorithm to correct the camera position, and then correct the camera point cloud data, and combining the surface reconstruction algorithm to generate a more accurate three-dimensional map model.

Benefits of technology

Improve the camera's point cloud correction accuracy at long distances, obtain a more accurate three-dimensional map model, combining the advantages of the camera and radar.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088423A_ABST
    Figure CN120088423A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-modal data perception fused large-scene space three-dimensional map reconstruction method. The method comprises the steps of performing registration on a first point cloud of a camera and a second point cloud of a radar to obtain an optimal rotation matrix and an optimal translation vector; correcting the initial pose of the camera according to the inverse matrix of the optimal rotation matrix and the reverse translation vector to obtain a corrected camera pose; converting a three-dimensional point of the first point cloud from a camera coordinate system to a world coordinate system according to the initial pose, and then converting the three-dimensional point of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud; fusing the plurality of first correction point clouds to obtain a three-dimensional point cloud of the whole map; and performing surface reconstruction operation on the map point cloud according to a surface reconstruction algorithm. According to the embodiment of the invention, the camera can obtain more accurate point cloud data on the basis of keeping the long-distance advantage, so that a more accurate three-dimensional map model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and specifically relates to a method for reconstructing a three-dimensional map of a large-scale scene space that integrates multi-modal data perception. Background Art

[0002] In the field of low-altitude economy, through three-dimensional map reconstruction, accurate navigation information can be provided for low-altitude aircraft, including airspace routes, takeoff and landing point positions, etc., thereby improving flight safety and efficiency. At the same time, the three-dimensional map can also be used for flight supervision to monitor the position and status of the aircraft in real time to ensure that flight activities comply with regulations. In urban planning, the three-dimensional map can provide detailed spatial information to help planners better understand the urban terrain, building layout, etc., so as to make more reasonable planning decisions. In addition, the three-dimensional map can also be used for the simulation and prediction of urban construction to evaluate the impact of different construction plans on the urban environment and traffic. Even when natural disasters or emergencies occur, the three-dimensional map can quickly provide information such as the terrain and building distribution of the disaster area to provide decision-making support for emergency rescue. At the same time, by comparing the three-dimensional maps before and after the disaster, the losses and impacts caused by the disaster can be evaluated to provide a scientific basis for post-disaster reconstruction.

[0003] Current three-dimensional reconstruction data comes from video data (RGB and IR) of cameras and point cloud data based on radar. Specifically, when performing three-dimensional reconstruction based on camera image data, the image data can be reconstructed into three-dimensional point cloud data through three-dimensional coordinate point recovery operations, and then the corresponding three-dimensional model can be determined by performing surface reconstruction operations on the three-dimensional point cloud data according to a preset surface reconstruction algorithm. Or directly convert the point cloud data of the radar into a three-dimensional model based on a surface reconstruction algorithm, etc. In the field of low-altitude economy, the shooting range of the camera is often far, reaching hundreds of meters or even several kilometers, while the point cloud generated by the radar is limited by cost within 200m. Generally, the range breadth of the point cloud generated by the camera is greater than that of the point cloud generated by the radar, but the accuracy of the point cloud generated by the radar is greater than that of the point cloud generated by the camera. How to make full use of the advantages of both to obtain a more accurate three-dimensional map model is a technical problem that technicians in this field have always had to solve. Summary of the Invention

[0004] Aiming at the problems in the prior art, this application provides a method for reconstructing a three-dimensional map of a large-scale scene space that integrates multi-modal data perception to improve the accuracy of three-dimensional map modeling in the field of low-altitude economy by utilizing the breadth advantage of camera image data and the accuracy advantage of radar point cloud data.

[0005] To solve at least one of the above problems, this application provides the following technical solutions: In the first aspect, this application provides a method for reconstructing a three-dimensional map of a large-scale scene space that integrates multi-modal data perception, including: Obtain a sequence of video frames captured by a camera, and calculate a first point cloud based on adjacent frame videos; and, obtain a second point cloud generated by a radar; the camera and the radar are adjacently arranged and have the same orientation; Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud; Convert multiple first corrected point clouds in the corrected camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

[0006] Further, the step of calculating the first point cloud based on adjacent frame videos includes: Perform feature point matching on the same target in adjacent frame videos to obtain feature point pairs of each same target; Construct a fundamental matrix based on the feature point pairs, and calculate an essential matrix according to the fundamental matrix and the internal parameters of the camera; Decompose the essential matrix to obtain a first rotation matrix and a first translation vector between two adjacent frame cameras; Determine the three-dimensional coordinates of each feature point through triangulation according to the internal parameters of the camera and the first rotation matrix and the first translation vector; the three-dimensional coordinates of the feature points of multiple targets form the first point cloud.

[0007] Further, the step that the three-dimensional coordinates of the feature points of multiple targets form the first point cloud includes: Take the center point of the camera rectangular screen range as the origin, draw an ellipse with X% of the rectangle length as the major axis and X% of the rectangle width as the minor axis, and select the feature points located inside the intersection area of the rectangle and the ellipse for feature point matching; the X% is 5%-15%.

[0008] Further, the step of registering the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector includes: For each point in the first point cloud, find the point in the second point cloud with the closest distance to it through the nearest neighbor search algorithm to form a matching point pair; Based on the principle of rigid transformation, calculate the second rotation matrix and the second translation vector that can align all the matching point pairs; Use the second rotation matrix and the second translation vector to transform the points in the first point cloud to make them closer to the points in the first point cloud that are closest to them; Repeat the above steps until the number of iterations reaches the upper limit, or when the distance change between the first point cloud and the second point cloud is less than a preset value, the obtained second rotation matrix and second translation vector are the optimal rotation matrix and the optimal translation vector.

[0009] Further, the step of calculating the second rotation matrix and the second translation vector that can align all the matching point pairs based on the principle of rigid transformation includes: In each iteration, randomly select some matching point pairs and calculate the transformation matrix for aligning these matching point pairs. The transformation matrix includes the second rotation matrix and the second translation vector; Apply the transformation matrix to the first point cloud and calculate the distances between the transformed first point cloud and other matching point pairs in the second point cloud; Regard the matching point pairs whose distances meet the set threshold as inliers and retain them, and regard other non-conforming matching points as outliers and remove them.

[0010] Further, the step of randomly selecting some matching point pairs in each iteration and calculating the transformation matrix for aligning these matching point pairs includes: Take the camera parameters, the points in the first point cloud, and the points in the second point cloud as variable nodes, and construct a factor graph with the reprojection error between the corresponding point coordinates in the first point cloud and the second point cloud as the observation edges; the camera parameters include internal parameters and external parameters; the external parameters include the second rotation matrix and the second translation vector; Perform graph optimization on the factor graph to obtain the transformation matrix; the factor graph uses sparse linear algebra and a solver for graph optimization.

[0011] Further, the method for obtaining the initial pose is: based on the known camera pose, calculate the initial pose of the camera for each frame of video captured by accumulating the relative poses frame by frame.

[0012] In a second aspect, the present application provides a large-scale scene space three-dimensional map reconstruction device that fuses multi-modal data perception, including: A point cloud generation module, configured to obtain a sequence of video frames captured by a camera and calculate a first point cloud based on adjacent frame videos; and obtain a second point cloud generated by a radar; the camera and the radar are adjacent and have the same orientation; The pose correction module is used to register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; The point cloud correction module is used to convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain the first corrected point cloud; The map construction module is used to convert multiple first corrected point clouds in the corrected camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

[0013] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception are implemented.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception are implemented.

[0015] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception are implemented.

[0016] As can be seen from the above technical solutions, the present application provides a method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception. The first point cloud is generated from the images captured by the camera, and the second point cloud is generated from the radar. By utilizing the more accurate advantage of the radar data, the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera is obtained, and the pose of the camera is corrected by this deviation value. Then, the three-dimensional points of the first point cloud are converted to the corrected camera coordinate system through coordinate transformation to obtain the first corrected point cloud. Compared with directly correcting the point cloud by using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of a large-scale scene space three-dimensional map reconstruction method that integrates multi-modal data perception in the embodiments of the present application; Figure 2 It is a selected point range area diagram of a large-scale scene space three-dimensional map reconstruction method that integrates multi-modal data perception in the embodiments of the present application; Figure 3 It is a structural diagram of a large-scale scene space three-dimensional map reconstruction device that integrates multi-modal data perception in the embodiments of the present application; Figure 4 It is a schematic structural diagram of an electronic device in the embodiments of the present application.

[0019] Reference numerals: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0021] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.

[0022] In view of the problems existing in the prior art, the present application provides a method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception. A first point cloud is generated from an image captured by a camera, and a second point cloud is generated by a radar. By taking advantage of the more accurate radar data, the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera is obtained, and the pose of the camera is corrected by this deviation value. Then, the three-dimensional points of the first point cloud are transformed to the corrected camera coordinate system through coordinate transformation to obtain a first corrected point cloud. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.

[0023] In order to obtain more accurate three-dimensional map modeling in a large range, the present application makes full use of the respective advantages of the camera and the radar, and provides an embodiment of a method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception. Refer to Figure 1 , the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception specifically includes the following content: Step S101: Obtain a video frame sequence captured by the camera, and calculate a first point cloud based on adjacent frame videos; and obtain a second point cloud generated by the radar; the camera and the radar are arranged adjacent to each other and have the same orientation.

[0024] Optionally, in this embodiment, the camera and the radar are arranged on the belly of a low-altitude aircraft, and the two are installed very close to each other in the physical space, or even share the same installation position or area. Therefore, the first point cloud can be corrected using the second point cloud, and the absolute pose of the camera can be corrected in the reverse direction using the obtained deviation value. The point cloud data of the camera is obtained again using the corrected pose, and finally, a first corrected point cloud data with higher accuracy than the first point cloud data is obtained.

[0025] Optionally, in this embodiment, the step of calculating the first point cloud based on adjacent frame videos includes: using image processing algorithms such as SIFT, SURF, and ORB to extract feature points in adjacent frame videos, and then using the nearest neighbor algorithm to match the feature points of the same target in adjacent frame videos to obtain feature point pairs of each same target; constructing an essential matrix based on the feature point pairs, and calculating an essential matrix according to the essential matrix and the internal parameters of the camera (focal length, principal point coordinates, distortion coefficients, etc.); by decomposing the essential matrix, obtaining a first rotation matrix and a first translation vector between two adjacent frame cameras, that is, the relative pose of two adjacent frame cameras; finally, according to the internal parameters of the camera and the first rotation matrix and the first translation vector, determining the three-dimensional coordinates of each feature point through triangulation; the three-dimensional coordinates of the feature points of multiple targets form a first point cloud.

[0026] Specifically, the fundamental matrix is defined as: F = K -T EK -1 Where F is the fundamental matrix, E is the essential matrix, and K is the intrinsic matrix of the camera.

[0027] The relationship between the essential matrix E and the relative pose of the camera (rotation matrix R and translation vector t) is: E = [t] × R Where [t] × is the skew-symmetric matrix of the translation vector t.

[0028] The calculation formula for the three-dimensional coordinates of the feature points is: P = K -1 p 1 × (RK -1 p 2 + t) Where p 1 and p 2 are the feature points in the camera coordinate system of two adjacent frames, K is the intrinsic matrix of the camera, (R, t) is the relative pose of the camera, and × is the cross product of vectors.

[0029] Optionally, in this embodiment, considering that the usage scenario of the present invention is a large spatial scale, small errors at the camera pose end will be amplified at the far end of the picture. Although distortion correction of the picture can restore the picture, there is still a distortion problem at the edge of the restored picture. Therefore, to eliminate the influence of lens distortion on the final result, in this embodiment, when selecting feature points, the feature points in the middle area of the picture are selected, and the edge feature points are ignored. At the same time, a part of the edge picture is discarded, and only the middle part of the picture is used to minimize the influence of lens distortion on calculating the camera pose.

[0030] The point selection range area mask is as Figure 2 shown. Taking the center point of the camera rectangular picture range as the origin, X% of the rectangle length as the major axis, and X% of the rectangle width as the minor axis to draw an ellipse, and select the feature points located inside the intersection area of the rectangle and the ellipse for feature point matching, and discard the remaining feature points. The rectangle is the original picture range. Exemplarily, if the camera is 16:9, the resolution is 1920 × 1080. X is determined according to the specific distortion situation, and the range is 5% - 15%. For example, 5%, 10%, 15% etc. are selected.

[0031] Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose.

[0032] Optionally, in this embodiment, the ICP (Iterative Closest Point) algorithm is used to register the first point cloud and the second point cloud. That is, the second point cloud of the radar is used as the target point cloud, and the first point cloud of the camera is used as the source point cloud to minimize the error function E between the first point cloud and the second point cloud:

[0033] where P i and Q i are the matching points in the first point cloud and the second point cloud respectively, and R and t are the rotation matrix and translation vector of the camera respectively.

[0034] That is, during the point cloud registration process, the ICP algorithm iteratively finds the closest point pairs between the two point clouds (the first point cloud and the second point cloud) and calculates the optimal rigid transformation (including the rotation matrix R and the translation vector t) so that the source point cloud (the first point cloud) can be aligned in the coordinate system of the target point cloud (the second point cloud). If each point in the first point cloud is transformed according to the rotation matrix R and the translation vector t, a corrected point cloud can be obtained. This corrected point cloud is closer to the second point cloud generated by the radar in space and can improve the accuracy of the first point cloud. Then, since the rotation matrix R and the translation vector t are the transformation matrices for transforming the first point cloud to the second point cloud, the reverse correction of the camera pose requires using the inverse matrix of R (or the transpose matrix, because R is an orthogonal matrix and its inverse matrix is equal to the transpose matrix) and -t (i.e., the opposite of the translation vector) to update the camera pose so that it can be aligned with the corrected point cloud. Finally, the point cloud data of the camera is obtained again using the corrected pose, and finally the first corrected point cloud data with higher accuracy than the first point cloud data is obtained.

[0035] Specifically, in this embodiment, the registration steps include: for each point in the first point cloud, find the point in the second point cloud that is closest to it through the nearest neighbor search algorithm to form a matching point pair; based on the principle of rigid transformation, calculate the second rotation matrix and the second translation vector that can align all the matching point pairs; use the second rotation matrix and the second translation vector to transform the points in the first point cloud to make them closer to the points in the first point cloud that are closest to them; repeat the above steps until the number of iterations reaches the upper limit or the distance change between the first point cloud and the second point cloud is less than the preset value. The obtained second rotation matrix and second translation vector are the optimal rotation matrix and optimal translation vector.

[0036] Optionally, in this embodiment, considering that the distortion of the central part of the camera's screen has the least impact, and the closer the radar is to the central area, the higher its accuracy. When performing ICP matching, a weight ratio can be increased for these central areas. The specific weight increase scheme is as follows: To make the weight of the middle area greater than that of the edge area, an appropriate σ value can be set so that the value of the Gaussian function is larger in the central area and smaller in the edge area. Use these weights to adjust the point cloud data so that the points in the central area have a greater influence during the registration process. The specific formula is as follows: Among them, d represents the distance from a point to a certain reference point (the center point of the screen in this example), and σ is the standard deviation, which is used to control the width of the Gaussian distribution.

[0037] In point cloud registration, the closer a point is to the central area (i.e., the smaller the d value), the greater its weight; the farther a point is from the central area, the smaller its weight. If the standard deviation σ is larger, the Gaussian distribution is wider, indicating a greater degree of dispersion of the data points; if the standard deviation σ is smaller, the Gaussian distribution is narrower, indicating that the data points are more concentrated.

[0038] Specifically, in this example, first check the accuracy of the radar data. If the deviation between the middle data error and the edge data error in the measurement result is greater, that is, the data in the center of the radar data is more accurate, then a smaller value should be selected for d. Exemplarily, in this example, the d value can be selected as 1.

[0039] Optionally, in this embodiment, considering that this point cloud registration is multi-point matching, the quality of the matching points is crucial for the registration result. If the matching point pairs are inaccurate, it will lead to a poor registration result. Therefore, this embodiment introduces the RANSAC algorithm (RANdom Sample Consensus) to optimize the registration, that is, using the RANSAC algorithm, by presetting a reasonable threshold to distinguish normal points (inliers) and abnormal points (outliers), and screening out the correct transformation parameters (optimal rotation matrix and optimal translation vector) from the matching points to reduce the influence of abnormal points on the registration and improve the robustness and accuracy of the registration.

[0040] Specifically, in this embodiment, the steps of calculating the second rotation matrix and the second translation vector that can align all the matching point pairs based on the principle of rigid transformation include: In each iteration, randomly select a part of the matching point pairs, preferably the minimum number of matching point pairs, such as 3 matching point pairs, and calculate the transformation matrix that aligns these matching point pairs. This transformation matrix includes the second rotation matrix and the second translation vector; Apply the transformation matrix to the first point cloud, and calculate the distance between the transformed first point cloud and other matching point pairs in the second point cloud to evaluate the quality of the transformation; Consider the matching point pairs whose distances meet the set threshold as inliers and retain them, and consider other non-conforming matching points as outliers and remove them. The RANSAC algorithm will continue to iterate and update the optimal transformation matrix until a transformation matrix that can best match the inliers is found.

[0041] Optionally, since in the RANSAC algorithm, it is usually necessary to find the optimal model in a large amount of data. Therefore, to improve the operation speed, in this embodiment, a factor graph is used to process these constraints in a more structured and efficient manner. That is, in the RANSAC algorithm of this embodiment, a model is estimated by randomly selecting samples from the matching points, and it is checked which point pairs conform to the model (i.e., inliers). By introducing the factor graph into the RANSAC algorithm, this process can be accelerated by modeling the constraints and optimization problems, especially in the case of applying to large-scale data.

[0042] A factor graph is a graphical model that consists of variable nodes and factor nodes. The factor nodes represent the constraints between variables. Usually, it is used to express the relationships between different variables in an optimization problem. Therefore, in the RANSAC algorithm of this embodiment, the factor graph can structurally represent the transformation (such as rotation and translation) between point clouds, and find the optimal transformation through optimization, so that this embodiment can efficiently find the optimal solution without directly brute-forcing through all possible matching points, and improve the search speed of the optimal solution.

[0043] Specifically, in this embodiment, the general framework of introducing the factor graph into the RANSAC algorithm is: Take the camera parameters (including internal parameters such as focal length, principal point coordinates, etc., and external parameters such as the camera pose described by the second rotation matrix and the second translation vector), the points of the first point cloud, and the points in the second point cloud as variable nodes, and add the BA (Bundle Adjustment) algorithm as a constraint to the factor graph to further optimize the camera pose and the three-dimensional coordinates of the feature points, for refining the geometric model obtained from the initial SIFT / SURF / ORB and other feature matches.

[0044] In this embodiment, the reprojection error between the corresponding point coordinates in the first point cloud and the second point cloud is used as the observed edge, or called the factor node, to construct the factor graph. Each matching point pair serves as a factor node, connecting two variable nodes (e.g., the transformation matrix corresponding to the first point cloud and the second point cloud, or called the transformation parameter). The task of the factor node is to measure the error of the point cloud matching under the current transformation. By minimizing the reprojection error, the positions of the three-dimensional points in the scene and the internal and external parameters of the camera are jointly optimized to obtain the most accurate geometric model and camera pose.

[0045] More specifically, for a known three-dimensional point P and camera parameters, the projected coordinate u of this point on the image can be calculated. If the position of the feature point observed by the radar is u′, the reprojection error is defined as the difference between the two: e = u′ - u. The goal of BA is to find a set of camera parameters and three-dimensional point coordinates that minimize the sum of the reprojection errors of all feature points on all images.

[0046] Therefore, in each iteration of the RANSAC algorithm in this embodiment, the estimated transformation parameters are used as variable nodes in the factor graph. Through optimization algorithms, such as the Gauss-Newton method, the Levenberg-Marquardt algorithm, etc., these transformation parameters are optimized. The goal of the optimization is to minimize the errors of the constraints in all factor nodes to obtain the optimal transformation matrix. Then, based on the current optimized transformation, the error of each matching point pair is calculated to determine whether the matching point is an inlier. If the error is less than a certain threshold, the matching point is considered an inlier, and the factor graph is updated. That is, after each iteration, the transformation parameters obtained by optimizing the factor graph are compared with the current optimal transformation, and the optimal transformation (the transformation with the most inliers) is updated. As a preferred solution, this embodiment uses sparse linear algebra and an efficient solver for graph optimization.

[0047] In the traditional RANSAC algorithm, for each fitting, all three-dimensional points need to be traversed to calculate the inlier or outlier status of each three-dimensional point. The factor graph can help identify and group three-dimensional points with similar constraint relationships, thereby accelerating the processing of these three-dimensional points. By decomposing and parallelizing these tasks, the factor graph can significantly improve the computational efficiency of the RANSAC algorithm. Moreover, through the factor graph, the RANSAC algorithm can more efficiently represent and utilize the relationship between inliers and outliers. This structured representation helps reduce redundant calculations when selecting inliers and outliers. And the factor graph can converge to a reasonable solution faster in each round of calculation through iterative optimization. Compared with the random selection and repeated trial of the traditional RANSAC algorithm, the factor graph quickly determines which three-dimensional points belong to inliers through a more refined optimization process, thus accelerating the convergence speed of the overall algorithm.

[0048] At the same time, the factor graph can not only select inliers and outliers in each RANSAC iteration, but also optimize the quality of the solution through a global optimization method. The factor graph can simultaneously process the constraint relationship of multiple sets of data, and correct the deviation through an optimization algorithm to reduce the influence of outliers on the final model, thereby improving the accuracy of the selection of inliers at the global level. In this way, the system does not need to repeatedly calculate the inliers and outliers of all points every time, while improving the calculation speed. Moreover, in traditional methods, the selection of inliers and outliers usually depends on local distance metrics, which may lead to large errors. In this embodiment, the factor graph helps to more accurately screen out inliers in three-dimensional points by utilizing more complex global constraints, such as pose constraints. This precise screening further reduces unnecessary calculations, especially when the amount of data is large. Therefore, since the factor graph can take into account the global constraints between all points, it is more accurate in the selection of inliers than the traditional RANSAC method.

[0049] Therefore, in this embodiment, through the factor graph, the RANSAC algorithm can use global constraints and optimization methods when selecting the interior points and exterior points of three-dimensional points to improve the accuracy of interior point selection and reduce redundant calculations, thereby significantly speeding up the calculation speed, improving robustness and accelerating the convergence process to adapt to the processing of large-scale point cloud data.

[0050] Optionally, in this embodiment, the method for obtaining the initial posture is: based on the known camera posture, the initial posture of the camera shooting each frame of the video is calculated by accumulating the relative posture frame by frame. That is, the posture of the camera when shooting the video (that is, the rotation matrix R and translation vector t of each frame of the camera) is calculated by analyzing the relative motion between adjacent frames of the camera. That is, assuming that the first frame posture of the camera is T 0 =[R 0 ∣t 0 ] is known and is usually set to the unit matrix (i.e. R 0 =I and t 0 = 0), the goal is to calculate the camera's pose T in each of the next frames i =[R i ∣t i ], where R i is the rotation matrix, t i is the translation vector, such as the first rotation matrix and the first translation vector in the above steps.

[0051] Specifically, the Lucas-Kanade optical flow algorithm or the dense optical flow method is used to estimate the motion of the camera using the motion of the pixels in the image. The rotation matrix R and the translation vector t obtained from the above method represent the relative position of the camera. In practical applications, if the initial position of the camera is known (for example, T 0), the formula for calculating the absolute pose of each frame by accumulating the relative poses frame by frame is: T i = T i-1 ·T rel where T i-1 is the camera pose of the previous frame, and T rel =[R rel ∣t rel is the relative pose from the previous frame to the current one. Each time a new camera pose is calculated, the absolute pose of the camera for the current frame is updated.

[0052] Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain the first corrected point cloud.

[0053] When calculating a three-dimensional point cloud, the triangulation method uses the relative position (baseline) between two frames of images and the parallax between corresponding points to calculate the position of each point in three-dimensional space. This process depends on the internal parameters (focal length, principal point coordinates, etc.) and external parameters (the pose of the camera, such as rotation and translation) of the camera to proceed. That is, the position of each three-dimensional point is reconstructed through the pose of the camera and the corresponding points in the image. The pose of the camera defines the conversion from the world coordinate system to the camera coordinate system. Therefore, if a three-dimensional point cloud has been obtained through the triangulation method, then the coordinates of these points are based on the camera coordinate system. If the camera pose is corrected, the three-dimensional point cloud needs to be converted from the old camera coordinate system to the new camera coordinate system.

[0054] Specifically, assume that the original three-dimensional point cloud P i =(X i ,Y i ,Z i ) is calculated relative to the initial camera pose T 0 =[R 0 ∣t 0 (where R 0 is the rotation matrix and t 0 is the translation vector). Now, the radar point cloud corrects the camera pose to obtain the corrected pose T 0 ′=[R 0 ′∣t 0 ′]. The conversion steps for converting all three-dimensional point clouds from the original coordinate system to the new coordinate system are as follows: 1. Convert the three-dimensional points from the camera coordinate system to the world coordinate system: Use the camera pose T 0 =[R 0 ∣t 0Convert each 3D point P i from the camera coordinate system to the world coordinate system: P i world =R 0 ·P i +t 0 2. Convert the 3D points from the world coordinate system to the new camera coordinate system: Use the new camera pose T 0 ′=[R 0 ′∣t 0 ′] to convert the point cloud to the new camera coordinate system: P i camera ′=R 0 ′ T ·(P i world t 0 ′) where R 0 ' T is the transpose matrix of R 0 ' (representing the rotation from the world coordinate system to the camera coordinate system).

[0055] 3. Combine the transformations: Combining these two steps, the 3D point cloud P 0 in the initial camera coordinate system T i can be used to calculate the 3D point cloud in the corrected camera coordinate system T 0 ′: P i camera ′=R 0 ' T ·(R 0 ·P i +t 0 -t 0 ′) Here, R 0 ·P i +t 0 is the result of converting the 3D points from the camera coordinate system to the world coordinate system. Subtracting t 0 ′ and multiplying by R 0 ' T completes the conversion of the 3D points from the world coordinate system to the corrected camera coordinate system.

[0056] When using cameras and radars to detect the low-altitude economic field, the point clouds formed by the cameras and radars only partially overlap within the fan-shaped radiation range near the main point of the device. If the point clouds are directly corrected using the transformation matrix after point cloud registration, as the distance increases, the correction error of the camera point cloud will become larger and larger, making it difficult to supplement the data range of the radar point cloud. Therefore, compared with traditional methods, the method of first correcting the camera pose using the radar point cloud and then correcting the camera point cloud using the corrected camera pose can greatly improve the correction accuracy of the camera points at a long distance, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.

[0057] Step S104: Convert the first corrected point clouds in multiple corrected camera coordinate systems into the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

[0058] Since multiple consecutive three-dimensional point clouds calculated by the triangulation method are respectively located in the three-dimensional coordinate systems of multiple camera frames, after calculating multiple consecutive three-dimensional point clouds using the triangulation method, it is necessary to align the three-dimensional coordinate systems of multiple camera frames into the same world coordinate system through point cloud fusion and merge multiple point clouds into a unified three-dimensional model.

[0059] The specific fusion method is as follows: Coordinate transformation: Perform coordinate transformation on each new point cloud (for example, the point cloud calculated in frame k), and convert it from the camera coordinate system of this frame to the world coordinate system. Referring to the above calculation process of the initial camera pose, obtain the rotation matrix R k and translation vector t k of the camera k relative to the reference frame (usually frame 0). Then, the formula for converting the point P k from the camera k coordinate system to the world coordinate system of the reference frame P 0 is: P 0 =R k ·P k +t k where R k is the coordinate of a certain three-dimensional point in the point cloud in the camera k coordinate system, and P 0 is the coordinate of this point in the world coordinate system of the reference frame.

[0060] Point cloud registration: Use point cloud registration algorithms such as the ICP algorithm or NDT algorithm to align the point clouds, and fuse the point clouds through stitching to obtain a complete and fused three-dimensional point cloud. Preferably, post-processing methods such as voxel filtering (Voxel Grid Filter) can also be used to reduce the density and noise of the point cloud.

[0061] Meshing: Perform meshing operations on the fused point cloud to obtain a three-dimensional grid model; map the surface information in the image onto the three-dimensional grid to form a three-dimensional map model with texture.

[0062] As can be seen from the above description, the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception provided by the embodiments of the present application generates a first point cloud from the images captured by the camera, generates a second point cloud from the radar, and utilizes the more accurate advantage of the radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera with this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation to obtain the first corrected point cloud. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.

[0063] In order to obtain more accurate three-dimensional map modeling in a larger range, the present application makes full use of the respective advantages of the camera and the radar, and provides an embodiment of a large-scale scene space three-dimensional map reconstruction device for implementing all or part of the content of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception. See Figure 3 , the large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception specifically includes the following content: Point cloud generation module 10, configured to obtain the video frame sequence captured by the camera, and calculate the first point cloud according to adjacent frame videos; and obtain the second point cloud generated by the radar; the camera and the radar are arranged adjacent to each other and have the same orientation; Pose correction module 20, configured to register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; Point cloud correction module 30, configured to convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain the first corrected point cloud; The map construction module 40 is configured to transform multiple first corrected point clouds in the corrected camera coordinate system into the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; and perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

[0064] As can be seen from the above description, the large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception provided by the embodiments of the present application generates a first point cloud from the images captured by the camera, generates a second point cloud from the radar, and takes advantage of the more accurate radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera by this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation to obtain the first corrected point cloud. Compared with directly correcting the point cloud using the transformation matrix and translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.

[0065] From the hardware level, in order to obtain more accurate three-dimensional map modeling in a larger range, the present application makes full use of the respective advantages of the camera and the radar, and provides an embodiment of an electronic device for implementing all or part of the content in the method for large-scale scene space three-dimensional map reconstruction for fusing multi-modal data perception. The electronic device specifically includes the following: A processor, a memory, a communication interface, and a bus; wherein, the processor, the memory, and the communication interface complete communication with each other through the bus; the communication interface is used to implement information transmission between the large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception and related devices such as the core business system, the user terminal, and the relevant database. The logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the method for large-scale scene space three-dimensional map reconstruction for fusing multi-modal data perception and the embodiments of the large-scale scene space three-dimensional map reconstruction device for fusing multi-modal data perception, and the content is incorporated herein, and the repeated parts will not be described again.

[0066] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0067] In practical applications, part of the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception can be executed on the side of the electronic device as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing power of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor.

[0068] The above-mentioned client device may have a communication module (i.e., a communication unit), which can be communicatively connected to a remote server to achieve data transmission with the server. The server may include a server on the side of the task scheduling center, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform communicatively linked to the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.

[0069] Figure 4 It is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of the present application. As Figure 4 shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 4 is exemplary; other types of structures may also be used to supplement or replace this structure to achieve telecommunication functions or other functions.

[0070] In one embodiment, the function of the method for reconstructing a three-dimensional map of a large-scale scene by fusing multi-modal data perception may be integrated into the central processing unit 9100. Among them, the central processing unit 9100 may be configured to perform the following controls: Step S101: Obtain a sequence of video frames captured by a camera, and calculate a first point cloud based on adjacent frame videos; and, obtain a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and have the same orientation; Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix, and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud; Step S104: Convert the multiple first corrected point clouds in the camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

[0071] As can be seen from the above description, the electronic device provided by the embodiment of the present application generates a first point cloud from the images captured by the camera, generates a second point cloud from the radar, and takes advantage of the more accurate radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera by this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation to obtain the first corrected point cloud. Compared with directly correcting the point cloud using the transformation matrix and translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud using the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.

[0072] In another embodiment, the large-scale scene space three-dimensional map reconstruction device that fuses multi-modal data perception can be separately configured from the central processing unit 9100. For example, the large-scale scene space three-dimensional map reconstruction device that fuses multi-modal data perception can be configured as a chip connected to the central processing unit 9100 to implement the functions of the large-scale scene space three-dimensional map reconstruction method through the control of the central processing unit.

[0073] As Figure 4 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include all the components shown in Figure 4 ; in addition, the electronic device 9600 may further include components not shown in Figure 4 , and reference can be made to the prior art.

[0074] As Figure 4 shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.

[0075] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. It can store the above-mentioned information related to failures, and can also store programs for executing relevant information. And the central processing unit 9100 can execute the program stored in the memory 9140 to achieve information storage or processing, etc.

[0076] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.

[0077] The memory 9140 can be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be a memory that stores information even when powered off, can be selectively erased, and has more data. Examples of this memory are sometimes referred to as EPROM, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 can include an application / function storage unit 9142, which is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.

[0078] The memory 9140 can also include a data storage unit 9143, which is used to store data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 can include various drivers for the communication function of the electronic device and / or for executing other functions of the electronic device (such as a messaging application, an address book application, etc.).

[0079] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which can be the same as in the case of a conventional mobile communication terminal.

[0080] Based on different communication technologies, in the same electronic device, multiple communication modules 9110 can be provided, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, so as to implement normal telecommunication functions. The audio processor 9130 can include any suitable buffers, decoders, amplifiers, etc. In addition, the audio processor 9130 is also coupled to a central processor 9100, so that it is possible to record on the local machine through the microphone 9132 and play the sound stored on the local machine through the speaker 9131.

[0081] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps in the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception with the execution subject being a server or a client in the above embodiments. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, all steps in the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception with the execution subject being a server or a client in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented: Step S101: Obtain a video frame sequence captured by a camera, and calculate a first point cloud based on adjacent frame videos; and obtain a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and have the same orientation; Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud; Step S104: Convert multiple first corrected point clouds in the corrected camera coordinate system to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform a surface reconstruction operation on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

[0082] As can be seen from the above description, the computer-readable storage medium provided by the embodiments of the present application generates a first point cloud from an image captured by a camera, generates a second point cloud from a radar, and takes advantage of the more accurate radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera by this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation, a first corrected point cloud is obtained. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.

[0083] The embodiments of the present application also provide a computer program product that can implement all the steps in the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception where the execution subject in the above embodiments is a server or a client. When the computer program / instructions are executed by a processor, the steps of the large-scale scene space three-dimensional map reconstruction method for fusing multi-modal data perception are implemented. For example, the computer program / instructions implement the following steps: Step S102: Register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; correct the initial pose of the camera according to the inverse matrix and the reverse translation vector to obtain the corrected camera pose; Step S103: Convert the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then convert the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose to obtain a first corrected point cloud; Step S104: Convert the first corrected point clouds in multiple corrected camera coordinate systems to the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

[0084] As can be seen from the above description, the computer program product provided by the embodiments of the present application generates a first point cloud from the images captured by a camera, generates a second point cloud from a radar, and takes advantage of the more accurate radar data. By obtaining the deviation between the second point cloud generated by the radar and the first point cloud generated by the camera, and correcting the pose of the camera with this deviation value, and then converting the three-dimensional points of the first point cloud to the corrected camera coordinate system through coordinate transformation to obtain a first corrected point cloud. Compared with directly correcting the point cloud using a transformation matrix and a translation vector, the method of first correcting the camera pose with the radar point cloud and then correcting the camera point cloud with the corrected camera pose can greatly improve the correction accuracy of the camera's long-distance points, enabling the camera to obtain more accurate point cloud data while maintaining the long-distance advantage, and thus obtaining a more accurate three-dimensional map model.

[0085] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, apparatus, or computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0086] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of these flows Figure 1 or a combination of multiple flows and / or blocks

[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of these flows Figure 1 or a combination of multiple flows and / or blocks

[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps of the process Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for the functions specified in one block or a plurality of blocks.

[0089] In the present invention, specific embodiments are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception, characterized in that: The method comprises: Acquire a video frame sequence captured by a camera, and calculate a first point cloud based on adjacent frames of video; and acquire a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and face the same direction; The first point cloud and the second point cloud are registered by a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; an inverse matrix of the optimal rotation matrix and an inverse translation vector of the optimal translation vector are calculated; an initial position and posture of the camera are corrected according to the inverse matrix and the inverse translation vector to obtain a corrected camera position and posture; Converting the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then converting the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose, to obtain a first corrected point cloud; The first corrected point clouds under multiple corrected camera coordinate systems are converted into the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; the map point cloud is subjected to surface reconstruction operation according to the surface reconstruction algorithm to obtain a three-dimensional map model.

2. The method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception according to claim 1, characterized in that: The step of calculating the first point cloud according to adjacent frame videos comprises: Matching feature points of the same target in adjacent frames of video to obtain feature point pairs of the same target; Constructing a basic matrix based on the feature point pairs, and calculating an essential matrix according to the basic matrix and the intrinsic parameters of the camera; Decomposing the essential matrix to obtain a first rotation matrix and a first translation vector between two adjacent frame cameras; The three-dimensional coordinates of each feature point are determined by triangulation according to the intrinsic parameters of the camera, the first rotation matrix, and the first translation vector; the three-dimensional coordinates of the feature points of multiple targets constitute the first point cloud.

3. The method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception according to claim 2, characterized in that: The step of composing the first point cloud with the three-dimensional coordinates of the feature points of the multiple targets comprises: An ellipse is drawn with the center point of the camera's rectangular image range as the origin, X% of the rectangle's length as the major axis, and X% of the rectangle's width as the minor axis, and feature points located inside the intersection of the rectangle and the ellipse are selected for feature point matching; the X% is 5%-15%.

4. The method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception according to claim 1, characterized in that: The step of registering the first point cloud and the second point cloud by a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector comprises: For each point in the first point cloud, the nearest neighbor search algorithm is used to find the point in the second point cloud that is closest to it to form a matching point pair; Based on the rigid transformation principle, calculating a second rotation matrix and a second translation vector capable of aligning all the matching point pairs; Using the second rotation matrix and the second translation vector, transform the point in the first point cloud so that the point is closer to the point in the first point cloud that is closest to the point in the first point cloud; Repeat the above steps until the number of iterations reaches an upper limit, or the distance change between the first point cloud and the second point cloud is less than a preset value, and the obtained second rotation matrix and second translation vector are the optimal rotation matrix and the optimal translation vector.

5. The method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception according to claim 4, characterized in that: The step of calculating the second rotation matrix and the second translation vector capable of aligning all the matching point pairs based on the rigid transformation principle comprises: In each iteration, randomly select some matching point pairs, and calculate a transformation matrix for aligning these matching point pairs, wherein the transformation matrix includes the second rotation matrix and the second translation vector; Applying the transformation matrix to the first point cloud, and calculating the distance between the transformed first point cloud and other matching point pairs in the second point cloud; Matching points whose distances meet the set threshold are considered as inliers and retained, while other matching points that do not meet the threshold are considered as outliers and removed.

6. The method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception according to claim 5, characterized in that: The steps of randomly selecting some matching point pairs in each iteration and calculating the transformation matrix for aligning these matching point pairs include: The camera parameters, points in the first point cloud and points in the second point cloud are used as variable nodes, and the reprojection errors between the coordinates of the points corresponding to the first point cloud and the second point cloud are used as observation edges to construct a factor graph; the camera parameters include intrinsic parameters and extrinsic parameters; the extrinsic parameters include the second rotation matrix and the second translation vector; The factor graph is graph optimized to obtain the transformation matrix; the factor graph is graph optimized using sparse linear algebra and a solver.

7. The method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception according to claim 1, characterized in that: The method for obtaining the initial posture is: based on the known camera posture, the initial posture of the camera shooting each frame of video is calculated by accumulating relative postures frame by frame.

8. A large scene space three-dimensional map reconstruction device integrating multimodal data perception, characterized in that: The device comprises: A point cloud generation module, used to obtain a video frame sequence captured by a camera, and calculate a first point cloud based on adjacent frames of video; and obtain a second point cloud generated by a radar; the camera and the radar are arranged adjacent to each other and face the same direction; A posture correction module is used to register the first point cloud and the second point cloud through a registration algorithm to obtain an optimal rotation matrix and an optimal translation vector; calculate the inverse matrix of the optimal rotation matrix and the reverse translation vector of the optimal translation vector; and correct the initial posture of the camera according to the inverse matrix and the reverse translation vector to obtain a corrected camera posture; a point cloud correction module, configured to transform the three-dimensional points of the first point cloud from the camera coordinate system to the world coordinate system according to the initial pose, and then transform the three-dimensional points of the first point cloud from the world coordinate system to the corrected camera coordinate system according to the corrected camera pose, so as to obtain a first corrected point cloud; The map construction module is used to convert the first corrected point clouds under multiple corrected camera coordinate systems into the world coordinate system for fusion to obtain the three-dimensional point cloud of the entire map; perform surface reconstruction operations on the map point cloud according to the surface reconstruction algorithm to obtain a three-dimensional map model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for reconstructing a large scene space three-dimensional map by integrating multimodal data perception as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional scene reconstruction method and device for reconstruction and extension site

    CN115100367A

  • Method, device and system for generating three-dimensional model point cloud of object to be modeled

    CN115830217A

  • Forest region positioning and three-dimensional reconstruction method and system based on multi-sensor fusion

    CN116228969A

  • Leiye space automatic registration method, system and terminal based on image feature learning

    CN118429402A

  • Point cloud three-dimensional reconstruction method based on secondary adjacent frame constraint pose map optimization

    CN118887345A

Cited By

  • Generative completion-based three-dimensional Gaussian splashing small sample new view synthesis method

    CN121392006A