3D Reconstruction Method for Complex Scenarios Based on Multi-Source Fusion

By combining NeRF and 3D GS, using SDF to associate multimodal data, optimize pose and scene expression, the problem of low three-dimensional reconstruction accuracy in complex scenes is solved, and high-precision three-dimensional reconstruction and new perspective rendering are achieved.

CN119006700BActive Publication Date: 2025-06-17INST OF AUTOMATION CHINESE ACAD OF SCI (LUOYANG) ROBOTICS & INTELLIGENT EQUIP INNOVATION INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410921977.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2025-06-17
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction technology has problems such as scale differences, field of view limitations, and insufficient texture in complex scenarios, resulting in inaccurate reconstruction results, especially in environments with poor performance in areas with lack of obvious structural characteristics.

Method used

A multi-source fusion-based approach is used, combining neural radiation field (NeRF) and three-dimensional Gaussian sputtering (3D GS), and multimodal data are associated with symbol distance field (SDF) as intermediate media to optimize the implicit expression of camera and lidar poses and scenes.

Benefits of technology

High-precision 3D reconstruction in complex scenes is realized, and the pose and scene expression can be jointly optimized given the initial pose, which improves the reconstruction accuracy and supports new perspective rendering and three-dimensional mesh model extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006700B_ABST
    Figure CN119006700B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional reconstruction method for complex scenes based on multi-source fusion, which realizes the information fusion of multi-sensors based on neural radiance fields and 3D Gaussian splatting. The entire scene is encoded in the signed distance field of NeRF, and then multi-modal data is associated through SDF as an intermediate medium. During the training process of NeRF, 3D GS is used as a guidance to explicitly supervise the estimated value of SDF and guide the sampling of NeRF rays. The present invention can jointly optimize the poses of cameras and lidars and the implicit expression of the scene under the condition of a given initial pose, obtain more accurate poses and three-dimensional reconstruction results of the scene, and achieve higher reconstruction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and relates to the technical field of three-dimensional reconstruction, and in particular to a method for three-dimensional reconstruction of complex scenes based on multi-source fusion. Background Art

[0002] Dense three-dimensional reconstruction of scenes has always been a popular research field. Currently, two main types of sensors are used, namely cameras and lidars.

[0003] In the past few decades, due to their cost-effectiveness, high resolution, and ability to capture rich semantic information, cameras have become the most commonly used sensors for three-dimensional reconstruction. In the traditional vision-based three-dimensional reconstruction process, methods such as Structure-from-Motion (SfM) or Simultaneous Localization and Mapping (SLAM) are usually used to calculate the camera pose and perform sparse reconstruction. Subsequently, the Multi-View Stereo (MVS) method is used for dense reconstruction. However, due to the limitations of camera sensors, pure vision methods have the following series of difficulties: First, the lack of depth information leads to a scale difference between the model reconstructed from images and the actual model as well as the camera motion; Second, vision-based methods require sufficient overlapping regions between images to accurately recover the camera motion, which is a challenge for cameras with a small field of view; Finally, the reconstructed environment needs to have rich textures and numerous visual features for reliable feature matching between images. These factors limit the use of vision-based three-dimensional reconstruction in some complex scenes to be challenging.

[0004] Lidars can capture accurate depth information by emitting laser beams and calculating the laser flight time, so they are hardly affected by light or the environment. The Lidar Odometry and Mapping (LOAM) method, as well as a series of improved algorithms such as LeGO-LOAM and F-LOAM, are the most widely used methods in current lidar-based three-dimensional reconstruction. Since lidars can detect objects in three-dimensional space, the LOAM method solves the scale singularity problem. In addition, the LOAM method uses the spatial structure as a feature, so three-dimensional reconstruction can still be performed in textureless scenes. However, this method performs poorly in some environments lacking obvious structural features, such as long corridors or playgrounds; in addition, lidar points are sparsely and unevenly distributed in space, that is, the farther away from the lidar, the fewer spatial points, resulting in difficulty in judging the similarity between two lidar point clouds. Therefore, the LOAM method has limitations in closed-loop detection and is easily affected by cumulative errors.

[0005] Therefore, combining lidar and cameras for 3D reconstruction can compensate for their respective disadvantages and obtain better reconstruction results. Most of the existing fusion-based reconstruction methods rely heavily on the strict synchronization of multi-modal data, and some of these methods also require additional auxiliary sensors such as inertial measurement units (IMUs) and global navigation satellite systems (GNSS). Additionally, traditional camera and lidar fusion is often insufficient. For example, some methods maintain a visual map and a lidar map simultaneously. When the input information is an image, the pose is solved using the visual map; when the input information is lidar, the lidar map is used for solution. This approach increases the system overhead and lacks information correlation between sensors, failing to fully integrate the information from different sensors.

[0006] The fundamental reason for the insufficient fusion of information from different sensors above lies in that it is very difficult to correlate multi-modal data. For images and lidar, the difficulties in correlating the two are in the following three aspects:

[0007] (1) Different data modalities; an image is essentially two-dimensional information, while lidar is three-dimensional information.

[0008] (2) Different data volumes; each image has approximately several million pixels, while lidar has only about one hundred thousand 3D points.

[0009] (3) Different sensor fields of view; the field of view of a camera is generally 60 - 100 degrees, while lidar has a 360-degree field of view in the horizontal direction.

[0010] These difficulties above result in the traditional fusion methods being in a "fighting separately" state and unable to obtain the best 3D reconstruction results. Moreover, the 3D reconstruction results obtained through fusion methods exist in the form of point clouds, that is, the entire reconstruction model consists of discrete 3D points in space. In many applications, since the discrete point cloud model cannot be infinitely subdivided, it often has to be converted into a triangular mesh model (Mesh) first, and errors will inevitably occur during this conversion process.

[0011] The proposal of the Neural Radiance Field (NeRF) technology has brought a powerful driving force to the research of 3D scene reconstruction and novel view synthesis. The core of the neural radiance field lies in using a MLP (Multi-Layer Perceptron) to represent the entire scene. Different from traditional explicit scene representations, the neural radiance field is an implicit representation of the scene, that is, the entire scene is stored in the neural network; information such as the color and transparency of each point in the scene can be obtained by inputting the query point into the neural network. However, since the neural radiance field relies on the photometric consistency between the rendered image and the real image as the supervision signal during training, the resulting scene representation tends to be more "photometrically consistent" rather than "geometrically correct"; in some cases, incorrect geometry can also result in a correct rendered image; therefore, using the original neural radiance field for 3D reconstruction often fails to obtain the correct result. Additionally, the neural radiance field requires the initialization of the network, and this initialization is only applicable to the case where the object is located at the center of the scene, such as image capture and reconstruction around a single object; in large-scale scenes, the above assumptions will be violated, leading to non-convergence of the neural network training.

[0012] The 3D Gaussian Splatting (3D GS) technology is another method for novel view synthesis. This method is a more explicit scene representation, and its core lies in using a set of three-dimensional ellipsoids to represent the entire 3D scene. Each Gaussian ellipsoid has attributes such as mean, variance, color, and transparency; by projecting these ellipsoids onto the camera plane, the rendered image can be obtained. Compared with NeRF, 3D GS runs faster and has higher accuracy in view synthesis. In terms of 3D reconstruction, since 3D GS is a more explicit representation and each ellipsoid has a center point, the point cloud of the scene can be obtained by extracting the center points of the ellipsoids. However, because 3D GS itself focuses on rendering, this results in the center of each Gaussian ellipsoid not being located on the surface of the scene but being distributed near the surface; this leads to relatively low accuracy in the results obtained by using the original 3D GS for 3D reconstruction.

[0013] Moreover, both the neural radiance field and 3D Gaussian splatting need to obtain the camera poses in advance using other methods (such as SfM). In the case of inaccurate poses, the performance of the above two methods will decline significantly, resulting in reduced accuracy when using both for 3D reconstruction. Summary of the Invention

[0014] To solve the above problems, the objective of the present invention is to provide a method for three-dimensional reconstruction of complex scenes based on multi-source fusion, which realizes the information fusion of multiple sensors based on Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3D GS); the entire scene is encoded in the Signed Distance Function (SDF) of NeRF, and then multi-modal data is associated through the SDF as an intermediate medium; during the training process of NeRF, 3D GS is used as a guidance to explicitly supervise the estimated value of the SDF and guide the sampling of NeRF rays.

[0015] To achieve the above object of the invention, the present invention adopts the following technical solutions:

[0016] A method for three-dimensional reconstruction of complex scenes based on multi-source fusion, which comprises the following steps:

[0017] S1. Collect the lidar point cloud data and camera image data of the target scene, use the lidar point cloud data for pose estimation to obtain the lidar pose information; use the camera image data for pose estimation to obtain the camera pose information;

[0018] S2. Based on the camera image data and camera pose information in step S1, calculate the loss function related to the three-dimensional ellipsoid, which is used to train the attributes of the three-dimensional ellipsoid and optimize the camera pose. The attributes of the three-dimensional ellipsoid include: color, position, transparency, rotation, scale, wherein rotation and scale are combined to obtain the variance of the three-dimensional Gaussian ellipsoid;

[0019] S3. Based on the camera image data and lidar point cloud data in step S1, as well as the camera and lidar poses, calculate the loss function related to the neural radiance field;

[0020] S4. Add all the losses in steps S2 and S3 together to obtain the final loss; perform backpropagation on the final loss to optimize the three-dimensional Gaussian ellipsoid, neural radiance field, and camera and lidar poses;

[0021] S5. Dynamically increase the Gaussian ellipsoid according to the gradient of the final loss propagated to the mean value of the Gaussian ellipsoid;

[0022] S6. If the transparency of a certain Gaussian ellipsoid is lower than 0.001, directly delete the Gaussian ellipsoid;

[0023] S7. Repeat the above steps S2 to S6 at least tens of thousands of times to complete the training and obtain the three-dimensional expression of the entire scene;

[0024] S8. Extract the three-dimensional reconstruction result of the scene according to the trained neural radiance field.

[0025] Furthermore, step S2 described above includes the following sub-steps:

[0026] S2.1. Randomly select an image, project the N Gaussian ellipsoids of the target scene into the image coordinates according to the pose of the image. The projection process includes projecting the mean and variance of the Gaussian ellipsoid respectively to obtain the coordinates of the Gaussian mean in the image plane and the projection of the variance on the two-dimensional plane, and determining the projected Gaussian data.

[0027] S2.2. Render a new image according to the projected Gaussian data. Each pixel p on the image i The rendered color is the weighted average of the colors of the M Gaussian ellipsoids, and N≥M.

[0028] S2.3. Obtain the rendered image After that, calculate the rendering loss between the rendered image and the real image

[0029]

[0030] where SSIM(a, b) represents the structural similarity of two images; λ ssim is a hyperparameter used to balance the two errors.

[0031] S2.4. Assuming there are N Gaussian ellipsoids, first obtain the scales of all Gaussian ellipsoids and take the minimum value τ of the scales of each Gaussian ellipsoid i , and the scale loss function of the Gaussian ellipsoid is expressed as

[0032]

[0033] Furthermore, in step S2.1 described above, the projection process of the Gaussian mean is as follows:

[0034] First, assume [x w , y w , z w T is the mean of the Gaussian ellipsoid in the world coordinate system. Transform the mean to the camera coordinate system, and the transformation formula is

[0035] [x c , y c , z c T =R[x w , y w , z w T +t​​​​

[0036] Among them, [x c , y c , z c T is the coordinate of the center (mean value) of the Gaussian ellipsoid in the camera coordinate system, and (R, t) is the pose of the camera; T represents the transpose;

[0037] Then, calculate the perspective projection matrix M persp

[0038]

[0039] Among them, f x , f y are the focal lengths of the camera in the horizontal and vertical directions of the image respectively; (c x , c y ) are the principal point coordinates of the camera; W and H are the width and height of the image respectively; f and n are the far plane and near plane of the frustum of the pyramid;

[0040] Finally, transform the Gaussian mean into the homogeneous coordinate system and multiply it by the perspective projection matrix M persp to obtain the coordinates (x ndc , y ndc , z ndc ) of the Gaussian mean in the NDC coordinate system; at this time, there are x ndc ∈[-1, 1], y ndc ∈[-1, 1], z ndc ∈[0, 1]; take the first two dimensions (x ndc , y ndc ) of the NDC coordinates and scale them to W and H to obtain the coordinates of the Gaussian mean in the image plane.

[0041] Furthermore, in the above step S2.1, the projection process of the Gaussian variance is as follows:

[0042] First, assume that [x w , y w , z w T is the mean value of the Gaussian ellipsoid in the world coordinate system, and transform the mean value to the camera coordinate system. The transformation formula is

[0043] [x c , y c , z c T =R[x w , y w , z w T +t

[0044] Among them, [x​​​​c , y c , z c T are the coordinates of the center (mean value) of the Gaussian ellipsoid in the camera coordinate system, and (R, t) is the pose of the camera; T represents the transpose;

[0045] Then, calculate the Jacobian matrix J according to the mean value;

[0046]

[0047] Assume that the variance of the Gaussian ellipsoid in the world coordinate system is Σ 3D , and transform the covariance with the Jacobian matrix and the rotation of the camera, Σ' = JRΣ 3D R T J; where R is the rotation of the camera;

[0048] Finally, take the first two rows and the first two columns of Σ' to obtain the projection Σ of the variance on the two-dimensional plane 2D .

[0049] Further, in the above step S2.2, the following sub-steps are included:

[0050] First, arrange the M Gaussian ellipsoids in ascending order of the distance from the camera plane;

[0051] Then, calculate the transparency β i of each Gaussian ellipsoid to the pixel p i,j ;

[0052]

[0053] where β' j is the transparency of the Gaussian ellipsoid itself; μ j is the projection position of the mean value of the Gaussian ellipsoid on the image plane; Σ j is the variance of the Gaussian ellipsoid projected on the image plane; β i,j represents the transparency of the j-th Gaussian ellipsoid to the pixel p i ;

[0054] Finally, after obtaining the transparency β i of each Gaussian ellipsoid j to each pixel p i,j , perform weighted averaging on the color C i,j of each Gaussian ellipsoid according to the weight w j to calculate the final color i of each pixel p

[0055]

[0056] Further, the above step S3 includes the following sub-steps:​

[0057] S3.1. Randomly select one of the Gaussian ellipsoids visible in the current image and emit a ray starting from the optical center and pointing to the center of the Gaussian ellipsoid; uniformly sample n1 points on this ray to achieve coarse sampling, so that the sampling points cover the entire scene; then, calculate the distance d from the Gaussian center point to the camera center, and uniformly sample n2 points within the distance range of [0.9d, 1.1d] to achieve fine sampling, which is mainly used to restore the geometric structure of the scene; make each ray generate (n1 + n2) sampling points; for the (n1 + n2) sampling points, estimate the color and SDF value of the sampling points to obtain the color of the entire ray.

[0058] S3.2. Repeat step S3.1 m times, that is, select m Gaussian ellipsoids within the field of view as the guidance for sampling, and then obtain m rays, with a total of m*(n1 + n2) sampling points; add up the losses of these rays and sampling points to obtain the final Gaussian ray loss and the Gaussian SDF loss

[0059] S3.3. Randomly sample K rays on each image, and the color of each ray is the color of the pixels of the sampling points of the ray.

[0060] S3.4. On each ray, randomly sample Q three-dimensional points in the order from near to far from the camera center, input the three-dimensional points into the neural network, and output the SDF value corresponding to each three-dimensional point and the color of each three-dimensional point Calculate the estimated color of each ray by the method in step S3.1

[0061] Then, calculate the estimated color and the true color C j to obtain the radiance field ray loss between them

[0062]

[0063] S3.5. Randomly select a three-dimensional point p in the lidar point cloud, and the distance from the radar point to the lidar center is d'; emit a ray from the radar center pointing to p, randomly sample Q three-dimensional points on this ray in the order from near to far from the lidar center, input the three-dimensional points into the neural network, and output the SDF value corresponding to each three-dimensional point Similarly, calculate the distance d of each three-dimensional point to the lidar center i ; according to the different positions of the three-dimensional points, calculate different SDF errors;

[0064] S3.6. Repeat step S3.5 for r times, i.e., select r LiDAR points, generate r rays, and add up the losses of these rays to obtain the radiation field SDF loss.

[0065] Furthermore, step S3.1 specifically includes:

[0066] Input the (n1 + n2) three-dimensional points into a neural network, which will output the SDF value and color of each three-dimensional point; arrange these three-dimensional points in the order from near to far from the camera center. Assume the SDF value of a certain point is and the color value is Obtain the weight w of each three-dimensional point through the SDF value i , and the formula is

[0067]

[0068] where is the SDF value of two adjacent three-dimensional sampling points; Sig(*) represents the Sigmoid function; after obtaining the weights, perform weighted averaging on the colors to obtain the color of the entire ray.

[0069]

[0070] Calculate the Gaussian ray loss between the estimated ray color and the true ray color C.

[0071] For the sampling points within the distance range of the ray [0.9d, 1.1d], i.e., n2 fine sampling points, additionally calculate the error between the estimated SDF value and the true SDF value of each point. Among them, the true value s of the SDF i is the distance from each point to the center of the Gaussian ellipsoid; if the sampling point is between the camera optical center and the Gaussian center, then s i is positive; if the sampling point is outside the Gaussian ellipsoid, then s i is negative.

[0072] Through the above definition method, the SDF value is divided into positive and negative, that is, the orientation of the object surface is specified: for the points outside the object surface, the SDF is positive; for the points inside the object surface, the SDF is negative; the Gaussian SDF loss between the estimated SDF value and the true value.

[0073] Furthermore, in step S4 above, add up all the losses in steps S2 and S3 to obtain the final loss.

[0074]

[0075] Among them, λ1, λ2, λ3, λ4, λ5, λ6 are weights;

[0076] is the rendering loss; is the scale loss; is the Gaussian ray loss; is the Gaussian SDF loss; is the radiation field ray loss; is the radiation field SDF loss.

[0077] Furthermore, in the above step S3.5, if the three-dimensional point is far from the radar point, that is, |d i - d'| > 0.3, then the SDF error of this point is

[0078] If the three-dimensional point is close to the radar point, that is, |d i - d'| ≤ 0.3, then the SDF error of this point is Among them, (d' - d i ) is the true value of the SDF of this point.

[0079] Further, the above step S8 includes the following sub-steps:

[0080] S8.1. First, preset the resolution of the Mesh, and divide the entire scene into Y segments along the three coordinate axes respectively. Therefore, the entire scene has Y 3 individual voxel blocks;

[0081] S8.2. Extract the coordinates of the midpoints of each voxel block, input these points into the neural network, and obtain the SDF value of each point;

[0082] S8.3. According to the SDF value of the point, use the marching tetrahedra algorithm, that is, obtain the three-dimensional reconstruction mesh model of the entire scene.

[0083] Due to the adoption of the above technical solution, the present invention has the following advantages:

[0084] The 3D reconstruction method for complex scenes based on multi-source fusion integrates Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS), with NeRF and 3DGS guiding each other; it can implicitly associate multi-modal data in space without explicit cross-modal feature matching; it can directly extract a 3D mesh model (Mesh), which is better for other downstream tasks; at the same time, it obtains the rendering results of new perspectives of the scene and the mesh model of the scene; during the 3D reconstruction process, under the condition of a given initial pose, it can jointly optimize the poses of the camera and lidar and the implicit expression of the scene, obtain more accurate poses and 3D reconstruction results of the scene, achieve higher reconstruction accuracy, and has good popularization and application value. Brief Description of the Drawings

[0085] Figure 1 is the flowchart of the 3D reconstruction method for complex scenes based on multi-source fusion of the present invention;

[0086] Figure 2 is the comparison diagram of the reconstruction results on a single object in one of the embodiments of the present invention;

[0087] Figure 3 is the comparison diagram of the reconstruction results on a single object in the second embodiment of the present invention;

[0088] Figure 4 is the comparison diagram of the reconstruction results in the outdoor scene of autonomous driving in the third embodiment of the present invention;

[0089] Figure 5 is the comparison diagram of the reconstruction results in the outdoor scene of autonomous driving in the fourth embodiment of the present invention. Detailed Description of the Embodiments

[0090] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments.

[0091] As Figure 1 shown, the 3D reconstruction method for complex scenes based on multi-source fusion includes the following specific steps:

[0092] S1. Collect continuous multi-frame point cloud data of the target scene through a lidar; use the point cloud data for pose estimation to obtain lidar pose information; collect continuous multi-frame image data of the target scene through a camera; use the image data for pose estimation to obtain camera pose information;

[0093] S2, based on the camera image data and camera pose information in step S1, calculating a loss function related to the three-dimensional ellipsoid, the loss function is used to train the properties of the three-dimensional ellipsoid and optimize the camera pose, the properties of the three-dimensional ellipsoid include: color, position (mean), transparency, rotation, scale, wherein the rotation and scale are combined to obtain the variance of the three-dimensional Gaussian ellipsoid; comprising the following sub-steps:

[0094] S2.1. Randomly select an image, and project N Gaussian ellipsoids of the target scene to the image coordinates according to the position and posture of the image. The projection process includes projecting the mean and variance of the Gaussian ellipsoids respectively, obtaining the coordinates of the Gaussian mean under the image plane and the projection of the variance on the two-dimensional plane, and determining the projected Gaussian data;

[0095] (1) The projection process of the Gaussian mean is:

[0096] First, assume that [x w ,y w ,z w ] T is the mean of the Gaussian ellipsoid in the world coordinate system. The mean is transformed to the camera coordinate system. The transformation formula is:

[0097] [x c ,y c ,z c ] T =R[x w ,y w ,z w ] T +t

[0098] Among them, [x c ,y c ,z c ] T is the coordinate of the center (mean) of the Gaussian ellipsoid in the camera coordinate system, (R, t) is the position of the camera; T represents transposition;

[0099] Then, calculate the perspective projection matrix M persp

[0100]

[0101] Among them, f x 、f y are the focal lengths of the camera in the horizontal and vertical directions of the image, respectively; (c x 、c y ) is the principal point coordinate of the camera; W and H are the width and height of the image respectively; f and n are the far plane and near plane of the viewing cone respectively;

[0102] Finally, the Gaussian mean is transformed into a homogeneous coordinate system and combined with the perspective projection matrix Mpersp Multiply to obtain the coordinates (x ndc , y ndc , z ndc ) of the Gaussian mean in the NDC coordinate system; at this time, x ndc ∈[-1, 1], y ndc ∈[-1, 1], z ndc ∈[0, 1]; take the first two dimensions (x ndc , y ndc ) of the NDC coordinates and scale them to W and H to obtain the coordinates of the Gaussian mean in the image plane;

[0103] (2) The projection process of the Gaussian variance is as follows:

[0104] First, assume that [x w , y w , z w T is the mean of the Gaussian ellipsoid in the world coordinate system. Transform the mean to the camera coordinate system, and the transformation formula is

[0105] [x c , y c , z c T =R[x w , y w , z w T +t

[0106] where [x c , y c , z c T is the coordinate of the center (mean) of the Gaussian ellipsoid in the camera coordinate system, and (R, t) is the pose of the camera; T represents the transpose;

[0107] Then, calculate the Jacobian matrix J according to the mean;

[0108]

[0109] Assume that the variance of the Gaussian ellipsoid in the world coordinate system is Σ 3D , and use the Jacobian matrix and the rotation of the camera to transform the covariance, Σ' = JRΣ 3D R T J; where R is the rotation of the camera;

[0110] Finally, take the first two rows and the first two columns of Σ' to obtain the projection Σ 2D of the variance on the two-dimensional plane;

[0111] S2.2. Render a new image according to the projected Gaussian data, and for each pixel p on the image​​​​i The rendered color is the weighted average of the colors of M Gaussian ellipsoids, and N ≥ M because the area of each Gaussian projection is finite, and only the Gaussian ellipsoids projected onto the position of each pixel are used when rendering each pixel; it includes the following sub-steps:

[0112] First, arrange the M Gaussian ellipsoids in ascending order of their distances from the camera plane;

[0113] Then, calculate the transparency β i of each Gaussian ellipsoid for pixel p i,j ;

[0114]

[0115] where β' j is the transparency of the Gaussian ellipsoid itself; μ j is the projected position of the mean value of the Gaussian ellipsoid on the image plane; Σ j is the variance of the Gaussian ellipsoid projected on the image plane; β i,j represents the transparency of the j-th Gaussian ellipsoid for pixel p i ;

[0116] Finally, after obtaining the transparency β i of each Gaussian ellipsoid j for each pixel p i,j , perform a weighted average on the color C i,j of each Gaussian ellipsoid according to the weight w j to calculate the final color i of each pixel p

[0117]

[0118] S2.3. Obtain the rendered image After that, calculate the rendering loss between the rendered image and the real image ;

[0119]

[0120] where SSIM(a, b) represents the structural similarity between two images; λ ssim is a hyperparameter used to balance the two errors. Generally, λ ssim = 0.2;

[0121] S2.4. Assume there are N Gaussian ellipsoids. First, obtain the scales of all Gaussian ellipsoids and take the minimum value τ i of the scales of each Gaussian ellipsoid. The scale loss function is expressed as

[0122]

[0123] The purpose of this loss is to minimize the smallest scale of the Gaussian ellipsoid. Since each Gaussian ellipsoid has three scales, by minimizing the smallest scale, the ellipsoid can be made as flat as possible, that is, the ellipsoid becomes a "disk", so that the Gaussian is closely attached to the surface of the scene object;

[0124] S3. Calculate the loss function related to the neural radiance field based on the camera image data and lidar point cloud data in step S1, as well as the camera pose and lidar pose; it includes the following sub-steps:

[0125] S3.1. Randomly select a Gaussian ellipsoid visible in the current image and emit a ray starting from the optical center and pointing to the center of the Gaussian ellipsoid; uniformly sample n1 points on this ray to achieve coarse sampling so that the sampling points cover the entire scene; then, calculate the distance d from the Gaussian center point to the camera center, and uniformly sample n2 points within the distance range of [0.9d, 1.1d] to achieve fine sampling, which is mainly used to restore the geometric structure of the scene; in this way, each ray will generate (n1 + n2) sampling points; the sampling points n1 and n2 are generally selected as powers of 2. Preferably, n1 is generally selected as 256 or 512, and n2 is generally selected as 64 or 128; for the (n1 + n2) sampling points, estimate the color and SDF value of the sampling points to obtain the color of the entire ray. The specific operations include:

[0126] Input the (n1 + n2) three-dimensional points into a neural network, and the network will output the SDF value and color of each three-dimensional point; arrange these three-dimensional points in the order from near to far from the camera center. Assume that the SDF value of a certain point is The color value is Obtain the weight w of each three-dimensional point through the SDF value i The formula is

[0127]

[0128]

[0129] where are the SDF values of two adjacent three-dimensional sampling points; Sig(*) represents the Sigmoid function; after obtaining the weights, perform weighted averaging on the colors to obtain the color of the entire ray

[0130]

[0131] Calculate the Gaussian ray loss between the estimated ray color and the true ray color C

[0132] For the sampling points within the distance range of the ray [0.9d, 1.1d], that is, n2 fine sampling points, additionally calculate the error between the estimated SDF value and the true SDF value for each point, where the true value s of the SDF i is the distance from each point to the center of the Gaussian ellipsoid; if the sampling point is between the camera optical center and the Gaussian center, then s i is a positive number; if the sampling point is outside the Gaussian ellipsoid (i.e., the Gaussian ellipsoid is between the sampling point and the camera optical center), then s i is a negative number;

[0133] By the above definition method, the SDF values are divided into positive and negative, that is, the orientation of the object surface is specified: for points outside the object surface, the SDF is positive; for points inside the object surface, the SDF is negative; the Gaussian SDF loss between the estimated SDF value and the true value

[0134] S3.2. Repeat step S3.1 m times, that is, select m Gaussian ellipsoids within the field of view as the guidance for sampling, and then obtain m rays, with a total of m*(n1 + n2) sampling points; add up the losses of these rays and sampling points to obtain the final Gaussian ray loss and the Gaussian SDF loss

[0135] S3.3. Randomly sample K rays on each image, and the color of each ray is the color of the pixels of the sampling points on the ray;

[0136] S3.4. On each ray, randomly sample Q three-dimensional points in the order from near to far from the camera center, input the three-dimensional points into the neural network, and output the SDF value corresponding to each three-dimensional point and the color of each three-dimensional point Calculate the estimated color of each ray by the method in step S3.1

[0137] Then, calculate the radiance field ray loss between the estimated color and the true color C j

[0138]

[0139] S3.5. Randomly select a three-dimensional point p in the lidar point cloud, and the distance from the lidar point to the lidar center is d'; emit a ray from the lidar center pointing to p, randomly sample Q three-dimensional points on the ray in the order from near to far from the lidar center, input the three-dimensional points into the neural network, and output the SDF value corresponding to each three-dimensional point​ Similarly, calculate the distance d from each 3D point to the lidar center i ; According to the different positions of the 3D points, calculate different SDF errors;

[0140] If the 3D point is far from the lidar point, that is, |d i - d'| > 0.3, then the SDF error of this point is

[0141] If the 3D point is close to the lidar point, that is, |d i - d'| ≤ 0.3, then the SDF error of this point is where (d' - d i ) is the true value of the SDF of this point;

[0142] S3.6. Repeat step S3.5 r times, that is, select r lidar points, generate r rays, and add up the losses of these rays to obtain the radiation field SDF loss

[0143] S4. Add up all the losses in steps S2 and S3 to obtain the final loss

[0144] where λ1, λ2, λ3, λ4, λ5, λ6 are weights used to balance different types of losses;

[0145] are different types of losses, is the rendering loss; is the scale loss; is the Gaussian ray loss; is the Gaussian SDF loss; is the radiation field ray loss; is the radiation field SDF loss;

[0146] Backpropagate the final loss to optimize the 3D Gaussian ellipsoid, neural radiation field, and the poses of the camera and lidar;

[0147] S5. Dynamically increase the Gaussian ellipsoid according to the gradient of the final loss propagated to the Gaussian ellipsoid mean;

[0148] If the gradient of a certain Gaussian mean is large and the scale of this Gaussian is small, then clone this Gaussian ellipsoid once and move the cloned Gaussian ellipsoid according to the gradient of the mean;

[0149] If the gradient of a Gaussian mean is large and the scale of the Gaussian is large, then split the Gaussian ellipsoid into two smaller Gaussian ellipsoids, with the scale of each sub-Gaussian ellipsoid being half of the original scale, and move one of the sub-Gaussians according to the gradient of the mean;

[0150] S6. If the transparency of a Gaussian ellipsoid is lower than 0.001, that is, the Gaussian ellipsoid is very transparent and has little effect on rendering, then directly delete the Gaussian ellipsoid;

[0151] S7. Repeat the above steps S2 to S6 at least tens of thousands of times, that is, complete the training to obtain the three-dimensional representation of the entire scene;

[0152] S8. Extract the three-dimensional reconstruction result of the target scene according to the trained neural radiance field, including the following sub-steps:

[0153] S8.1. First, preset the resolution of the Mesh, that is, divide the entire scene into Y segments along the three coordinate axes respectively. Therefore, the entire scene will have Y 3 voxel blocks; the higher the resolution, the finer the result obtained, but the slower the calculation speed and the larger the volume of the generated file; generally, 256 is selected, that is, divide the entire scene into 256 segments along the three coordinate axes respectively. Therefore, the entire scene will have 256 3 voxel blocks;

[0154] S8.2. Extract the coordinates of the midpoints of each voxel block, input these points into the neural network, and obtain the SDF value of each point;

[0155] S8.3. Use the Marching Cube algorithm according to the SDF value of the point, that is, obtain the three-dimensional reconstruction mesh model (Mesh) of the entire scene.

[0156] From Figure 2 、 3 、4, 5, it can be seen that the three-dimensional reconstruction method of complex scenes based on multi-source fusion of the present invention can relatively completely reconstruct the entire three-dimensional scene with high reconstruction accuracy.

[0157] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Those of ordinary skill in the art should understand that: the technical solutions recorded in the foregoing embodiments can be modified, or some of the technical features can be equivalently replaced; these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A complex scene 3D reconstruction method based on multi-source fusion, characterized by: It includes the following steps: S1. Collecting laser radar point cloud data and camera image data of the target scene, using the laser radar point cloud data to perform pose estimation to obtain laser radar pose information; using the camera image data to perform pose estimation to obtain camera pose information; S2, based on the camera image data and camera pose information in step S1, calculating a loss function related to the three-dimensional ellipsoid, the loss function is used to train the properties of the three-dimensional ellipsoid and optimize the camera pose, the properties of the three-dimensional ellipsoid include: color, position, transparency, rotation, scale, wherein the rotation and scale are combined to obtain the variance of the three-dimensional Gaussian ellipsoid; comprising the following sub-steps: S2.

1. Randomly select an image, and project N Gaussian ellipsoids of the target scene to the image coordinates according to the position and posture of the image. The projection process includes projecting the mean and variance of the Gaussian ellipsoids respectively, obtaining the coordinates of the Gaussian mean under the image plane and the projection of the variance on the two-dimensional plane, and determining the projected Gaussian data; S2.

2. Based on the projected Gaussian data, a new image is rendered. Each pixel p on the image i The rendered color is the weighted average of the colors of M Gaussian ellipsoids, and N ≥ M; S2.

3. Obtaining a rendered image After that, the rendered image and the real image are calculated Error Among them, SSIM(a,b) represents the structural similarity of two images; λ ssim is a hyperparameter used to balance the two errors; S2.

4. Assuming there are N Gaussian ellipsoids, first obtain the scales of all Gaussian ellipsoids and take out the minimum value τ in the scale of each Gaussian ellipsoid i , the loss function of the Gaussian ellipsoid scale is expressed as S3, based on the camera image data and the lidar point cloud data in step S1, as well as the camera and lidar poses, calculating a loss function related to the neural radiation field; S4, adding all the losses in step S2 and step S3 together to obtain the final loss; back-propagating the final loss to optimize the three-dimensional Gaussian ellipsoid, the neural radiation field, and the camera and lidar poses; S5. Dynamically increase the Gaussian ellipsoid according to the gradient of the final loss propagated to the mean of the Gaussian ellipsoid; S6. If the transparency of a Gaussian ellipsoid is lower than 0.001, directly delete the Gaussian ellipsoid; S7, repeating the above steps S2 to S6 at least tens of thousands of times, completing the training, and obtaining a three-dimensional expression of the entire scene; S8. Extract the three-dimensional reconstruction result of the scene based on the trained neural radiation field.

2. The method for 3D reconstruction of complex scenes based on multi-source fusion according to claim 1, characterized in that: In step S2.1, the projection process of the Gaussian mean is: First, assume that [x w ,y w ,z w ] T is the mean of the Gaussian ellipsoid in the world coordinate system. The mean is transformed to the camera coordinate system. The transformation formula is: [x c ,y c ,z c ] T =R[x w ,y w ,z w ] T +t Among them, [x c ,y c ,z c ] T is the coordinate of the center of the Gaussian ellipsoid in the camera coordinate system, (R, t) is the position of the camera; T represents transposition; Then, calculate the perspective projection matrix M persp Among them, f x 、f y are the focal lengths of the camera in the horizontal and vertical directions of the image, respectively; (c x 、c y ) is the principal point coordinate of the camera; W and H are the width and height of the image respectively; f and n are the far plane and near plane of the viewing cone respectively; Finally, the Gaussian mean is transformed into a homogeneous coordinate system and combined with the perspective projection matrix M persp Multiply them to get the coordinates of the Gaussian mean in the NDC coordinate system (x ndc ,y ndc ,z ndc ); At this time, there is x ndc ∈[-1,1],y ndc ∈[-1,1],z ndc ∈[0,1]; take the first two dimensions of the NDC coordinates (x ndc ,y ndc ) and scale it to W and H to obtain the coordinates of the Gaussian mean in the image plane.

3. The method for 3D reconstruction of complex scenes based on multi-source fusion according to claim 1, characterized in that: In step S2.1, the projection process of Gaussian variance is: First, assume that [x w ,y w ,z w ] T is the mean of the Gaussian ellipsoid in the world coordinate system. The mean is transformed to the camera coordinate system. The transformation formula is: [x c ,y c ,z c ] T =R[x w ,y w ,z w ] T +t Among them, [x c ,y c ,z c ] T is the coordinate of the center of the Gaussian ellipsoid in the camera coordinate system, (R, t) is the position of the camera; T represents transposition; Then, the Jacobian matrix J is calculated based on the mean; Assume that the variance of the Gaussian ellipsoid in the world coordinate system is Σ 3D , transform the covariance using the Jacobian matrix and the camera rotation, Σ'=JRΣ 3D R T J; where R is the rotation of the camera; Finally, take the first two rows and columns of Σ' to get the projection of the variance on the two-dimensional plane Σ 2D .

4. The method for 3D reconstruction of complex scenes based on multi-source fusion according to claim 1, characterized in that: The step S2.2 includes the following sub-steps: First, arrange the M Gaussian ellipsoids from near to far according to their distance from the camera plane; Then, calculate each Gaussian ellipsoid for pixel p i Transparency of β i,j ; Among them, β' j is the transparency of the Gaussian ellipsoid itself; μ j is the projection position of the mean of the Gaussian ellipsoid on the image plane; Σ j is the variance of the Gaussian ellipsoid projection on the image plane; β i,j Represents the jth Gaussian ellipsoid for pixel p i transparency; Finally, we get the Gaussian ellipsoid j for each pixel p i Transparency of β i,j Then, according to the weight w i,j For each Gaussian ellipsoid color C j Perform weighted averaging and calculate each pixel p i Final Color 5. The method for 3D reconstruction of complex scenes based on multi-source fusion according to claim 1, characterized in that: The step S3 comprises the following sub-steps: S3.

1. Randomly select one Gaussian ellipsoid visible in the current image and send out a ray starting from the optical center and pointing to the center of the Gaussian ellipsoid; uniformly sample n1 points on this ray to achieve coarse sampling so that the sampling points cover the entire scene; then calculate the distance d from the Gaussian center point to the camera center, uniformly sample n2 points within the distance range of [0.9d, 1.1d] to achieve fine sampling, which is mainly used to restore the geometric structure of the scene; each ray generates (n1+n2) sampling points; for the (n1+n2) sampling points, estimate the color and SDF value of the sampling points to obtain the color of the entire ray; S3.2, repeat step S3.1 m times, that is, select m Gaussian ellipsoids within the field of view as sampling guides, and then obtain m rays, a total of m*(n1+n2) sampling points; add the losses of these rays and sampling points together to obtain the final Gaussian ray loss and Gaussian SDF loss S3.3, randomly sample K rays on each image, the color of each ray being the color of the pixel at the sampling point of the ray; S3.

4. On each ray, randomly sample Q 3D points in the order from near to far from the camera center, input the 3D points into the neural network, and output the SDF value corresponding to each 3D point And the color of each 3D point The estimated color of each ray is calculated by the method in step S3.1 Next, calculate the estimated color and the true value color C j The error between S3.

5. Randomly select a 3D point p in the laser radar point cloud, the distance between the radar point and the laser radar center is d'; emit a ray pointing to p from the radar center, randomly sample Q 3D points on the ray in the order from near to far from the laser radar center, input the 3D points into the neural network, and output the SDF value corresponding to each 3D point Also calculate the distance d from each 3D point to the center of the laser radar i ; Calculate different SDF errors according to the different positions of the 3D points; S3.6, repeat step S3.5 r times, that is, select r lidar points, generate r rays, and add the losses of these rays together to get the final loss 6. The method for 3D reconstruction of complex scenes based on multi-source fusion according to claim 5, characterized in that: The step S3.1 comprises: The (n1+n2) 3D points are input into a neural network, which outputs the SDF value and color of each 3D point. The 3D points are arranged in order from near to far from the camera center. Assuming that the SDF value of a certain point is Color value The weight w of each 3D point is obtained by the SDF value i , the formula is in, is the SDF value of two adjacent 3D sampling points; Sig(*) represents the Sigmoid function; after obtaining the weight, the weighted average of the colors can be used to obtain the color of the entire ray. Calculate the error between the estimated ray color and the true ray color C For the sampling points within the distance range of the ray [0.9d, 1.1d], that is, n2 fine sampling points, the error between the estimated SDF value and the true SDF value of each point is additionally calculated, where the true value of SDF s i is the distance from each point to the center of the Gaussian ellipsoid; if the sampling point is between the camera optical center and the Gaussian center, then s i is a positive number; if the sampling point is outside the Gaussian ellipsoid, then s i is a negative number; Through the above definition, the SDF value is divided into positive and negative, that is, the direction of the object surface is specified: at a point outside the object surface, the SDF is positive; at a point inside the object surface, the SDF is negative; the error between the estimated SDF value and the true value is 7. The method for 3D reconstruction of complex scenes based on multi-source fusion according to claim 5, characterized in that: In step S3.5, if the 3D point is far from the radar point, that is, |d i -d'|>0.3, then the SDF error of this point is If the 3D point is close to the radar point, that is, |d i -d'|≤0.3, then the SDF error of this point is Among them, (d'-d i ) is the true value of the SDF at that point.

8. The method for 3D reconstruction of complex scenes based on multi-source fusion according to claim 1, characterized in that: The step S8 comprises the following sub-steps: S8.

1. First, preset the resolution of the Mesh and divide the entire scene into Y segments along the three coordinate axes. 3 Individual voxel blocks; S8.2, extract the coordinates of the midpoint of each voxel block, input these points into the neural network, and obtain the SDF value of each point; S8.

3. Using the marching tetrahedron algorithm according to the SDF value of the point, a three-dimensional reconstructed mesh model of the entire scene is obtained.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method, system and device in low-illumination scene and medium

    CN116824051A

  • Shielding removing method based on nerve radiation field

    CN116977360A