Three-dimensional reconstruction method, equipment, device, medium and product
By constructing a spatial hash grid and local symbol distance field SDF, combining geometric parameters and MLP parameter optimization, the problem of expressiveness and lightweight in surface reconstruction of large scenes is solved, and an efficient and lightweight three-dimensional reconstruction method is realized.
Patent Information
- Application Number
- CN202510185193.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing technology, in the large-scene surface reconstruction method, it is difficult to achieve high expressiveness and lightweight simultaneously, resulting in huge memory consumption.
The spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local symbol distance fields SDFs are constructed, the global SDF is constructed based on the local SDF, and the mesh surface is extracted through geometric parameters and MLP parameters optimization, combined with the mobile cube algorithm.
It realizes three-dimensional reconstruction that is both expressive and lightweight while efficiently handling large-scale scenarios, and can flexibly represent complex geometric structures, reducing computational and memory overhead.
Smart Images

Figure CN120339493A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a three-dimensional reconstruction method, device, apparatus, medium and product. Background Art
[0002] Reconstructing large scenes from point clouds is a key research topic in the fields of computer vision and graphics, and has great application potential especially in the fields of autonomous driving, autonomous robots and mixed reality. However, traditional methods usually model the scene as a discrete distance grid or an implicit field based on predefined basis functions, and convert it into a 3D field through TSDF fusion technology or finite element analysis. To adapt to large scenes, these methods widely adopt sparse data structures, but still need to allocate a large number of data nodes to capture fine geometry, resulting in huge memory consumption.
[0003] In recent years, learning-based methods represent the scene as a continuously differentiable implicit field through neural networks, and introduce sparse feature grids to enhance scalability and expressiveness. These methods use data-driven neural kernel functions for scene reconstruction, which are more expressive than traditional basis functions. However, although these methods reduce the dependence on a large number of data nodes, they still rely on high-dimensional latent feature vectors to fit fine geometry, thus increasing additional memory consumption.
[0004] Although previous studies have made significant progress in large-scale scene surface reconstruction, most of the existing methods focus on the expressiveness of map representation while ignoring the optimization of memory occupancy. Therefore, developing a map representation method that is both expressive and lightweight is an important direction of current research. Summary of the Invention
[0005] The present invention provides a three-dimensional reconstruction method, device, apparatus, medium and product to solve the defect that the large-scale scene surface reconstruction method in the prior art cannot be both expressive and lightweight.
[0006] The present invention provides a three-dimensional reconstruction method, the method comprising: Constructing a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and a plurality of local signed distance fields SDF, and constructing a global SDF based on the plurality of local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters; Collecting a plurality of sampling points from the spatial hash grid, and determining the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the index of the local SDF intersected by each voxel in the regular grid, and the regular grid is constructed based on a plurality of candidate local SDFs, and the plurality of candidate local SDFs are selected based on the observation range of the sensor for collecting the point cloud of the map scene to be reconstructed; Based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, determine the predicted SDF value corresponding to each sampling point, and based on the true SDF value and the predicted SDF value of each sampling point, update the geometric parameters and MLP parameters until the convergence condition is reached to obtain the target global SDF; Extract the mesh surface from the target global SDF based on the marching cubes algorithm to perform 3D reconstruction on the map scene to be reconstructed.
[0007] According to a 3D reconstruction method provided by the present invention, the method further includes: At every preset number of iterations, remove the first local SDF with an SDF value greater than the preset SDF threshold from the multiple local SDFs, and determine the second local SDF with an absolute value of the gradient on the plane greater than the preset gradient threshold among the multiple local SDFs after removing the first local SDF, where the plane is perpendicular to the normal vector of the support point of the local SDF; In the case where the maximum scale of the second local SDF on the plane is less than the scale threshold, clone a third local SDF along the gradient direction of the second local SDF on the plane; In the case where the maximum scale of the second local SDF on the plane is not less than the scale threshold, split the second local SDF into a fourth local SDF and a fifth local SDF.
[0008] According to a 3D reconstruction method provided by the present invention, the mesh index is constructed in the following manner: Based on the observation range of the sensor that collects the point cloud of the map scene to be reconstructed, determine the candidate local SDFs among the multiple local SDFs; Based on the bounding box ranges corresponding to all the candidate local SDFs, construct a regular grid; the bounding box range is determined based on the geometric parameters of the candidate local SDF and the radius of the base SDF; Determine the local SDF array corresponding to each voxel in the regular grid, where the local SDF array includes all the local SDFs that intersect with the voxel; Based on the instance array corresponding to each voxel, establish the index of the local SDFs intersected by each voxel.
[0009] According to a 3D reconstruction method provided by the present invention, the determining the predicted SDF value corresponding to each sampling point based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point includes: For each target local SDF that intersects with the sampling point, based on the geometric parameters of the target local SDF, determine the local coordinates of the sampling point in the local coordinate system of the target local SDF; Determine the local SDF value corresponding to the local coordinates of the sampling point in the local coordinate system of the target local SDF based on the MLP parameters of the target local SDF; Perform weighted averaging on the local SDF values corresponding to all target local SDFs intersecting with the sampling point to obtain the predicted SDF value corresponding to the sampling point.
[0010] According to a three-dimensional reconstruction method provided by the present invention, updating the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point includes: When the absolute value of the true SDF value of the sampling point is less than the preset truncation distance, update the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point based on the true SDF value and the predicted SDF value of the sampling point; When the absolute value of the true SDF value of the sampling point is not less than the preset truncation distance, update the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point based on the predicted SDF value of the sampling point and the preset truncation distance.
[0011] According to a three-dimensional reconstruction method provided by the present invention, the method further includes: Perform voxel downsampling on the point cloud of the map scene to be reconstructed to obtain the initial support points corresponding to each local SDF; Initialize the base SDF as an implicit plane, and initialize the scaling vectors of the initial support points corresponding to each local SDF; Determine the normal vectors of the initial support points corresponding to each local SDF based on principal component analysis, and initialize the rotation vectors of the initial support points corresponding to each local SDF with the normal vectors.
[0012] The present invention also provides a three-dimensional reconstruction device, and the device includes: A first reconstruction module, configured to construct a spatial hash grid and a plurality of local signed distance fields SDF corresponding to the point cloud of the map scene to be reconstructed, and construct a global SDF based on the plurality of local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters; A second reconstruction module, configured to collect a plurality of sampling points from the spatial hash grid, and determine the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the indices of the local SDFs intersected by each voxel in the regular grid, and the regular grid is constructed based on a plurality of candidate local SDFs, and the plurality of candidate local SDFs are selected based on the observation range of the sensor for collecting the point cloud of the map scene to be reconstructed; A third reconstruction module, configured to determine a predicted SDF value corresponding to each sampling point based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until a convergence condition is reached, so as to obtain a target global SDF; A fourth reconstruction module, configured to extract a mesh surface from the target global SDF based on the marching cubes algorithm for three-dimensional reconstruction of the map scene to be reconstructed.
[0013] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the three-dimensional reconstruction method as described in any one of the above is implemented.
[0014] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the three-dimensional reconstruction method as described in any one of the above is implemented.
[0015] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the three-dimensional reconstruction method as described in any one of the above is implemented.
[0016] The three-dimensional reconstruction method, device, apparatus, medium, and product provided by the present invention construct a spatial hash grid and multiple local signed distance fields (SDFs) corresponding to the point cloud of the map scene to be reconstructed, and construct a global SDF based on the multiple local SDFs. Among them, the local SDF is constructed based on geometric parameters and MLP parameters. In this way, the local SDF is only parameterized by a small multi-layer perceptron (MLP), does not rely on high-dimensional feature vectors, and the state of each local SDF is controlled by geometric parameters, which enables the local SDF to adapt to complex geometries. Multiple sampling points are collected from the spatial hash grid, and the target local SDF corresponding to each sampling point is determined based on the voxel where each sampling point is located and the grid index. Among them, the grid index includes the index of the local SDF intersected by each voxel in the regular grid. In this way, the local SDF intersected with the sampling point can be quickly and parallelly located through the grid index, improving the query efficiency. The geometric parameters and MLP parameters are locally optimized and updated through the true SDF value and the predicted SDF value corresponding to each sampling point, reducing the computational complexity brought by global optimization. Finally, a high-quality mesh surface is extracted from the optimized global SDF, avoiding additional memory overhead. Through technologies such as localized representation, spatial hash grid, grid index, and MLP parameterization, the present invention realizes both expressive and lightweight three-dimensional reconstruction, while efficiently processing large-scale scenes and being able to flexibly represent complex geometric structures. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 is a schematic flowchart of the 3D reconstruction method provided by the present invention; Figure 2 is one of the schematic diagrams of the scene of the 3D reconstruction method provided by the present invention; Figure 3 is another schematic diagram of the scene of the 3D reconstruction method provided by the present invention; Figure 4 is yet another schematic diagram of the scene of the 3D reconstruction method provided by the present invention; Figure 5 is still another schematic diagram of the scene of the 3D reconstruction method provided by the present invention; Figure 6 is a schematic structural diagram of the 3D reconstruction device provided by the present invention; Figure 7 is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the scope of protection of the present invention.
[0020] Figure 1 is a schematic flowchart of the 3D reconstruction method provided by an embodiment of the present invention. As Figure 1 shown, the method includes: Step 110, constructing a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local signed distance fields (SDFs), and constructing a global SDF based on the multiple local SDFs; wherein, the local SDFs are constructed based on geometric parameters and MLP parameters; Here, the point cloud of the map scene to be reconstructed is obtained by a sensor (such as a lidar or a depth camera). To reduce the computational amount, the point cloud can be downsampled (such as voxelization). For example, the point cloud is divided into regular voxel grids, and only one point is retained in each voxel.
[0021] Here, the spatial hashing grid is a three-dimensional data structure for efficiently storing and querying point clouds. It divides the three-dimensional space into regular grid cells (voxels), and each voxel stores the point cloud data associated with it.
[0022] In this embodiment, the appropriate voxel size can be selected according to the density of the point cloud and the complexity of the scene, and the dimension of the grid can be determined. Then, a three-dimensional array (or list) is created, where each element corresponds to a voxel, and each voxel stores the point cloud that intersects it. Finally, the point cloud is traversed, and each point is assigned to the voxel it belongs to, and its index in the grid is calculated through the coordinates of the point.
[0023] Furthermore, based on the spatial hashing grid, multiple local SDFs are initialized and combined into a global SDF to describe the geometric structure of the entire scene.
[0024] In this embodiment, each local SDF describes the geometric shape of a local region in the scene and is constructed from geometric parameters (such as position, rotation vector, scaling vector) and MLP parameters θ.
[0025] In one example, the global SDF is composed of local SDFs, that is: , where each local SDF is derived from a base SDF with a radius , and the base SDF is parameterized by a small MLP with parameters . Each local SDF is anchored at a support point with position , and its shape is controlled by a rotation vector and a scaling vector .
[0026] It should be understood that the base SDF is the basic form of the local SDF, which defines a simple signed distance field. The signed distance field is a function representing geometric shapes that assigns a value to each spatial point, indicating the distance from the point to the nearest geometric surface. If the point is inside the surface, the value is negative; outside the surface, the value is positive; on the surface, the value is zero.
[0027] In this embodiment, the range of the base SDF is restricted to a spherical region with a radius of 3, thus avoiding the complexity of global calculations while ensuring that the local SDF can efficiently represent local geometry. The base SDF is parameterized by a small multi-layer perceptron (MLP) with parameters θ . By parameterizing the base SDF with an MLP, complex geometric shapes can be flexibly represented while avoiding the dependence on high-dimensional feature vectors.
[0028] Step 120: Collect multiple sampling points from the spatial hash grid, and determine the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the index of the local SDF intersected by each voxel in the regular grid, and the regular grid is constructed based on multiple candidate local SDFs among the multiple local SDFs; In this step, multiple sampling points can be collected from the spatial hash grid along the rays emitted by a sensor (such as a camera or lidar). For example, for each ray, starting from the origin, gradually calculate the intersection points of the ray and the grid until the ray leaves the grid range. Select a sampling point at each intersection point. The sampling point can be the intersection point itself or a position near the intersection point. The selection of sampling points can be uniformly distributed or adaptively sampled according to the complexity of the scene. For each sampling point, quickly find the local SDF intersected with it through the grid index.
[0029] Here, the regular grid is constructed based on multiple candidate local SDFs, and the candidate local SDFs are selected according to the observation range of the sensor and are used to cover the key areas in the scene. It should be understood that the observation range of the sensor (such as a camera or lidar) defines which local SDFs are relevant. For example, if the sensor can only observe a part of the scene, then only the local SDFs intersecting with this part of the area will be selected as candidates.
[0030] In this embodiment, record the index of the local SDF intersected by each voxel in the regular grid. For example, use an array or hash table to store the index information of each voxel to ensure that the local SDF related to a certain voxel can be quickly found during query.
[0031] Step 130: Determine the predicted SDF value corresponding to each sampling point based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until the convergence condition is reached to obtain the target global SDF; Here, the true SDF value of the sampling point can be locally calculated by the sampling point along the ray to the surface. This value represents the true distance from the sampling point to the geometric surface and is the target value during the optimization process.
[0032] For each sampling point, calculate its predicted SDF value based on the geometric parameters and MLP parameters of the target local SDF. Update the geometric parameters and MLP parameters of the local SDF corresponding to each sampling point according to the loss between the predicted SDF value and the true SDF value of each sampling point. For example, gradient descent or other optimization algorithms can be used to update the geometric parameters and MLP parameters until the loss converges. Finally, reconstruct the global SDF through the optimized local SDF.
[0033] Step 140: Extract a mesh surface from the target global SDF based on the marching cubes algorithm to perform 3D reconstruction on the map scene to be reconstructed.
[0034] In this embodiment, the global SDF is first voxelized, that is, the three-dimensional space is divided into regular cubic grids. The vertices of each cube store the SDF values. Then, according to the sign of the SDF value, the vertices of each cube are classified as "inside" (SDF value is negative) or "outside" (SDF value is positive). According to the classification of the vertices (inside or outside), the configuration of the cube is determined. According to the configuration, triangles are generated to approximate the isosurface. Finally, the triangles generated by all the cubes are combined to form a complete mesh surface to complete the 3D reconstruction.
[0035] The 3D reconstruction method proposed in this embodiment constructs a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local signed distance fields (SDFs), and constructs a global SDF based on the multiple local SDFs. Among them, the local SDF is constructed based on geometric parameters and MLP parameters. In this way, the local SDF is only parameterized by a small multi-layer perceptron (MLP) and does not rely on high-dimensional feature vectors. The state of each local SDF is controlled by geometric parameters, which enables the local SDF to adapt to complex geometries. Multiple sampling points are collected from the spatial hash grid, and the target local SDF corresponding to each sampling point is determined based on the voxel where each sampling point is located and the grid index. Among them, the grid index includes the index of the local SDF intersected by each voxel in the regular grid. In this way, the local SDF intersected by the sampling point can be quickly and parallelly located through the grid index, improving the query efficiency. The geometric parameters and MLP parameters are locally optimized and updated through the true SDF value and the predicted SDF value corresponding to each sampling point, reducing the computational cost brought by global optimization. Finally, a high-quality mesh surface is extracted from the optimized global SDF, avoiding additional memory overhead. Through technologies such as localized representation, spatial hash grid, grid index, and MLP parameterization, the present invention realizes both expressive and lightweight 3D reconstruction, and can flexibly represent complex geometric structures while efficiently processing large-scale scenes.
[0036] It should be noted that each embodiment of this application can be freely combined, reordered, or executed independently, and does not need to rely on or depend on a fixed execution order.
[0037] In some embodiments, the method further includes: At every preset number of iterations, remove the first local SDF with an SDF value greater than a preset SDF threshold from the multiple local SDFs, and determine a second local SDF among the multiple local SDFs after removing the first local SDF whose absolute value of the gradient on a plane is greater than a preset gradient threshold, where the plane is perpendicular to the normal vector of the support point of the local SDF; When the maximum scale of the second local SDF on the plane is less than the scale threshold, clone a third local SDF along the gradient direction of the second local SDF on the plane; When the maximum scale of the second local SDF on the plane is not less than the scale threshold, split the second local SDF into a fourth local SDF and a fifth local SDF.
[0038] In this embodiment, during the iterative training process, in order to optimize the map representation and reduce unnecessary calculations and memory occupancy, a pruning-expansion strategy is executed every preset number of iterations. Specifically, if the absolute value of the SDF value of the center point of a certain local SDF is greater than a preset threshold dth , it is considered that this local SDF may be located in a region far from the geometric surface, or the geometric information it represents is inaccurate, so it is removed.
[0039] After removing these local SDFs, it may lead to the lack of geometric information in the local area, thus damaging the integrity of the reconstruction. Therefore, in this embodiment, new local SDFs are also dynamically generated in the under-reconstructed areas.
[0040] Specifically, as shown in Figure 2 , for each local SDF, define a plane perpendicular to the normal vector of the support point e . This plane is used to detect the change of the local SDF in this direction. Calculate the gradient e of the local SDF on the plane , and the gradient represents the change rate of the local SDF on this plane. If the absolute value of the gradient is greater than a preset threshold , it means that the local SDF in this area changes greatly and there may be under-reconstruction.
[0041] Next, as shown in Figure 3 , according to the conditions of the gradient and the scale, decide how to generate new local SDFs. Specifically, if the maximum scale e on the plane se is less than the scale threshold sth , then clone a new local SDF along the gradient direction . The cloned local SDF inherits the parameters of the original local SDF and is appropriately adjusted in the gradient direction. If on the plane eThe maximum scale on se is greater than or equal to the scale threshold sth , then the original local SDF is split into two smaller local SDFs. The split local SDFs can better cover complex geometric structures.
[0042] The 3D reconstruction method proposed in this embodiment dynamically optimizes the distribution of local SDFs during the training process in the above manner, which not only ensures the integrity and accuracy of the reconstruction but also improves the efficiency of map representation.
[0043] In some embodiments, the grid index is constructed in the following manner: Based on the observation range of the sensor that acquires the point cloud of the map scene to be reconstructed, determine the candidate local SDFs among the multiple local SDFs; Based on the bounding box ranges corresponding to all the candidate local SDFs, construct a regular grid; the bounding box ranges are determined based on the geometric parameters of the candidate local SDFs and the radius of the base SDF; Determine the local SDF array corresponding to each voxel in the regular grid, and the local SDF array includes all the local SDFs that intersect with the voxel; Based on the instance array corresponding to each voxel, establish the index of the local SDFs intersected by each voxel.
[0044] In this embodiment, those local SDFs that intersect with the sensor observation range are selected from all the local SDFs as candidate local SDFs. These candidate local SDFs are considered relevant to reconstructing the current scene. Each candidate local SDF is simplified to an oriented bounding box, and its shape and size are determined by the geometric parameters (support point position, rotation vector, and scaling vector) and the radius of the base SDF m Determined. By simplifying the local SDF with the bounding box, it can be quickly determined whether the local SDF intersects with a certain area (such as the sensor observation range). Finally, as shown in Figure 4 (a), determine the boundary of the regular grid according to the bounding box ranges of the candidate local SDFs.
[0045] For each local SDF, determine which voxels in the regular grid it intersects with. Assign a key to each local SDF according to the ID of the intersecting voxels for quickly locating the local SDFs that intersect with a certain voxel. Package the parameters of each local SDF into an instance and associate it with the assigned key. Then sort all the instantiated local SDFs according to the assigned key (voxel ID) and save the sorted local SDF instances into an array. For each voxel, record the positions of the first and last local SDF instances that intersect with it. Through sorting and indexing, all the local SDFs that intersect with a certain voxel can be quickly queried, thus accelerating subsequent parallel processing.
[0046] Based on the above steps, for each voxel in the regular grid, a local SDF array can be established to store all the local SDFs that intersect with the voxel. Refer to Figure 4 As shown in (b), through indexing, all the local SDFs that intersect with a certain voxel can be quickly located, thereby achieving efficient query.
[0047] The 3D reconstruction method proposed in this embodiment combines the simplified representation of local SDF and the efficient query ability of the regular grid, which can improve the running efficiency in the 3D reconstruction of large-scale scenes.
[0048] In some embodiments, determining the predicted SDF value corresponding to each sampling point based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point includes: For each target local SDF that intersects with the sampling point, based on the geometric parameters of the target local SDF, determine the local coordinates of the sampling point in the local coordinate system of the target local SDF; Based on the MLP parameters of the target local SDF, determine the local SDF value corresponding to the local coordinates of the sampling point in the local coordinate system of the target local SDF; Perform a weighted average on the local SDF values corresponding to all the target local SDFs that intersect with the sampling point to obtain the predicted SDF value corresponding to the sampling point.
[0049] For each target local SDF that intersects with the sampling point p it is necessary to transform the sampling point from the world coordinate system to the local coordinate system of the local SDF. Specifically, refer to the following formula: p ; ; where is the coordinate of the sampling point p in the local coordinate system; is the inverse of the rotation vector of the target local SDF for aligning directions; is the position of the sampling point relative to the center point of the target local SDF; is the exponential form of the scaling vector to ensure that the scale is always positive.
[0050] In this embodiment, each local SDF is parameterized by an MLP with parameters θ . The MLP is used to calculate the SDF value from the local coordinates. Therefore, the local SDF value can be obtained by querying the basis SDF parameterized by the MLP: , where is the sampling point pAt the target local SDF The SDF value is the SDF function parameterized by the MLP
[0051] Finally, through weighted averaging, combining the SDF values of all target local SDFs intersecting with the sampling point, the final predicted SDF value of the sampling point is obtained: ; where is the weight, used to balance the contributions of different target local SDFs to the final SDF value
[0052] The 3D reconstruction method proposed in this embodiment can quickly and accurately calculate the SDF value of the sampling point in the following way, and is applicable to the 3D reconstruction and optimization of large-scale scenes
[0053] In some embodiments, updating the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point includes: When the absolute value of the true SDF value of the sampling point is less than the preset truncation distance, based on the true SDF value and the predicted SDF value of the sampling point, update the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point; When the absolute value of the true SDF value of the sampling point is not less than the preset truncation distance, based on the predicted SDF value of the sampling point and the preset truncation distance, update the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point
[0054] In this embodiment, the loss calculation between the true SDF value and the predicted SDF value of the sampling point is divided into two cases. Specifically, the first: the absolute value of the true SDF value is less than the preset truncation distance ; In this case, the loss function includes two parts: the difference in SDF value and the regularization term . Here, the difference in SDF value is to measure the absolute difference between the predicted SDF value S and the true SDF value , used to directly optimize the accuracy of the SDF value. The regularization term is used to regularize the gradient of the SDF to ensure that the magnitude of its gradient is close to 1. This is an important property of the signed distance field, that is, the magnitude of the gradient on the surface should be 1. Here, represents the gradient of the predicted SDF value S at the sampling point p , λ is the weight coefficient, used to balance the influence of the regularization term
[0055] Therefore, when the When , the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point are updated through the following loss function: L = .
[0056] The first type: the absolute value of the true SDF value Not less than the preset cutoff distance ; In this case, the absolute value of the true SDF value is large, indicating that the sampling point is far from the surface. Therefore, in order to avoid excessive errors that have a negative impact on the optimization process, the cutoff distance is used in this embodiment. The loss function is calculated as a reference value, and the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point are updated through the following loss function: L = .
[0057] The three-dimensional reconstruction method proposed in this embodiment combines direct optimization of the SDF value and gradient regularization in the above manner, and avoids excessive errors by truncation distance, thereby improving the accuracy and stability of three-dimensional reconstruction.
[0058] In some embodiments, the method further comprises: Perform voxel downsampling on the point cloud of the map scene to be reconstructed to obtain the initial support points corresponding to each local SDF; Initialize the base SDF as an implicit plane, and initialize the scaling vector of the initial support point corresponding to each local SDF; The normal vector of the initial support point corresponding to each of the local SDFs is determined based on principal component analysis, and the normal vector is used to initialize the rotation vector of the initial support point corresponding to each of the local SDFs.
[0059] Specifically, we first define voxels The size of is used to control the resolution of downsampling, divide the point cloud of the map scene to be reconstructed into a regular voxel grid, and select a representative point in each voxel (such as the average point or midpoint in the voxel). The downsampled points are used as the initial support points of the local SDF , these points are the center locations of the local SDF.
[0060] Then, a scaling vector is initialized for each local SDF to control the scale of the local SDF. Specifically, the scaling vector is initialized to , is a scaling factor related to the voxel size.
[0061] In this embodiment, the normal vector of each initial support point is estimated by principal component analysis (PCA)n and correct its direction to point to the sensor that collects the point cloud of the scene to be reconstructed. Then initialize the rotation vector in the following way : , wherein, represents the z axis in the global coordinate system, and the rotation angle α is the dot product of n . Based on the above initialization process, the normal vector n can be aligned to the z axis of the local coordinate system of the local SDF.
[0062] In addition, the base SDF in this embodiment is initialized as an implicit plane. In the local coordinate system, this plane can be expressed as z =0.
[0063] The three-dimensional reconstruction method proposed in this embodiment provides a basis for subsequent optimization and three-dimensional reconstruction through the above initialization process.
[0064] In one example, as shown in Figure 5 , first use the point cloud to reconstruct a hash grid near the geometric surface; then initialize the local SDF to generate an initial global SDF, and this global SDF is constrained near the geometric surface by the constructed hash grid; then, execute the local SDF detection algorithm to quickly detect the local SDF where the sampling points in the hash grid are located; subsequently, through the geometric parameters of the local SDF and the parameters of a small MLP, the predicted SDF value of each sampling point can be obtained S ; finally, construct a loss function between the predicted SDF value S and the true SDF value for training; during the training process, execute a pruning-expansion strategy every certain number of iterations. After training, use the marching cubes algorithm to extract the grid surface from the optimized global SDF to complete the reconstruction.
[0065] The three-dimensional reconstruction method proposed in this embodiment can be regarded as a neural implicit expression method based on basis functions. Compared with the pre-defined basis functions in traditional methods, the base SDF in this embodiment is driven by MLP parameters and thus has more expressiveness. Compared with the neural kernel functions derived from feature vectors in learning-based methods, this embodiment does not rely on feature vectors and thus is more compact.
[0066] Based on any of the above embodiments, the present invention also provides a model training device, Figure 6 which is a schematic structural diagram of the model training provided by the present invention, as shown in Figure 6As shown, the device includes: A first reconstruction module 610, configured to construct a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local signed distance fields (SDFs), and construct a global SDF based on the multiple local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters; A second reconstruction module 620, configured to collect multiple sampling points from the spatial hash grid, and determine a target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the indexes of the local SDFs intersected by each voxel in the regular grid, and the regular grid is constructed based on multiple candidate local SDFs, and the multiple candidate local SDFs are selected based on the observation range of the sensor that collects the point cloud of the map scene to be reconstructed; A third reconstruction module 630, configured to determine a predicted SDF value corresponding to each sampling point based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until a convergence condition is reached, so as to obtain a target global SDF; A fourth reconstruction module 640, configured to extract a grid surface from the target global SDF based on the marching cubes algorithm to perform three-dimensional reconstruction on the map scene to be reconstructed.
[0067] The device provided by the embodiment of the present invention constructs a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local signed distance fields (SDFs), and constructs a global SDF based on the multiple local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters, so that the local SDF is only parameterized by a small multi-layer perceptron (MLP) and does not depend on high-dimensional feature vectors, and the state of each local SDF is controlled by geometric parameters, which enables the local SDF to adapt to complex geometries; collects multiple sampling points from the spatial hash grid, and determines a target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the indexes of the local SDFs intersected by each voxel in the regular grid, so that the local SDF intersected with the sampling point can be quickly and parallelly located through the grid index, improving the query efficiency; locally optimizes and updates the geometric parameters and MLP parameters through the true SDF value and the predicted SDF value corresponding to each sampling point, reducing the computational amount brought by global optimization; finally extracts a high-quality grid surface from the optimized global SDF, avoiding additional memory overhead. The present invention realizes expressive and lightweight three-dimensional reconstruction through technologies such as localized representation, spatial hash grid, grid index, and MLP parameterization, and can flexibly represent complex geometric structures while efficiently processing large-scale scenes.
[0068] The 3D reconstruction device described in this embodiment can be correspondingly referred to the 3D reconstruction method provided by the present invention described above, and will not be specifically described here.
[0069] Figure 7 An exemplary physical structure diagram of an electronic device is shown in Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete communication with each other through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the 3D reconstruction method, and the method includes: Construct a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and a plurality of local signed distance fields SDF, and construct a global SDF based on the plurality of local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters; Collect a plurality of sampling points from the spatial hash grid, and determine the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the index of the local SDF intersected by each voxel in the regular grid, and the regular grid is constructed based on a plurality of candidate local SDFs, and the plurality of candidate local SDFs are selected based on the observation range of the sensor for collecting the point cloud of the map scene to be reconstructed; Based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, determine the predicted SDF value corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until the convergence condition is reached, and obtain the target global SDF; Extract the grid surface from the target global SDF based on the marching cubes algorithm to perform 3D reconstruction on the map scene to be reconstructed.
[0070] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0071] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the three-dimensional reconstruction method provided by the above-mentioned various methods. The method includes: Construct a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local signed distance fields (SDFs), and construct a global SDF based on the multiple local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters; Collect multiple sampling points from the spatial hash grid, and determine the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the indexes of the local SDFs intersected by each voxel in the regular grid, and the regular grid is constructed based on multiple candidate local SDFs, and the multiple candidate local SDFs are selected based on the observation range of the sensor that collects the point cloud of the map scene to be reconstructed; Based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, determine the predicted SDF value corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until the convergence condition is reached to obtain the target global SDF; Extract the mesh surface from the target global SDF based on the marching cubes algorithm to perform three-dimensional reconstruction on the map scene to be reconstructed.
[0072] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the three-dimensional reconstruction method provided by the above-mentioned various methods. The method includes: Construct a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local signed distance fields (SDFs), and construct a global SDF based on the multiple local SDFs; wherein, the local SDFs are constructed based on geometric parameters and MLP parameters. Collect multiple sampling points from the spatial hash grid, and determine the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the index of the local SDF intersected by each voxel in the regular grid, and the regular grid is constructed based on multiple candidate local SDFs, and the multiple candidate local SDFs are selected based on the observation range of the sensor that collects the point cloud of the map scene to be reconstructed. Based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, determine the predicted SDF value corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until the convergence condition is reached to obtain the target global SDF. Extract the grid surface from the target global SDF based on the marching cubes algorithm to perform 3D reconstruction on the map scene to be reconstructed.
[0073] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0074] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical coding features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A three-dimensional reconstruction method, characterized in that, Including: Construct a spatial hash grid corresponding to the point cloud of the map scene to be reconstructed and multiple local signed distance fields (SDFs), and construct a global SDF based on the multiple local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters; Collect multiple sampling points from the spatial hash grid, and determine the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the index of the local SDF intersected by each voxel in the regular grid, and the regular grid is constructed based on multiple candidate local SDFs, and the multiple candidate local SDFs are selected based on the observation range of the sensor for collecting the point cloud of the map scene to be reconstructed; Based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, determine the predicted SDF value corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until the convergence condition is reached to obtain the target global SDF; Extract the grid surface from the target global SDF based on the marching cubes algorithm to perform three-dimensional reconstruction on the map scene to be reconstructed.
2. The three-dimensional reconstruction method according to claim 1, characterized in that, The method further includes: At every preset number of iterations, remove the first local SDF with an SDF value greater than the preset SDF threshold from the multiple local SDFs, and determine the second local SDF with an absolute value of the gradient on the plane greater than the preset gradient threshold among the multiple local SDFs after removing the first local SDF, where the plane is perpendicular to the normal vector of the support point of the local SDF; Clone a third local SDF along the gradient direction of the second local SDF on the plane when the maximum scale of the second local SDF on the plane is less than the scale threshold; Split the second local SDF into a fourth local SDF and a fifth local SDF when the maximum scale of the second local SDF on the plane is not less than the scale threshold.
3. The three-dimensional reconstruction method according to claim 1, wherein The grid index is constructed in the following manner: Based on the observation range of the sensor for collecting the point cloud of the map scene to be reconstructed, determine the candidate local SDFs among the multiple local SDFs; Construct a regular grid based on the bounding box ranges corresponding to all the candidate local SDFs; the bounding box range is determined based on the geometric parameters of the candidate local SDF and the radius of the base SDF; Determine the local SDF array corresponding to each voxel in the regular grid, where the local SDF array includes all the local SDFs intersected by the voxel; Based on the instance array corresponding to each voxel, establish the index of the local SDF intersected by each voxel.
4. The three-dimensional reconstruction method according to claim 1, wherein Determining the predicted SDF value corresponding to each sampling point based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point includes: For each target local SDF intersected by the sampling point, determine the local coordinates of the sampling point in the local coordinate system of the target local SDF based on the geometric parameters of the target local SDF; Determine the local SDF value corresponding to the local coordinates of the sampling point in the local coordinate system of the target local SDF based on the MLP parameters of the target local SDF; Perform weighted averaging on the local SDF values corresponding to all target local SDFs intersecting with the sampling point to obtain the predicted SDF value corresponding to the sampling point.
5. The three-dimensional reconstruction method according to claim 1, wherein Updating the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point includes: When the absolute value of the true SDF value of the sampling point is less than the preset truncation distance, update the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point based on the true SDF value and the predicted SDF value of the sampling point; When the absolute value of the true SDF value of the sampling point is not less than the preset truncation distance, update the geometric parameters and MLP parameters of the target local SDF corresponding to the sampling point based on the predicted SDF value of the sampling point and the preset truncation distance.
6. The three-dimensional reconstruction method according to claim 1, characterized in that, The method further includes: Perform voxel downsampling on the point cloud of the map scene to be reconstructed to obtain the initial support points corresponding to each local SDF; Initialize the base SDF as an implicit plane, and initialize the scaling vectors of the initial support points corresponding to each local SDF; Determine the normal vector of the initial support point corresponding to each local SDF based on principal component analysis, and initialize the rotation vector of the initial support point corresponding to each local SDF with the normal vector.
7. A three-dimensional reconstruction device, characterized in that, Includes: The first reconstruction module is used to construct a spatial hash grid and multiple local signed distance fields SDF corresponding to the point cloud of the map scene to be reconstructed, and construct a global SDF based on the multiple local SDFs; wherein, the local SDF is constructed based on geometric parameters and MLP parameters; The second reconstruction module is used to collect multiple sampling points from the spatial hash grid, and determine the target local SDF corresponding to each sampling point based on the voxel where each sampling point is located and the grid index; wherein, the grid index includes the index of the local SDF intersected by each voxel in the regular grid, and the regular grid is constructed based on multiple candidate local SDFs, and the multiple candidate local SDFs are selected based on the observation range of the sensor for collecting the point cloud of the map scene to be reconstructed; The third reconstruction module is used to determine the predicted SDF value corresponding to each sampling point based on the geometric parameters and MLP parameters of the target local SDF corresponding to each sampling point, and update the geometric parameters and MLP parameters based on the true SDF value and the predicted SDF value of each sampling point until the convergence condition is reached to obtain the target global SDF; The fourth reconstruction module is used to extract the grid surface from the target global SDF based on the marching cubes algorithm to perform three-dimensional reconstruction on the map scene to be reconstructed.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the three-dimensional reconstruction method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional reconstruction method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional reconstruction method according to any one of claims 1 to 6.
Citation Information
Cited By
Triangular patch data management method and device and XR equipment
CN120953544A
Three-dimensional model surface reconstruction method
CN121095465A