A Point Cloud Reconstruction Method and System Based on Autoencoder

By adopting a point cloud reconstruction method based on autoencoder in three-dimensional reconstruction, the dependence problem on texture and lighting in the prior art is solved, high-quality three-dimensional reconstruction is achieved, and the reconstruction speed and generalization are improved.

CN114663600BActive Publication Date: 2025-05-30NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210400962.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-05-30
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

The prior art has problems of relying on texture and lighting in three-dimensional reconstruction, especially in the absence of texture or insufficient lighting, and deep learning-based algorithms lack generalization and cannot process incremental point cloud data.

Method used

The point cloud reconstruction method based on the autoencoder is adopted, and the continuous multi-frame point clouds collected by lidar are divided spatially, and the local directed distance function field is generated using the trained autoencoder network, and the local scene surface is generated through the isosurface extraction algorithm, and the local scene surface is finally spliced ​​into a complete scene.

Benefits of technology

It realizes high-quality reconstruction of complete scenes without relying on texture and lighting, reducing storage overhead, improving reconstruction speed, and having better generalization and the ability to process incremental point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663600B_ABST
    Figure CN114663600B_ABST
Patent Text Reader

Abstract

The present invention relates to a point cloud reconstruction method and system based on an autoencoder. The method includes obtaining current data of a lidar system; performing spatial partitioning on the current data to determine a plurality of data blocks; determining a local distance function field according to the data blocks by using a trained autoencoder network; determining a local scene surface by using an isosurface extraction algorithm according to the local distance function field; and splicing the local scene surfaces to determine a scene surface. The present invention can incrementally reconstruct a high-quality complete scene from a continuous number of frames of point clouds collected by a lidar, and at the same time, the scene storage overhead is small, the reconstruction speed is fast, and it can handle the reconstruction work of most outdoor scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and computer graphics, and particularly to a point cloud reconstruction method and system based on an autoencoder. Background Art

[0002] 3D reconstruction refers to the process of generating a 3D model from 2D image data or 3D point cloud data. 3D reconstruction has very important applications in the fields of computer vision, computer graphics, and robotics, and is the basis for applications such as autonomous driving, terrain generation, and augmented reality.

[0003] 3D reconstruction is mainly divided into two categories. One is 3D reconstruction based on camera images, and the other is 3D reconstruction based on lidar. 3D reconstruction based on images mainly includes several steps such as camera calibration, feature point extraction, pose calculation by feature point matching, depth estimation by dense matching, and surface reconstruction. However, 3D reconstruction based on images is more dependent on texture and lighting, and cannot work in the case of missing texture or insufficient lighting at night. In particular, there will be large errors in the two steps of pose calculation by feature point matching and depth estimation.

[0004] Currently, for 3D reconstruction based on multi-line lidar, the distance perception accuracy is high and it is not affected by lighting changes. The stability and robustness are greatly improved compared to 3D reconstruction based on images. This point cloud reconstruction algorithm is mainly divided into two categories. One is the traditional reconstruction algorithm, including using the iterative closest point algorithm to solve the pose and using the voxel fusion algorithm to reconstruct the surface. This method can achieve incremental 3D reconstruction, but for large-scale scenes, when updating the map, the storage cost increases significantly. The other is the algorithm based on deep learning, which can store the distance function field with shape encoding. However, the existing networks need to train a specific network for a specific object, do not have generalization, and cannot process incremental point cloud data either.

[0005] To solve the problems existing in the prior art in this field, a new point cloud reconstruction method needs to be proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a point cloud reconstruction method and system based on an autoencoder, which can incrementally reconstruct a high-quality complete scene from several consecutive frames of point clouds collected by a lidar, while having a small storage cost for the scene and a fast reconstruction speed, and being able to handle the reconstruction work of most outdoor scenes.

[0007] To achieve the above purpose, the present invention provides the following solution:

[0008] A point cloud reconstruction method based on an autoencoder, comprising:

[0009] Obtain the current data of the lidar system; the current data includes: 3D point cloud data of the current environment obtained by multi-line lidar scanning, spatial coordinates obtained by the global positioning system, and current acceleration obtained by the inertial navigation component;

[0010] Perform spatial partitioning on the current data to determine a plurality of data blocks; the data block is a continuous number of frames of point clouds in the local space-time domain;

[0011] According to the data block, use the trained autoencoder network to determine the local distance function field; the trained autoencoder network takes the data block as the input and the local distance function field as the output; the trained autoencoder network includes: a shape encoder, max pooling, and a shape decoder; the shape encoder is used to encode data blocks at different time series into shape encodings in the geometric shape domain; the max pooling is used to obtain the local spatial shape encoding from the shape encodings of different time series frames; the shape decoder is used to obtain the local spatial distance function field from the local spatial shape encoding;

[0012] According to the local distance function field, use the isosurface extraction algorithm to determine the local scene surface;

[0013] Stitch the local scene surfaces to determine the scene surface.

[0014] Optionally, the step of using the trained autoencoder network to determine the local distance function field according to the data block specifically includes:

[0015] Obtain the historical data of the lidar system;

[0016] Perform spatial partitioning on the historical data;

[0017] Use the traditional voxel fusion algorithm to determine the local distance function field for the partitioned historical data;

[0018] Determine the training set according to the partitioned historical data and the corresponding local distance function field;

[0019] Use the training set to train the autoencoder network to determine the trained autoencoder network.

[0020] Optionally, the loss function of the trained autoencoder network includes:

[0021] The training stage includes: the directed distance function field distance loss function, the directed distance field direction loss function, and the directed distance field occupancy loss function;

[0022] The testing stage includes: the directed distance function field distance loss function, the directed distance field direction loss function, the directed distance field occupancy loss function, and the point cloud to patch loss function.

[0023] Optionally, the determination of the scene surface by stitching the local scene surfaces specifically includes:

[0024] Using the spatial absolute coordinates obtained by the global positioning system as the initial value, integrating the acceleration obtained by the inertial navigation component to obtain the current external sensor parameters;

[0025] Determine the scene surface according to the stitching of the local scene surfaces using the current external sensor parameters.

[0026] A point cloud reconstruction system based on an autoencoder, comprising:

[0027] A current data acquisition module, configured to acquire the current data of the lidar system; the current data includes: 3D point cloud data of the current environment obtained by a multi-line lidar, spatial coordinates obtained by a global positioning system, and the current acceleration obtained by an inertial navigation component;

[0028] A current data partitioning module, configured to perform spatial partitioning on the current data to determine a plurality of data blocks; the data blocks are a continuous number of frames of point clouds in a local spatial and temporal domain;

[0029] A local distance function field determination module, configured to determine a local distance function field according to the data blocks using a trained autoencoder network; the trained autoencoder network takes the data blocks as input and outputs the local distance function field; the trained autoencoder network includes: a shape encoder, a max pooling layer, and a shape decoder; the shape encoder is configured to encode data blocks at different time series into shape encodings in a geometric shape domain; the max pooling layer is configured to obtain local spatial shape encodings from shape encodings of different time series frames; the shape decoder is configured to obtain a local spatial distance function field from the local spatial shape encoding;

[0030] A local scene surface determination module, configured to determine a local scene surface according to the local distance function field using an isosurface extraction algorithm;

[0031] A scene surface determination module, configured to determine the scene surface by stitching the local scene surfaces.

[0032] Optionally, the local distance function field determination module specifically includes:

[0033] A historical data acquisition unit, configured to acquire the historical data of the lidar system;

[0034] A historical data partitioning unit, configured to perform spatial partitioning on the historical data;

[0035] A local distance function field determination unit, configured to determine a local distance function field for the partitioned historical data using a traditional voxel fusion algorithm;

[0036] A training set determination unit, configured to determine a training set according to the divided historical data and the corresponding local distance function field;

[0037] A trained autoencoder network determination unit, configured to use the training set to train an autoencoder network to determine a trained autoencoder network.

[0038] Optionally, the loss function of the trained autoencoder network includes:

[0039] The training stage includes: a directed distance function field distance loss function, a directed distance field direction loss function, and a directed distance field occupancy loss function;

[0040] The testing stage includes: a directed distance function field distance loss function, a directed distance field direction loss function, a directed distance field occupancy loss function, and a point cloud to patch loss function.

[0041] Optionally, the scene surface determination module specifically includes:

[0042] A current sensor extrinsic parameter data determination unit, configured to take the spatial absolute coordinates obtained by the global positioning system as an initial value, integrate the acceleration obtained by the inertial navigation element, and obtain the current sensor extrinsic parameter data;

[0043] A scene surface determination unit, configured to determine a scene surface according to the local scene surfaces stitched together by using the current sensor extrinsic parameter data.

[0044] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0045] A point cloud reconstruction method and system based on an autoencoder provided by the present invention first divides continuous multi-frame point clouds collected by a lidar in the spatial domain, and then generates a local directed distance function field for the local spatial point clouds through an autoencoder network, including a shape encoder that maps local point cloud blocks of different time series frames to corresponding geometric shape domains, a max pooling layer that fuses geometric shape domains of different time series frames, and a shape decoder that maps geometric shape domains to corresponding local directed distance function fields. Then, a local scene surface is generated through an isosurface extraction algorithm, and finally, the local scene surfaces are stitched together into a complete scene. The present invention can generate a high-quality scene surface, and compared with the traditional voxel fusion algorithm, the storage overhead is greatly reduced; at the same time, compared with other neural network-based point cloud reconstruction algorithms, there is no need to train a specific network for a specific object, so it can be more general, with a faster reconstruction speed and richer local details. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 Schematic flowchart of a point cloud reconstruction method based on an autoencoder provided by the present invention;

[0048] Figure 2 Schematic diagram of the principle of a point cloud reconstruction method based on an autoencoder provided by the present invention;

[0049] Figure 3 Schematic diagram of the structure of a point cloud reconstruction system based on an autoencoder provided by the present invention. Detailed implementation manners

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0051] The object of the present invention is to provide a point cloud reconstruction method and system based on an autoencoder, which can incrementally reconstruct a high-quality complete scene from a continuous number of frames of point clouds collected by a lidar, while having a small storage overhead for the scene and a fast reconstruction speed, and being able to handle the reconstruction work of most outdoor scenes.

[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will further describe the present invention in detail with reference to the drawings and specific implementation manners.

[0053] Figure 1 Schematic flowchart of a point cloud reconstruction method based on an autoencoder provided by the present invention, Figure 2 Schematic diagram of the principle of a point cloud reconstruction method based on an autoencoder provided by the present invention, as Figure 1 and Figure 2 shown, the point cloud reconstruction method based on an autoencoder provided by the present invention includes:

[0054] S101, obtaining the current data of the lidar system; the current data includes: three-dimensional point cloud data of the current environment scanned by a multi-line lidar, spatial coordinates obtained by a global positioning system, and current acceleration obtained by an inertial navigation element;

[0055] S102, perform spatial partitioning on the current data to determine multiple data blocks; the data block is a continuous number of frames of point clouds in the local spatio-temporal domain; the data block is a cube with a side length of 10 meters, containing a continuous number of frames of point clouds in the time domain; different time-sequence frames include: when the number of point clouds in a single moment in the local cube space is too small, the system will automatically filter them out, while most of the local cube spaces maintain about 3 consecutive frames of point clouds received.

[0056] S103, according to the data block, use the trained autoencoder network to determine the local distance function field; the trained autoencoder network takes the data block as the input and the local distance function field as the output; the trained autoencoder network includes: a shape encoder, max pooling, and a shape decoder; the shape encoder is used to encode the data blocks at different time sequences into shape encodings in the geometric shape domain; the max pooling is used to obtain the local space shape encoding from the shape encodings of different time-sequence frames; the shape decoder is used to obtain the local space distance function field from the local space shape encoding;

[0057] S103 specifically includes:

[0058] Obtain the historical data of the lidar system;

[0059] Perform spatial partitioning on the historical data;

[0060] Use the traditional voxel fusion algorithm to determine the local distance function field for the partitioned historical data;

[0061] Determine the training set according to the partitioned historical data and the corresponding local distance function field;

[0062] Use the training set to train the autoencoder network to determine the trained autoencoder network.

[0063] The loss function of the trained autoencoder network includes:

[0064] The training stage includes: a signed distance function field distance loss function, a signed distance field direction loss function, and a signed distance field occupancy loss function; the shape decoder D is mainly composed of eight fully connected layers, namely a multi-layer perceptron (MLP), and after each fully connected layer, there is a batch normalization layer (BatchNorm) connected. The last two fully connected layers of the shape decoder D are connected to two outputs, and the signed distance function field is output through the tanh activation layer and the surface occupancy field is output through the sigmoid layer.

[0065] Among them, the constraints of the autoencoder in the training stage include a signed distance function field distance loss function, a signed distance field direction loss function, and a signed distance field occupancy loss function, that is:

[0066] L = ω 1 *LSDF +ω 2 *L sign +ω 3 *L occu ; where, [ω 1 , ω 2 , ω 3 are the corresponding weighting coefficients;

[0067] The distance loss function of the directed distance function field is:

[0068] L sign = sigmoid(-ω 20 *f θ (x)*s);

[0069] where, ω 20 is often set to a large value, such as 10000, to control that when f θ and s have different signs, the constraint is maximal, while when f θ and s have the same sign, the constraint is maximal.

[0070] The occupancy loss function of the directed distance field is:

[0071] L occu = -(O net *logO real + (1 - O net )*log(1 - O real ));

[0072] where, O net is the occupancy probability of vertex x output by the network, and O real ∈ {0, 1} is the true occupancy probability of vertex x.

[0073] Among them, the test phase includes the distance loss function of the directed distance function field, the direction loss function of the directed distance field, the occupancy loss function of the directed distance field, and the point cloud to patch loss function, that is:

[0074] L = ω 1 *L SDF + ω 2 *L sign + ω 3 *L occu + ω 4 *L recon ;

[0075] where, [ω 1 , ω 2 , ω 3 , ω 4 are the corresponding weighting coefficients.

[0076] where, [ω 1 , ω2 , ω 3 , ω 4 are the corresponding weighting coefficients.

[0077] The reconstruction constraints are introduced below, that is, the loss function from point cloud to patch. Since the isosurface extraction is implemented outside the network, this loss is only used to measure the reconstruction effect after the network training is completed, that is:

[0078]

[0079] Among them, M is the reconstructed surface, y is the input true vertex coordinates, and Δ(M, y) calculates the distance from all input point clouds to the scene surface.

[0080] Among them, the specific traditional voxel fusion algorithm is:

[0081] The point clouds of different frames are transformed into the same scale space by the external sensor parameter data. Based on the inverse ray casting algorithm, the signed distance value of the voxel field from the implicit surface and the signed distance value weighted fusion based on Gaussian weights are used to obtain the final distance function field.

[0082] S104. According to the local distance function field, use the isosurface extraction algorithm to determine the local scene surface;

[0083] S105. Stitch the local scene surfaces to determine the scene surface.

[0084] S105 specifically includes:

[0085] Taking the spatial absolute coordinates obtained by the global positioning system as the initial value, integrating the acceleration obtained by the inertial navigation component to obtain the current external sensor parameter data;

[0086] According to the current external sensor parameter data, stitch the local scene surfaces to determine the scene surface.

[0087] Figure 3 is the structural schematic diagram of a point cloud reconstruction system based on an autoencoder provided by the present invention. As Figure 3 shown, a point cloud reconstruction system based on an autoencoder provided by the present invention includes:

[0088] The current data acquisition module 301 is used to acquire the current data of the lidar system; the current data includes: the three-dimensional point cloud data of the current environment scanned by the multi-line lidar, the spatial coordinates obtained by the global positioning system, and the current acceleration obtained by the inertial navigation component;

[0089] The current data partitioning module 302 is used to partition the current data in space to determine a plurality of data blocks; the data blocks are several consecutive frames of point clouds in the local space-time domain;

[0090] The local distance function field determination module 303 is configured to determine a local distance function field according to a data block by using a trained autoencoder network; the trained autoencoder network takes the data block as an input and the local distance function field as an output; the trained autoencoder network includes: a shape encoder, a max pooling, and a shape decoder; the shape encoder is configured to encode data blocks at different time series into shape encodings in a geometric shape domain; the max pooling is configured to obtain local spatial shape encodings from the shape encodings of different time series frames; the shape decoder is configured to obtain a local spatial distance function field from the local spatial shape encodings;

[0091] The local scene surface determination module 304 is configured to determine a local scene surface according to the local distance function field by using an isosurface extraction algorithm;

[0092] The scene surface determination module 305 is configured to splice the local scene surfaces to determine a scene surface.

[0093] Specifically, the local distance function field determination module 303 includes:

[0094] A historical data acquisition unit configured to acquire historical data of a lidar system;

[0095] A historical data partitioning unit configured to perform spatial partitioning on the historical data;

[0096] A local distance function field determination unit configured to determine a local distance function field for the partitioned historical data by using a traditional voxel fusion algorithm;

[0097] A training set determination unit configured to determine a training set according to the partitioned historical data and the corresponding local distance function field;

[0098] A trained autoencoder network determination unit configured to train an autoencoder network by using the training set to determine a trained autoencoder network.

[0099] The loss function of the trained autoencoder network includes:

[0100] The training stage includes: a signed distance function field distance loss function, a signed distance field direction loss function, and a signed distance field occupancy loss function;

[0101] The testing stage includes: a signed distance function field distance loss function, a signed distance field direction loss function, a signed distance field occupancy loss function, and a point cloud to patch loss function.

[0102] Specifically, the scene surface determination module 304 includes:

[0103] The current external sensor parameter data determination unit is configured to integrate the acceleration obtained by the inertial navigation element with the spatial absolute coordinates obtained by the global positioning system as the initial value to obtain the current external sensor parameter data;

[0104] The scene surface determination unit is configured to determine the scene surface according to the local scene surfaces stitched using the current external sensor parameter data.

[0105] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method section.

[0106] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A point cloud reconstruction method based on an autoencoder, characterized in that, it includes: Obtain the current data of the lidar system; The current data includes: 3D point cloud data of the current environment scanned by a multi-line lidar, spatial coordinates obtained by a global positioning system, and current acceleration obtained by an inertial navigation component; Perform spatial partitioning on the current data to determine multiple data blocks; the data blocks are several consecutive frames of point clouds in the local space-time domain; According to the data blocks, use the trained autoencoder network to determine the local distance function field; the trained autoencoder network takes the data blocks as input and the local distance function field as output; the trained autoencoder network includes: a shape encoder, max pooling, and a shape decoder; the shape encoder is used to encode the shapes of data blocks at different time series; the max pooling is used to map the shape encodings of different time series frames to local space shape encodings in the geometric shape domain; the shape decoder is used to obtain the local space distance function field from the local space shape encoding; According to the local distance function field, use an isosurface extraction algorithm to determine the local scene surface; Stitch the local scene surfaces to determine the scene surface; The step of using the trained autoencoder network to determine the local distance function field according to the data blocks specifically includes: Obtain the historical data of the lidar system; Perform spatial partitioning on the historical data; Use a traditional voxel fusion algorithm on the partitioned historical data to determine the local distance function field; Determine a training set according to the partitioned historical data and the corresponding local distance function field; Use the training set to train the autoencoder network to determine the trained autoencoder network.

2. The point cloud reconstruction method based on an autoencoder according to claim 1, characterized in that, The loss function of the trained autoencoder network includes: The training stage includes: a signed distance function field distance loss function, a signed distance field direction loss function, and a signed distance field occupancy loss function; The test stage includes: a signed distance function field distance loss function, a signed distance field direction loss function, a signed distance field occupancy loss function, and a point cloud to patch loss function.

3. The point cloud reconstruction method based on an autoencoder according to claim 1, characterized in that, The step of stitching the local scene surfaces to determine the scene surface specifically includes: Use the spatial absolute coordinates obtained by the global positioning system as the initial value, integrate the acceleration obtained by the inertial navigation component, and obtain the current sensor extrinsic parameter data; Stitch the local scene surfaces according to the current sensor extrinsic parameter data to determine the scene surface.

4. A point cloud reconstruction system based on an autoencoder, characterized in that, it includes: A current data acquisition module for acquiring the current data of the lidar system; The current data includes: 3D point cloud data of the current environment scanned by a multi-line lidar, spatial coordinates obtained by a global positioning system, and current acceleration obtained by an inertial navigation component; A current data partitioning module for performing spatial partitioning on the current data to determine multiple data blocks; the data blocks are several consecutive frames of point clouds in the local space-time domain; The local distance function field determination module is used to determine the local distance function field according to the data block by using the trained autoencoder network; the trained autoencoder network takes the data block as the input and the local distance function field as the output; the trained autoencoder network includes: a shape encoder, a max pooling, and a shape decoder; the shape encoder is used to encode the data blocks at different time series into shape encodings in the geometric shape domain; the max pooling is used to obtain the local spatial shape encoding from the shape encodings of different time series frames; the shape decoder is used to obtain the local spatial distance function field from the local spatial shape encoding; The local scene surface determination module is used to determine the local scene surface according to the local distance function field by using the isosurface extraction algorithm; The scene surface determination module is used to splice the local scene surfaces to determine the scene surface; The local distance function field determination module specifically includes: The historical data acquisition unit is used to acquire the historical data of the lidar system; The historical data partitioning unit is used to perform spatial partitioning on the historical data; The local distance function field determination unit is used to determine the local distance function field for the partitioned historical data by using the traditional voxel fusion algorithm; The training set determination unit is used to determine the training set according to the partitioned historical data and the corresponding local distance function field; The trained autoencoder network determination unit is used to train the autoencoder network by using the training set to determine the trained autoencoder network.

5. The point cloud reconstruction system based on an autoencoder according to claim 4, wherein, The loss function of the trained autoencoder network includes: The training stage includes: the directed distance function field distance loss function, the directed distance field direction loss function, and the directed distance field occupancy loss function; The test stage includes: the directed distance function field distance loss function, the directed distance field direction loss function, the directed distance field occupancy loss function, and the point cloud to patch loss function.

6. The point cloud reconstruction system based on an autoencoder according to claim 4, wherein, The scene surface determination module specifically includes: The current sensor extrinsic parameter data determination unit is used to integrate the acceleration obtained by the inertial navigation element with the spatial absolute coordinates obtained by the global positioning system as the initial value to obtain the current sensor extrinsic parameter data; The scene surface determination unit is used to splice the local scene surfaces according to the current sensor extrinsic parameter data to determine the scene surface.

Citation Information

Patent Citations

  • Real-time three-dimensional reconstruction method and system for large-scale scene based on line-of-sight updating algorithm

    CN107862733A

  • Three-dimensional reconstruction method and device, electronic equipment and storage medium

    CN113487739A