Large-scene point cloud fast up-sampling method based on implicit neural network and spatial hash
Through the method based on implicit neural network and spatial hashing, the problem of low computing efficiency in the existing technology when processing point cloud data with large-scale and complex geometric structures is solved, and efficient and real-time point cloud upsampling is achieved, which is suitable for the processing of large-scale real-world data.
Patent Information
- Application Number
- CN202411860368.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-17
AI Technical Summary
When the prior art processes point cloud data with large-scale and complex geometric structures, it has challenges such as low computing efficiency, reliance on a large number of parameter adjustments, and high hardware resource requirements, making it difficult to effectively process large-scale real-world data.
Using an implicit neural network and spatial hashing method, the fast upsampling of large-scene point clouds is achieved through point cloud implicit surface coding, spatial hashing construction, adaptive spatial feature representation and high-fidelity point cloud generation.
Significantly improves the quality and efficiency of upsampling, supports real-time applications such as autonomous navigation and augmented reality, and can effectively process large-scale real-world data.
Smart Images

Figure CN119963767A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of point cloud upsampling, and in particular to a large scene point cloud fast upsampling method based on implicit neural network and spatial hashing. Background Art
[0002] With the rapid development of 3D scanning technology, LiDAR, depth cameras and other sensors, point cloud data has been widely used in autonomous driving, virtual reality, augmented reality, medical imaging and other fields. However, due to the resolution limitation of the acquisition equipment, view occlusion, motion blur and other factors, the point cloud actually obtained often has problems such as sparseness and irregular distribution, which seriously affects the accuracy and efficiency of subsequent 3D modeling, object recognition and scene understanding. In order to solve this problem, point cloud upsampling has become a key research direction, aiming to improve the resolution and geometric details of the reconstructed model by increasing the number of points and optimizing the distribution of points, while enhancing the continuity and stability of the surface. Although traditional interpolation methods, optimization methods and deep learning-based methods have made certain progress, they still face challenges such as low computational efficiency, reliance on a large number of parameter adjustments and high hardware resource requirements when processing large-scale and complex geometric data. Therefore, developing an efficient, high-quality point cloud upsampling method suitable for large-scale real-world data is of great significance to promote the development of 3D vision technology.
[0003] Although the existing traditional point cloud upsampling methods have solved the problem of sparse and irregular distribution of point clouds to a certain extent, they still have significant limitations in practical applications. Traditional interpolation methods (such as radial basis function RBF and moving least squares MLS) are prone to overfitting or high computational complexity when dealing with complex geometric structures, resulting in low computational efficiency; although optimization-based methods (such as Poisson reconstruction) can generate topologically correct surfaces, their computational cost is high, especially on large-scale data sets, where efficiency becomes a bottleneck, and they rely on normal information, while the normal estimation in actual point clouds may be inaccurate, affecting the reconstruction effect. In addition, although deep learning-based methods (such as DeepSDF and CONet) can learn complex geometric structures, the training process is time-consuming, the hardware resources are high, and most of them focus on explicit surface reconstruction, while relatively few studies are conducted on efficient representation and upsampling of implicit surfaces. These methods generally rely on a large number of parameter adjustments, which increases the difficulty of use and limits their applicability in automated applications, making it difficult to effectively process large-scale real-world data. Summary of the invention
[0004] In order to overcome the shortcomings of the existing technology, the present invention provides a large-scene point cloud fast upsampling method based on implicit neural network and spatial hashing, which is suitable for the fast processing of large-scale point cloud data, especially when processing large-scale, real-world data, can significantly improve the quality and efficiency of upsampling and support real-time applications such as autonomous navigation and augmented reality.
[0005] The technical solution adopted by the present invention to solve its technical problem is:
[0006] A fast upsampling method for large scene point clouds based on implicit neural network and spatial hashing, comprising the following steps:
[0007] Step S1, collecting a large scene point cloud data set, a handheld device scans the real scene to obtain point cloud data, and then normalizes the obtained point cloud data;
[0008] Step S2, point cloud implicit surface encoding: convert the input sparse point cloud into an implicit surface representation, and use a deep learning model to encode the point cloud to capture its intrinsic geometric structure;
[0009] Step S3, spatial hash construction: based on the implicit surface representation of the point cloud, an efficient spatial hash table is established, which is used to accelerate the fast search between the query point and its nearest neighbor;
[0010] Step S4, adaptive spatial feature representation: using the spatial hash table, adaptively represent the spatial features around the query point, and dynamically adjust the feature representation according to the position and local geometric structure of the query point to better adapt to the change of point cloud density in different areas;
[0011] Step S5, high-fidelity point cloud generation: Generate new point cloud data based on the optimized query point position and adaptive spatial feature representation. The newly generated point cloud not only retains the geometric features of the original point cloud, but also has higher resolution and details.
[0012] Furthermore, the method further comprises the following steps:
[0013] Step S6, performance evaluation and optimization: Continuously optimize algorithm parameters and network structure to adapt to different application scenarios and data sets to ensure optimal performance.
[0014] The beneficial effects of the present invention are mainly manifested in: being able to significantly improve the quality and efficiency of upsampling and supporting real-time applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flowchart of a fast upsampling method for large scene point clouds based on implicit neural networks and spatial hashing. DETAILED DESCRIPTION
[0016] The present invention will be further described below in conjunction with the accompanying drawings.
[0017] Reference Figure 1 , a large scene point cloud fast upsampling method based on implicit neural network and spatial hashing, the method comprising the following steps:
[0018] Step S1: Large scene point cloud data set collection: The handheld device scans the real scene to obtain point cloud data, and then normalizes the obtained point cloud data to give a point cloud The hardware used in the present invention is composed of the following: Livox's latest generation 3D laser radar Mid-360 and THUNDERROT-MIX host.
[0019] It is processed through the following process to achieve efficient and high-quality upsampling:
[0020] First, the original point cloud data is meshed to divide the point cloud space into uniform small three-dimensional voxel grids; in this way, the seed points can be densely and uniformly sampled in the three-dimensional space, thereby obtaining the corresponding dense and approximately uniform projection points;
[0021] Next, normalize the original point cloud p and calculate the point cloud ratio p scale =max(X max -X min ,Y max -Y min ,Z max -Z min )and The normalized point cloud is represented as In order to deepen the understanding of point cloud structure, nor The random sampling point is denoted as P rs , use the nearest neighbor search to calculate the distance between each point, denoted as P true ;
[0022] Then, query the neighbor distance d=P of the 51 points around each point true query(P nor ,51), using random sampling and spatial index structure, the distance between each point in the point cloud is calculated; this provides key data for subsequent processing; the more points there are, the more effective the subsequent field construction is.
[0023] Step S2: implicit surface coding of point cloud. The process is as follows:
[0024] There are three methods for point cloud data mapping: Signed Distance Function (SDF), Single Resolution Feature Mesh (Single Resolution Feature Mesh), and Multi-Resolution Feature Mesh.
[0025] The SDF method represents geometric shapes by learning the distance from each point to the nearest surface, and improves the density and quality of point clouds by generating high-resolution point cloud data; the single-resolution feature grid stores feature vectors in a three-dimensional grid with a fixed resolution, which is suitable for processing uniformly distributed data; the multi-resolution feature grid captures features of different scales through multiple layers of grids with different resolutions, which is suitable for processing complex geometric figures and ensures that high-quality point cloud data can be obtained both globally and locally; in the SDF method, the signed distance function SDF is used to represent the scene geometry. For a given three-dimensional point P rs , this continuous function f returns the distance from the point to the nearest surface,
[0026]
[0027] Where p is a three-dimensional point and s is the corresponding SDF value;
[0028] Step S3, spatial hash construction, based on the implicit surface representation of the point cloud, establishes an efficient spatial hash table, which is used to accelerate the fast search between the query point and its nearest neighbor. It divides the three-dimensional space into several cells and assigns each point to the corresponding cell. The process is as follows:
[0029] Step S31, spatial hash construction: The core idea of spatial hash is to divide the three-dimensional space into cubic cells of fixed size. Suppose there is a three-dimensional space And define a cubic cell with a side length of h. For any point p = (x, y, z), the cell where it is located is calculated by the following formula:
[0030]
[0031] in, Indicates the rounding down operation, h is the side length of the cell, which is selected according to the density and geometric complexity of the point cloud; a smaller h can improve the query accuracy but increase memory consumption; a larger h can reduce memory usage but may cause more hash collisions;
[0032] Step S32, dynamically adjust the cell size: In order to adapt to point cloud areas with different densities, a mechanism for dynamically adjusting the cell size can be introduced. The value of h is adaptively adjusted according to the local point cloud density. For each query point P, the number of points N within a certain range (such as radius r) around it is calculated, and the cell size h is dynamically adjusted according to N. The following formula is used:
[0033]
[0034] Where V is the volume of the query region and α is an adjustable parameter that controls the scaling of the cell size;
[0035] Step 4: Adaptive spatial feature representation. The process is as follows:
[0036] Step 4.1, step-by-step optimization strategy: In order to more accurately predict the unsigned distance value and learn more local details, this method proposes a progressive learning strategy, which takes the intermediate results of the mobile query as additional priors. For the original point cloud represented by the surface discrete representation, the closer the query position is to the given point cloud, the smaller the error of searching for the target point on the given point cloud. On this basis, two regions are established: a high-confidence region with small error and a low-confidence region with large error. The query points are sampled in the high-confidence region to help train the network, and the auxiliary points are sampled in the low-confidence region. After the network converges in the current stage, it moves to the estimated surface position through the network gradient, and the moved auxiliary points are used as the surface priority for the next stage.
[0037] Step 4.2, fusion of hierarchical representation of feature information: In previous SDF work, MLP (multi-layer perceptron) is usually used to regress the distribution of the object surface. MLP is a neural network architecture that can learn to map the input spatial position to a value related to the surface distance. However, the traditional SDF method mainly focuses on generating an accurate representation of the object surface, and in this process, the density information is usually not fully considered; the method of the present invention adds density information and makes it related to the spatial position. Considering the density information not only provides information on the surface shape, but also provides the density distribution inside and outside the object, thereby providing a more comprehensive object representation; in order to effectively demonstrate different implicit neural field methods, three different implicit neural field methods are designed: SDF, single-resolution feature grid and multi-resolution feature grid.
[0038] In the SDF method, the scene geometry is represented as a signed distance function (SDF). A signed distance function is a continuous function f that returns the distance from a given 3D point to the nearest surface.
[0039]
[0040] Where x is a 3D point and s is the corresponding SDF value. We use a learnable parameter θ to parameterize the SDF function and investigate several different design choices to represent the function, explicitly as a dense grid of learnable SDF values, hybrid MLP single-resolution or multi-resolution feature grids;
[0041] The most straightforward way to parameterize an SDF is to store the SDF values directly in the discrete volume In each cell of , the resolution is RH×RW×RD. In order to query the SDF value of any point x from the dense SDF grid Using interpolation,
[0042]
[0043] The single-resolution characteristic grid method combines these two parameterization methods using the characteristic-conditioned MLPf θ and resolution R 3 The characteristic grid Φ θ , where each grid cell stores a feature vector instead of directly storing the SDF value;
[0044] Multi-resolution feature grids use a single feature grid Φ θ , you can use a resolution of R l Multi-resolution feature grid Sampling resolution in geometric space, combining features at different frequencies,
[0045]
[0046] Among them, R min and R max are the coarsest and finest resolutions, respectively. With the cubic growth of the total number of grid cells, a fixed number of parameters are used to store the feature grids, and a spatial hash function is used to index the feature vectors at finer levels; each grid contains at most T feature vectors with dimension F, in At the coarse level, the feature grids are densely stored, and at a finer level, i.e. Use the hash function to index the corresponding feature vector,
[0047]
[0048] in, is a bitwise XOR operation, π i is the only large prime number; use the default values Rmin=16,Rmax=2048,L=16,F=2,T=2 19 ,
[0049]
[0050] As the total number of grid cells grows cubically, a fixed number of parameters are used to store the feature grids, and a spatial hash function is used to index the feature vectors;
[0051] Step 5: High-fidelity point cloud generation: After completing the query point position optimization and feature representation, new point cloud data is generated using methods such as upsampling to ensure that the new points are evenly distributed on the surface and add more details; then, post-processing steps such as smoothing and patching holes are performed to further improve the visual quality and integrity of the point cloud;
[0052] Step 6. The design of the loss function is crucial. It not only determines the training direction of the model, but also directly affects the quality and fidelity of the generated point cloud. This method introduces multiple loss functions to ensure that the model can effectively learn implicit surface representation and generate high-quality high-resolution point clouds. The process is as follows:
[0053] Step 6.1, CD loss L1 and CD loss L2;
[0054] Point clouds A and B contain n points and m points respectively. For each point z i ∈A, find the point p that is closest to it j ∈B, and calculate the distance between them For each point p j ∈B, find the point z that is closest to it i ∈A, and calculate the distance between them ChamferL2 is defined as the average of these two sets of distances,
[0055]
[0056] Step 6.2, KL_loss and Geoloss;
[0057] Assume that there are two discrete probability distributions P and Q, which correspond to the distribution of two point clouds. In the calculation of KL divergence between two point clouds, the sample space can be regarded as a set of points. P(i) represents the probability of a point in the first point cloud, and Q(i) represents the probability of the corresponding point in the second point cloud. KL Divergence loss measures the difference between two distributions. The larger the value, the greater the difference between them, and the smaller the value, the smaller the difference between them. Sinkhorn distance is used as a loss function to calculate the distance between two point clouds.
[0058]
[0059] D s (P,Q)=min γ∈U(a,b) <γ,C>;
[0060] D s (P,Q) represents the Sinkhorn distance, γ is a joint probability distribution that satisfies the edge distribution of point cloud P and point cloud Q, U(a,b) is the set of joint probability distributions whose edge distribution is point cloud P and point cloud Q, and C is the cost matrix that measures the loss in going from a point cloud in point cloud P to a point cloud in point cloud Q.
[0061] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.
Claims
1. A fast upsampling method for large scene point clouds based on implicit neural network and spatial hashing, characterized in that: The method comprises the following steps: Step S1, collecting a large scene point cloud data set, a handheld device scans the real scene to obtain point cloud data, and then normalizes the obtained point cloud data; Step S2, point cloud implicit surface encoding: convert the input sparse point cloud into an implicit surface representation, and use a deep learning model to encode the point cloud to capture its intrinsic geometric structure; Step S3, spatial hash construction: based on the implicit surface representation of the point cloud, an efficient spatial hash table is established, which is used to accelerate the fast search between the query point and its nearest neighbor; Step S4, adaptive spatial feature representation: using the spatial hash table, adaptively represent the spatial features around the query point, and dynamically adjust the feature representation according to the position and local geometric structure of the query point to better adapt to the change of point cloud density in different areas; Step S5, high-fidelity point cloud generation: Generate new point cloud data based on the optimized query point position and adaptive spatial feature representation. The newly generated point cloud not only retains the geometric features of the original point cloud, but also has higher resolution and details.
2. The fast upsampling method for large scene point cloud based on implicit neural network and spatial hashing as claimed in claim 1, characterized in that: The method further comprises the following steps: Step S6, performance evaluation and optimization: Continuously optimize algorithm parameters and network structure to adapt to different application scenarios and data sets to ensure optimal performance.
3. The fast upsampling method for large scene point cloud based on implicit neural network and spatial hashing as claimed in claim 1 or 2, characterized in that: The process of step S1 is as follows: First, the original point cloud data is meshed to divide the point cloud space into uniform small three-dimensional voxel grids; in this way, the seed points can be densely and uniformly sampled in the three-dimensional space, thereby obtaining the corresponding dense and approximately uniform projection points; Next, normalize the original point cloud p and calculate the point cloud ratio p scale =max(X max -X min ,Y max -Y min ,Z max -Z min )and The normalized point cloud is represented as In order to deepen the understanding of point cloud structure, from P nor The random sampling point is denoted as P rs , use the nearest neighbor search to calculate the distance between each point, denoted as P true ; Then, query the neighbor distance d=P of the 51 points around each point true query(P nor ,51), uses random sampling and spatial index structure to calculate the distance between points in the point cloud.
4. The large scene point cloud fast upsampling method based on implicit neural network and spatial hashing as claimed in claim 1 or 2, characterized in that: The process of step S2 is as follows: three methods are used for point cloud data mapping: signed distance field (SDF), single-resolution feature grid and multi-resolution feature grid; the SDF method represents geometric shapes by learning the distance from each point to the nearest surface, and improves the point cloud density and quality by generating high-resolution point cloud data; the single-resolution feature grid stores feature vectors in a three-dimensional grid with a fixed resolution, which is suitable for processing uniformly distributed data; the multi-resolution feature grid captures features of different scales through multiple layers of grids with different resolutions, which is suitable for processing complex geometric figures, ensuring that high-quality point cloud data can be obtained both globally and locally.
5. The fast upsampling method for large scene point cloud based on implicit neural network and spatial hashing as claimed in claim 1 or 2, characterized in that: The process of step S3 is as follows: Step S31, spatial hash construction: The core idea of spatial hash is to divide the three-dimensional space into cubic cells of fixed size. Suppose there is a three-dimensional space And define a cubic cell with a side length of h. For any point p = (x, y, z), the cell where it is located is calculated by the following formula: in, represents the rounding down operation, h is the side length of the cell, which is selected according to the density and geometric complexity of the point cloud; Step S32, dynamically adjust the cell size: In order to adapt to point cloud areas with different densities, a mechanism for dynamically adjusting the cell size can be introduced. The value of h is adaptively adjusted according to the local point cloud density. For each query point P, the number of points N within the set range around it is calculated, and the cell size h is dynamically adjusted according to N. The following formula is used: Where V is the volume of the query region and α is an adjustable parameter that controls the scaling of the cell size.
6. The fast upsampling method for large scene point cloud based on implicit neural network and spatial hashing as claimed in claim 1 or 2, characterized in that: The process of step S4 is as follows: Step 4.1, step-by-step optimization strategy: two regions are established: a high confidence region with small error and a low confidence region with large error. Query points are sampled in the high confidence region to help train the network, and auxiliary points are sampled in the low confidence region. After the network converges in the current stage, it is moved to the estimated surface position through the network gradient. The moved auxiliary points are used as the surface priority for the next stage. Step 4.2: Hierarchical representation fusion of feature information.
7. The fast upsampling method for large scene point cloud based on implicit neural network and spatial hashing as claimed in claim 2, characterized in that: In step 6, multiple loss functions are introduced to ensure that the model can effectively learn implicit surface representation and generate high-quality high-resolution point clouds. The process is as follows: Step 6.1, CD loss L1 and CD loss L2; Point clouds A and B contain n points and m points respectively. For each point z i ∈A, find the point p that is closest to it j ∈B, and calculate the distance between them For each point p j ∈B, find the point z that is closest to it i ∈A, and calculate the distance between them ChamferL2 is defined as the average of these two sets of distances, Step 6.2, KL_loss and Geoloss; Assume that there are two discrete probability distributions P and Q, which correspond to the distribution of two point clouds. In the calculation of KL divergence between two point clouds, the sample space can be regarded as a set of points. P(i) represents the probability of a point in the first point cloud, and Q(i) represents the probability of the corresponding point in the second point cloud. D KL Divergence loss measures the difference between two distributions. The larger the value, the greater the difference between them, and the smaller the value, the smaller the difference between them. Sinkhorn distance is used as a loss function to calculate the distance between two point clouds. D s (P,Q)=min γ∈U(a,b) <γ,C>; D s (P,Q) represents the Sinkhorn distance, γ is a joint probability distribution that satisfies the edge distribution of point cloud P and point cloud Q, U(a,b) is the set of joint probability distributions whose edge distribution is point cloud P and point cloud Q, and C is the cost matrix that measures the loss in going from a point cloud in point cloud P to a point cloud in point cloud Q.
Citation Information
Patent Citations
Point cloud fusion method and system based on Riemannian geometric constraint
CN115601494A
Arbitrary resolution point cloud up-sampling method based on neural implicit function
CN116468610A
Border-free scene new view angle synthesis method based on mixed neural radiation field
CN116977536A
Generative adversarial network-based point cloud up-sampling method
CN117011132A
Multi-view neural implicit surface reconstruction method based on point cloud guidance
CN117689747A
Cited By
High-quality implicit surface reconstruction method based on adaptive search radius
CN120997411A
High-quality implicit surface reconstruction method based on adaptive search radius
CN120997411B
Multi-scale point cloud attention defect detection method for key parts of aerospace equipment
CN121305140A
Surface shape measurement method based on implicit neural modeling and meta-learning
CN121708085A
A surface shape measurement method based on implicit neural modeling and meta-learning
CN121708085B