Three-dimensional point cloud visibility rapid prediction method and system based on neural network

By constructing a neural network architecture based on octree convolution U-Net and lightweight MLP, the computational efficiency and accuracy issues of 3D point cloud visibility determination are solved, and efficient and robust visibility prediction is achieved, which is suitable for scenarios such as real-time rendering and robot navigation.

CN120689857AActive Publication Date: 2025-09-23PEKING UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510682384.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-23
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Existing 3D point cloud visibility determination methods have bottlenecks in computational efficiency, robustness, and accuracy. In particular, they perform poorly in the face of noise interference, uneven point density, and concave structures. They also lack differentiability and are difficult to embed into an end-to-end optimization framework.

Method used

A U-Net architecture based on octree convolution is constructed for multi-scale feature extraction, combined with a lightweight multi-layer perceptron (MLP) for visibility prediction. The system efficiency is improved through multi-viewpoint parallel reasoning and dynamic feature reuse mechanism, and a synthetic data-driven end-to-end training strategy is designed to generate real visibility labels.

Benefits of technology

It achieves significant improvements in judgment accuracy and robustness while maintaining computational efficiency, has excellent cross-dataset generalization capabilities, is suitable for scenarios such as real-time rendering and robot navigation, and supports end-to-end optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689857A_ABST
    Figure CN120689857A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional point cloud visibility rapid prediction method and system based on a neural network, and the method comprises the steps: constructing a neural network architecture which comprises a feature extractor and a visibility predictor, and enabling the feature extractor to adopt a 3D U-Net structure, and to be used for extracting view-independent features of a point cloud; the visibility predictor adopts a lightweight multi-layer perceptron and is used for predicting the visibility of each point in the three-dimensional point cloud according to the extracted features and view directions. A multi-view parallel reasoning and dynamic feature multiplexing mechanism is designed, so that the reasoning speed in a multi-view scene is remarkably improved; and designing an end-to-end training strategy driven by synthetic data, and generating training data with a real visibility label. The method has high efficiency, robustness and generalization, can effectively overcome the technical problems of noise, low density and complex geometric structures in three-dimensional point cloud visibility prediction, remarkably improves the accuracy and efficiency of point cloud visibility prediction, and is suitable for application scenes such as real-time rendering, view angle optimization and surface reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer graphics technology and artificial intelligence technology, and relates to a method and system for quickly predicting point cloud visibility based on a neural network. Background Art

[0002] Visibility determination for 3D point clouds is a core problem in computer graphics, 3D vision, and robotics. It aims to determine the visibility of each point in a 3D point cloud from a given viewpoint, providing fundamental support for real-time rendering, scene understanding, and robotic navigation. However, due to the inherent discreteness and unstructured nature of point cloud data, visibility determination presents fundamental challenges: mathematically, the infinitesimal size of a single point renders traditional geometric occlusion-based computational methods ineffective due to the near-zero probability of accurate occlusion determination.

[0003] Existing solutions are mainly divided into two categories: indirect judgment methods based on surface reconstruction, and direct visibility analysis methods based on geometric transformation. The former reconstructs continuous surfaces from point clouds and then performs traditional visibility detection, while the latter, represented by the hidden point removal (HPR) algorithm, directly screens visible points through spherical coordinate transformation and convex hull calculation. This study focuses on the technological breakthrough of the second type of method, aiming to develop an efficient point cloud visibility judgment system that does not require surface reconstruction, and overcome the bottlenecks of traditional methods in terms of computational efficiency and scene adaptability. Compared with methods that rely on complex surface reconstruction, although direct visibility judgment needs to deal with the fuzzy spatial topological relationship caused by the discrete characteristics of point clouds, it avoids the uncertainty and high complexity of the reconstruction process and is more suitable for application scenarios with high real-time requirements; while the HPR method avoids reconstruction but relies on convex hull calculation, the time complexity is still high, and it is difficult to cope with large-scale point clouds and complex geometric structures.

[0004] Although existing methods can achieve certain results under ideal conditions, the accuracy and robustness of visibility determination are still significantly limited when faced with noise interference, uneven point density, and concave structures. Surface reconstruction-based methods are limited by open issues in the reconstruction algorithm itself, such as sensitivity to non-uniform sampling and noise, and the computational complexity increases sharply with the size of the point cloud. Although the HPR method directly operates on the point cloud, the time complexity of the convex hull calculation is on the order of O(nlogn), and it faces a more serious performance bottleneck when the number of points exceeds 10^5. It also has poor processing effect on geometric concave areas with large curvature. In addition, existing methods generally lack differentiability and are difficult to embed into modern learning frameworks for end-to-end optimization.

[0005] In recent years, 3D deep learning technology has made significant progress in the field of point cloud processing. Representative works include voxel-based 3D convolutional networks, sparse voxel networks, and point cloud feature extraction networks (such as PointNet++, PointCNN, O-CNN, etc.). The recent introduction of the Transformer architecture has further improved the modeling capabilities of 3D geometric features. These technologies provide new ideas for point cloud visibility determination: replacing hand-designed geometric rules with data-driven feature learning can effectively overcome the theoretical limitations of traditional methods. However, no research has attempted to apply neural networks to point cloud visibility prediction. In contrast, this study innovatively constructs a 3D U-Net architecture, realizes multi-scale feature extraction through octree convolution, and combines it with a lightweight shared MLP to realize viewpoint-dependent point cloud visibility reasoning, significantly improving the determination accuracy while maintaining computational efficiency.

[0006] Of particular note is that learning-based methods can implicitly learn complex occlusion patterns and geometric priors through massive data, which is difficult to achieve through traditional explicit geometric analysis. This study is the first to combine deep feature learning with viewpoint parameterization to construct an end-to-end visibility determination framework. Experiments have shown that this method surpasses the existing optimal methods in terms of prediction accuracy, noise robustness and computational efficiency, while also having excellent cross-dataset generalization capabilities, providing efficient and reliable visibility priors for downstream tasks such as point cloud rendering, surface reconstruction, and normal estimation. In addition, the differentiable nature of this framework allows it to be seamlessly integrated into the 3D reconstruction optimization pipeline, opening up a new path for the collaborative optimization of point cloud processing and neural networks. Summary of the Invention

[0007] To overcome the shortcomings of the existing technology, the present invention proposes a three-dimensional point cloud visibility prediction technology based on neural networks, designs an efficient and accurate point cloud visibility prediction paradigm, and given a three-dimensional point cloud and several perspectives, can quickly and accurately predict the visibility of each point in the three-dimensional point cloud under the corresponding perspective.

[0008] For the convenience of explanation, the present invention agrees on the following definitions of terms:

[0009] U-Net: A convolution-based deep learning model architecture proposed in the literature (Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. In Intern).

[0010] HPR: A hidden point removal algorithm proposed in the literature (Sagi Katz, Ayellet Tal, and Ronen Basri. 2007. Direct visibility of point sets. In SIGGRAPH.) to calculate the visibility of a point cloud from a given viewpoint.

[0011] First, this paper transforms the point cloud visibility determination problem into a binary classification task given a three-dimensional point cloud and a viewing direction. Using a similar approach to point cloud segmentation, it directly predicts the visibility of each point using a deep learning model. Secondly, to address the inefficiency of repeated computations for a single point cloud under different viewpoints, this paper decouples point cloud feature extraction from the visibility prediction network. A neural network architecture is constructed, consisting of a feature extractor and a visibility predictor. The feature extractor uses a 3D U-Net structure to extract view-independent features from the point cloud. The visibility predictor uses a lightweight multi-layer perceptron (MLP) as input, first calculating the difference between the point cloud coordinates and the viewpoint position and normalizing it to obtain the viewing direction of each point. The viewing direction is then positionally encoded and multiplied by the extracted features. After passing through the MLP, the visibility of each point in the point cloud is predicted. Finally, this paper proposes a data preparation scheme for generating synthetic training data with realistic visibility labels. Data augmentation is used to improve the model's robustness to noise and point cloud density.

[0012] The technical solutions provided by the present invention are as follows:

[0013] The neural network-based 3D point cloud visibility prediction method proposed in this paper includes: constructing a point cloud neural network (feature extractor) based on octree convolution for multi-scale feature extraction of point clouds; constructing a lightweight multi-layer perceptron predictor for predicting point cloud visibility from its feature vectors and viewing direction; designing a multi-viewpoint parallel inference and dynamic feature reuse mechanism for rapid response under multi-viewpoint input conditions of a single point cloud; and designing a synthetic data-driven end-to-end training strategy for rapid generation of training data with real visibility labels. The method includes the following steps:

[0014] 1) Construct an octree-convolution-based U-Net (feature extractor) to perform multi-scale feature extraction. For point cloud data obtained by methods such as lidar, an octree is constructed and octree convolution and interpolation are performed to extract the feature vector corresponding to the original point cloud.

[0015] This paper applies a multi-scale feature extraction method based on an octree convolutional neural network. The core of this method is to use the octree data structure to efficiently encode point clouds, combined with the powerful feature extraction capabilities of the U-Net architecture, to achieve efficient extraction of multi-scale geometric features from point clouds.

[0016] The feature extractor follows the method described in the literature (Wang PS, Liu Y, Guo YX, et al. O-cnn: Octree-based convolutional neural networks for 3D shape analysis [J]. ACM Transactions OnGraphics (TOG), 2017, 36(4): 1-11.), taking a point cloud as input and outputting a feature vector for each point in the point cloud. Specifically, the input point cloud is first normalized to ensure that it is within the unit cube. Then, the unit cube space is divided into voxel grids of different sizes. The point cloud is converted into an octree by recursively subdividing the non-empty voxels occupied by the point cloud until the maximum depth is reached. The finest octree node contains the average 3D spatial coordinates of each point in the corresponding voxel, which serves as the input signal of the U-Net. If there is normal vector information corresponding to the point cloud, we also concatenate the average normal vector and the average coordinates to form the input signal. This octree structure can not only effectively represent the geometric information of the point cloud, but also significantly reduce the computational complexity.

[0017] During feature extraction, octree nodes serve as input signals and are processed by an octree-convolution-based neural network. The entire octree-convolution-based point cloud neural network consists of stacked residual blocks, downsampling and upsampling layers, and skip connections. Each residual block contains two octree-based convolutional layers, interspersed with batch normalization and ReLU activation functions. Through these residual blocks, the network learns the multi-scale geometric features of the point cloud. The decoder restores the spatial resolution through upsampling layers and skip connections, ensuring that the feature map is spatially aligned with the original point cloud. Finally, the octree leaf node features are mapped back to the original point cloud through linear interpolation, generating a viewpoint-independent point-by-point feature vector. This design not only significantly improves feature representation efficiency but also supports the processing of large-scale point clouds, avoiding the computational bottlenecks of traditional surface reconstruction methods.

[0018] Furthermore, the feature extractor of this invention boasts excellent scalability and flexibility. By adjusting the depth of the octree and the number of residual blocks, it can adapt to point cloud data of varying scale and complexity. Furthermore, this module can be combined with other deep learning models to implement more complex point cloud processing tasks, such as point cloud classification, segmentation, and registration. In summary, the multi-scale feature extraction method based on octree convolutional U-Net provides an efficient, robust, and flexible solution for point cloud processing, with broad application prospects.

[0019] 2) Construct a viewpoint-aware lightweight MLP predictor to convert point cloud visibility into a binary classification task based on feature vectors and viewpoint direction; perform multi-layer perceptron (MLP) inference on the feature vectors of the point cloud and the corresponding viewpoint direction to obtain the visibility of each point in the point cloud.

[0020] In point cloud visibility determination, viewpoint information fusion is a key step. This paper proposes a viewpoint-aware lightweight multi-layer perceptron (MLP) predictor. By combining feature vectors with viewpoint direction encoding, this predictor transforms point cloud visibility into a binary classification task based on feature vectors and viewpoint direction, achieving end-to-end binary classification.

[0021] The viewpoint direction is represented by a unit vector and mapped to a high-dimensional space through high-frequency position encoding. Assume that the viewpoint direction is represented by a three-dimensional vector p = (p1, p2, p3). Each component of the viewpoint direction is embedded according to the method in the literature (Mildenhall B, Srinivasan PP, Tancik M, et al. Nerf: Representing scenes as neural radiancefields for view synthesis [J]. Communications of the ACM, 2021, 65 (1): 99-106.):

[0022] γ(p i )=[sin(2 0 πp i ),cos(2 0 πp i ),…,sin(2 L-1 πp i ),cos(2 L-1 πp i )]

[0023] Where L represents the number of frequencies used to encode the viewpoint direction, p i represents the i-th component of the viewpoint direction p, γ(p i) is the embedded vector. The three components of vector p are embedded separately and then concatenated. This high-frequency position encoding can effectively capture subtle changes in viewpoint direction, providing rich viewpoint information for subsequent visibility prediction.

[0024] The encoded viewpoint vector is element-wise multiplied by the feature vector to generate viewpoint-specific features. These feature vectors are input into a lightweight MLP to predict visibility probabilities. The MLP network architecture consists of two hidden layers, each followed by batch normalization and a ReLU activation function. This lightweight design enables the network to accurately predict visibility for each point while maintaining efficient computation.

[0025] The loss function uses cross entropy loss, which is defined as follows:

[0026]

[0027] Among them, y i is the true visibility label, is the predicted visibility probability, and N is the number of points in the point cloud. By minimizing the cross entropy loss, the network can learn the visibility pattern of the point cloud and achieve accurate visibility prediction.

[0028] The lightweight MLP predictor of the present invention has multiple advantages. First, its lightweight design supports multi-viewpoint feature reuse, and a single feature extraction can be adapted to any viewpoint, significantly outperforming the viewpoint-by-viewpoint convex hull calculation of the traditional HPR method. Second, the predictor can effectively fuse viewpoint information and point cloud features to achieve end-to-end visibility determination, avoiding the complex geometric calculations used in traditional methods. Finally, through high-frequency position encoding and a lightweight network structure, the predictor can accurately predict the visibility of point clouds while maintaining efficient computation, showing good robustness and generalization capabilities.

[0029] In practice, the present invention performs multi-viewpoint parallel reasoning and dynamic feature reuse for multi-viewpoint scenarios. For multiple viewpoint positions of the same point cloud, the point cloud feature vectors are cached, and the multiple viewpoint direction vectors are concatenated and predicted to obtain the point cloud visibility from multiple viewpoints.

[0030] The present invention proposes a multi-viewpoint parallel reasoning and dynamic feature reuse mechanism, which significantly improves the efficiency of the system. The core of this mechanism is dynamic feature reuse, that is, the output of the feature extractor (viewpoint-independent features) is cached in the memory. When the viewpoint changes, only the viewpoint encoding needs to be updated and the MLP forward calculation needs to be re-executed, avoiding repeated feature extraction. This design not only reduces the amount of calculation, but also improves the response speed of the system. Specifically, when the viewpoint changes, the system only needs to encode the new viewpoint direction, fuse the encoded viewpoint vector with the cached feature vector, and then perform visibility prediction through MLP. This process avoids re-feature extraction of the entire point cloud, significantly reducing computational overhead.

[0031] For multi-viewpoint scenarios, the present invention supports GPU parallel computing. By batch-inputting the encoding vectors of different viewpoints into MLP, the system can synchronously output the visibility masks of all viewpoints through matrix operations. This parallel computing method fully utilizes the parallel computing capabilities of the GPU and significantly improves the reasoning speed in multi-viewpoint scenarios. Experiments show that the reasoning speed of this mechanism is 58 times faster than that of HPR at a point cloud scale of 130k, and the memory usage is linearly related to the number of points, with good scalability and practicality. The multi-viewpoint parallel reasoning and dynamic feature reuse mechanism provides an efficient, flexible and practical solution for point cloud processing, which significantly improves the performance of the system in multi-viewpoint scenarios.

[0032] 4) In practice, the present invention designs a synthetic data-driven end-to-end training strategy. For a 3D dataset (such as ShapeNet), dataset restoration, point cloud sampling, visibility calculation, and data enhancement are performed to obtain a dataset for network training.

[0033] This paper proposes a data preparation scheme for generating synthetic training data with real visibility labels. We first repair the meshes in ShapeNet using the method proposed in the literature (Peng-Shuai Wang, Yang Liu, and Xin Tong. 2022. Dual OctreeGraph Networks for Learning Adaptive Volumetric Shape Representations. ACM Trans. Graph. (SIGGRAPH) 41, 4 (2022)) to ensure that they are sealed and manifold. Next, we uniformly sample 20,000 points on the mesh surface and then reduce the number of points to 8192 by farthest point sampling. We randomly and uniformly sample viewpoints on the surface of the bounding sphere of the point cloud. For each point in the point cloud, we connect it to the viewpoint to form a line segment and calculate the intersection with all triangular faces. If an intersection occurs (excluding itself), the point is considered invisible; otherwise, it is visible. This process is repeated for all points in the point cloud to generate corresponding visibility labels.

[0034] Finally, the generated point cloud data can be enhanced according to the needs of the actual application scenario, including random addition of Gaussian noise, point density perturbation, and affine transformation. These data enhancement strategies not only increase the diversity of the data but also improve the robustness of the model.

[0035] Based on the above method, the present invention realizes a neural network-based three-dimensional point cloud visibility rapid prediction system, which includes a feature extraction module, a visibility prediction module, a multi-viewpoint reasoning module and a data generation and training module; wherein, the feature extraction module is used to extract multi-scale geometric features from the input three-dimensional point cloud, and provide a basic feature vector for subsequent visibility prediction; the visibility prediction module is used to combine the feature vector and viewpoint direction generated by the feature extraction module to predict the visibility state of each point; the multi-viewpoint reasoning module is used to achieve rapid response under the conditions of multi-viewpoint input of a single point cloud, and improve system efficiency through dynamic feature reuse and parallel computing; the data generation and training module is used to generate synthetic training data with real visibility labels, and optimize the entire system through an end-to-end training strategy to ensure the robustness and generalization ability of the model.

[0036] The synthetic data-driven end-to-end training strategy of the present invention has multiple advantages. First, through synthetic data generation and restoration, large-scale, high-quality annotated data can be obtained, solving the problem of scarce real data. Second, data augmentation methods such as Gaussian noise significantly improve the robustness and diversity of the model. Finally, the training strategy can adapt to point cloud data of different scales and complexities and has good generalization capabilities. In short, the synthetic data-driven end-to-end training strategy provides an efficient, robust, and practical solution for point cloud processing, which has broad application prospects.

[0037] The present invention has the following technical advantages and innovations:

[0038] 1. Efficiency: Octree sparse convolution and feature reuse mechanism achieves O(1) complexity, which is significantly better than HPR's O(n log n) and is suitable for real-time interactive scenarios.

[0039] 2. Robustness: Data-driven feature learning effectively overcomes problems such as noise, low density, and concave structure.

[0040] 3. Generalization: On the ShapeNet dataset, the accuracy loss when generalizing to unseen categories is <2%.

[0041] 4. Differentiability: The entire process is gradient-differentiable and can be embedded in an end-to-end optimization framework. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flowchart of the point cloud visibility prediction method provided by the present invention.

[0043] Figure 2 It is a network structure diagram of the feature extractor of the present invention.

[0044] Figure 3 This is a comparison of the visibility prediction results of the present invention and the existing methods on the ShapeNet dataset.

[0045] Figure 4 This is a comparison of the visibility prediction results of the present invention and the existing methods on a real scan dataset.

[0046] Figure 5 This is a comparison of the visibility prediction results of the present invention and the existing method on noisy point clouds.

[0047] Figure 6 This is the visibility prediction result of the present invention on the residual defect cloud.

[0048] Figure 7 This is a comparison of the prediction efficiency of the present invention and existing methods on point clouds of different scales.

[0049] Figure 8This is the result of applying the present invention to view-dependent surface reconstruction.

[0050] Figure 9 This is the result of applying the present invention to point cloud normal vector prediction and surface reconstruction.

[0051] Figure 10 This is the result of applying the present invention to the selection of the optimal viewing angle.

[0052] Figure 11 This is the result of applying the present invention to shadow rendering. DETAILED DESCRIPTION

[0053] The present invention will be further described below by way of examples in conjunction with the accompanying drawings, but the scope of the present invention is not limited in any way.

[0054] The present invention provides a point cloud visibility prediction method and system based on a neural network, the process of which is as shown in the attached figure. Figure 1 As shown in the figure, it consists of two parts, the first of which is a multi-scale feature extractor based on octree convolution U-Net. Its specific network structure is shown in the attached figure. Figure 2 As shown in the figure, the second part is the viewpoint-aware lightweight MLP predictor, and its specific network structure is shown in the attached figure. Figure 1 As shown in , there are two hidden layers and the number of output channels is 2.

[0055] 1) Construct an octree-based convolution U-Net for multi-scale feature extraction. For point cloud data obtained by methods such as lidar, an octree is constructed and octree convolution and interpolation are performed to extract the feature vector corresponding to the original point cloud.

[0056] Point cloud data, as an important form of 3D data representation, has a wide range of applications in fields such as computer graphics, computer vision, and robotics. However, due to the sparsity and lack of explicit connectivity of point clouds, extracting useful geometric information directly from them is a challenging task. Traditional point cloud processing methods typically rely on surface reconstruction or complex geometric calculations, which are often inefficient when processing large point clouds and have poor robustness to noise and incomplete data.

[0057] To address these issues, this paper applies a multi-scale feature extraction method based on octree convolutional U-Net. The core of this method is to efficiently encode point clouds using the octree data structure, combined with the powerful feature extraction capabilities of the U-Net architecture, to achieve efficient extraction of multi-scale geometric features from point clouds.

[0058] The feature extractor takes a point cloud as input and outputs a feature vector for each point. Specifically, the input point cloud is first normalized to ensure that it is within the unit cube, and then converted into an octree by recursively subdividing non-empty voxels until the maximum depth is reached. The thinnest octree node contains the average coordinates of the points within each corresponding voxel, which serves as the input signal of the U-Net. If normals exist, we also concatenate the average normals with the average coordinates to form the input signal. This octree structure can not only effectively represent the geometric information of the point cloud, but also significantly reduce the computational complexity.

[0059] During feature extraction, octree nodes serve as input signals and are processed by an encoder-decoder U-Net. The entire U-Net consists of stacked residual blocks, downsampling and upsampling layers, and skip connections. Each residual block contains two octree-based convolutional layers, interspersed with batch normalization and ReLU activation functions. Through these residual blocks, the network is able to learn the multi-scale geometric features of the point cloud. The decoder restores the spatial resolution through upsampling layers and skip connections, ensuring that the feature map is spatially aligned with the original point cloud. Finally, an interpolation module maps the octree leaf node features back to the original point cloud, generating a viewpoint-independent point-by-point feature vector. This design not only significantly improves feature representation efficiency but also supports the processing of large-scale point clouds, avoiding the computational bottlenecks of traditional surface reconstruction methods.

[0060] Furthermore, the feature extraction module of the present invention offers excellent scalability and flexibility. By adjusting the depth of the octree and the number of residual blocks, it can adapt to point cloud data of varying scales and complexities. Furthermore, the module can be combined with other deep learning models to implement more complex point cloud processing tasks, such as point cloud classification, segmentation, and registration. In summary, the multi-scale feature extraction method based on octree convolutional U-Net provides an efficient, robust, and flexible solution for point cloud processing, with broad application prospects.

[0061] 2) Construct a viewpoint-aware lightweight MLP predictor and perform MLP inference on the feature vectors of the point cloud and the corresponding viewing direction to obtain the visibility of each point in the point cloud.

[0062] In the task of determining point cloud visibility, the fusion of viewpoint information is a key step. Traditional point cloud processing methods typically use hand-crafted features and complex geometric calculations to determine point visibility. These methods are not only computationally complex but also have poor robustness to noise and incomplete data. To address these issues, this paper proposes a viewpoint-aware lightweight multi-layer perceptron (MLP) predictor. By combining feature vectors with viewpoint direction encoding, this predictor converts point cloud visibility into a binary classification task based on feature vectors and viewpoint direction, achieving end-to-end binary classification.

[0063] The viewpoint direction is represented by a unit vector and mapped to a high-dimensional space through high-frequency position encoding. Assume that the viewpoint direction is represented by a three-dimensional vector p = (p1, p2, p3). Each component of the viewpoint direction is embedded as follows:

[0064] γ(p i )=[sin(2 0 πp i ),cos(2 0 πp i ),…,sin(2 L-1 πp i ),cos(2 L-1 πp i )]

[0065] Where L represents the number of frequencies used to encode the viewpoint direction. The three components of vector p are embedded separately and then concatenated. This high-frequency position encoding effectively captures subtle changes in viewpoint direction, providing rich viewpoint information for subsequent visibility prediction.

[0066] The encoded viewpoint vector is element-wise multiplied by the feature vector to generate viewpoint-specific features. These feature vectors are input into a lightweight MLP to predict visibility probabilities. The MLP network architecture consists of two hidden layers, each followed by batch normalization and a ReLU activation function. This lightweight design enables the network to accurately predict visibility for each point while maintaining efficient computation.

[0067] The loss function uses cross entropy loss, which is defined as follows:

[0068]

[0069] Among them, y i is the true visibility label, is the predicted visibility probability, and N is the number of points in the point cloud. By minimizing the cross entropy loss, the network can learn the visibility pattern of the point cloud and achieve accurate visibility prediction.

[0070] The lightweight MLP predictor of the present invention has multiple advantages. First, its lightweight design supports multi-viewpoint feature reuse, and a single feature extraction can be adapted to any viewpoint, significantly outperforming the viewpoint-by-viewpoint convex hull calculation of the traditional HPR method. Second, the predictor can effectively fuse viewpoint information and point cloud features to achieve end-to-end visibility determination, avoiding the complex geometric calculations used in traditional methods. Finally, through high-frequency position encoding and a lightweight network structure, the predictor can accurately predict the visibility of point clouds while maintaining efficient computation, showing good robustness and generalization capabilities.

[0071] 3) Perform multi-viewpoint parallel reasoning and dynamic feature reuse. For multiple viewpoint positions of the same point cloud, the point cloud feature vectors are cached, and the multiple view direction vectors are spliced ​​and predicted to obtain the point cloud visibility from multiple viewpoints.

[0072] In practical applications, point cloud data often needs to be processed from multiple viewpoints, such as in real-time robot navigation and autonomous driving. Traditional point cloud processing methods in multi-view scenarios typically require separate feature extraction and calculation for each viewpoint, which is computationally complex and difficult to meet real-time requirements. To address these issues, this paper proposes a multi-viewpoint parallel inference and dynamic feature reuse mechanism, significantly improving system efficiency.

[0073] The core of this mechanism is dynamic feature reuse, that is, the output of the feature extraction module (viewpoint-independent features) is cached in memory. When the viewpoint changes, only the viewpoint encoding needs to be updated and the MLP forward calculation needs to be re-executed, avoiding repeated feature extraction. This design not only reduces the amount of calculation, but also improves the response speed of the system. Specifically, when the viewpoint changes, the system only needs to encode the new viewpoint direction, fuse the encoded viewpoint vector with the cached feature vector, and then perform visibility prediction through MLP. This process avoids re-extracting features from the entire point cloud, significantly reducing computational overhead.

[0074] For multi-viewpoint scenarios, the present invention supports GPU parallel computing. By batch-inputting the encoding vectors of different viewpoints into MLP, the system can synchronously output the visibility masks of all viewpoints through matrix operations. This parallel computing method fully utilizes the parallel computing capabilities of the GPU and significantly improves the reasoning speed in multi-viewpoint scenarios. Experiments show that the reasoning speed of this mechanism is 58 times faster than that of HPR at a point cloud scale of 130k, and the memory usage is linearly related to the number of points, with good scalability and practicality. The multi-viewpoint parallel reasoning and dynamic feature reuse mechanism provides an efficient, flexible and practical solution for point cloud processing, which significantly improves the performance of the system in multi-viewpoint scenarios.

[0075] The implementation method of dynamic feature reuse and parallel reasoning is:

[0076] Establish a feature buffer pool to store the viewpoint-independent feature matrix F∈R (N×D) ;

[0077] When processing K new viewpoints {v1,...,v K}, execute:

[0078] (1) Parallel calculation of all viewpoint codes γ(v K );

[0079] (2) Broadcast the feature matrix F to K copies, and multiply each viewpoint code element by element to obtain F K ;

[0080] (3) F K Reshape into a (K×N)×D tensor and obtain a K×N×2-dimensional prediction score through a single MLP forward calculation;

[0081] (4) Take argmax of the last dimension to obtain a K×N dimensional visibility matrix to achieve multi-viewpoint parallel reasoning.

[0082] In practice, the present invention designs a synthetic data-driven end-to-end training strategy. For the 3D ShapeNet dataset, dataset restoration, point cloud sampling, visibility calculation, and data enhancement are performed to obtain a dataset for network training.

[0083] One challenge that remains is how to obtain ground-truth visibility labels for the point clouds used to train the network. Although there are some freely available 3D datasets, such as ShapeNet, the meshes in these datasets are usually of low quality, have holes or self-intersection problems, and do not contain visibility annotations.

[0084] To address this issue, we propose a data preparation scheme for generating synthetic training data with real visibility labels. We first use the method proposed in the paper (Peng-Shuai Wang, Yang Liu, and Xin Tong. 2022. Dual Octree Graph Networks for Learning Adaptive Volumetric ShapeRepresentations. ACM Trans. Graph. (SIGGRAPH) 41, 4 (2022)) to repair the meshes in ShapeNet to ensure that they are sealed and manifold. Next, we uniformly sample 20,000 points on the mesh surface and then reduce the number of points to 8192 by farthest point sampling. We randomly and uniformly sample viewpoints on the surface of the bounding sphere of the point cloud. For each point in the point cloud, we connect it to the viewpoint to form a line segment and calculate the intersection with all triangular faces. If an intersection occurs (excluding itself), the point is considered invisible; otherwise, it is visible. This process is repeated for all points in the point cloud to generate the corresponding visibility labels.

[0085] Finally, the generated point cloud data can be enhanced according to the needs of the actual application scenario, including random addition of Gaussian noise, point density perturbation, and affine transformation. These data enhancement strategies not only increase the diversity of the data but also improve the robustness of the model.

[0086] Synthetic data generation strategies include:

[0087] Mesh repair stage: fill holes and correct non-manifold edges on the original mesh to ensure surface closure;

[0088] Point cloud sampling stage: uniformly sample the repaired mesh surface and downsample to the target number N by sampling the farthest point;

[0089] Viewpoint sampling stage: uniformly sample the surface of the bounding sphere with a radius of R to generate M viewpoints, each of which points to the center of the sphere;

[0090] Visibility annotation stage: For point p∈P and viewpoint v∈V, calculate the nearest intersection of ray pv and the grid. If the Euclidean distance between the intersection point and p is greater than the threshold, it is marked as invisible;

[0091] Data enhancement stage: random Gaussian noise and density perturbation are applied to the point cloud.

[0092] The specific implementation method includes the following steps:

[0093] 1) Training data processing and enhancement.

[0094] In this example, we use the ShapeNet dataset as the main data source. ShapeNet is a large-scale 3D model library that contains object models of various categories. We selected 13 categories of object models from ShapeNet, totaling approximately 50,000 models. To generate point cloud data, we first repaired the surface of each model to ensure that it was a closed and manifold mesh. Then, we uniformly sampled 20,480 points on the surface of each model and reduced it to 8,192 points through farthest point sampling (FPS). In addition, we also obtained real scanned point cloud data from the Google ScannedObjects dataset to verify the generalization ability of the model in real scenes. The dataset, number of models, and number of sampling points can be adjusted according to the application scenario and are not limited to the scope selected in this example.

[0095] To improve the robustness of the model, we perform various data augmentation operations on the point cloud data during training, including: adding random noise with a size not exceeding 2% of the point cloud bounding box; point density perturbation, randomly selecting 2048 to 8192 points as input; and incompleteness simulation, randomly selecting three viewpoints and removing points that are not visible from any of these three viewpoints.

[0096] 2) Network Structure Construction and Parameter Selection: The maximum depth of the octree is set to 8; the U-Net feature extractor is divided into 4 levels, with a convolution kernel size of 3*3; the input layer viewpoint direction encoding dimension of the MLP predictor is 63 (frequency L = 10), and the two hidden layer sizes are set to 128 and 64, respectively.

[0097] 3) Use the processed training data to train the above network, using the AdamW optimizer, with an initial learning rate of 1×10^-4 and a 10-fold decay every 30 epochs.

[0098] 4) Use rendering tools such as Blender to render the prediction results into images at the corresponding perspective.

[0099] During testing, that is, when the user uses it, the point cloud and several viewpoint positions are input. The present invention sends the normalized point cloud and viewing direction into the system, and obtains the corresponding point cloud visibility prediction result through the feature extractor and lightweight predictor, and can further obtain the rendering result through the rendering tool.

[0100] Based on the above method, the present invention realizes a neural network-based three-dimensional point cloud visibility rapid prediction system, which includes a feature extraction module, a visibility prediction module, a multi-viewpoint reasoning module and a data generation and training module; wherein, the feature extraction module is used to extract multi-scale geometric features from the input three-dimensional point cloud, and provide a basic feature vector for subsequent visibility prediction; the visibility prediction module is used to combine the feature vector and viewpoint direction generated by the feature extraction module to predict the visibility state of each point; the multi-viewpoint reasoning module is used to achieve rapid response under the conditions of multi-viewpoint input of a single point cloud, and improve system efficiency through dynamic feature reuse and parallel computing; the data generation and training module is used to generate synthetic training data with real visibility labels, and optimize the entire system through an end-to-end training strategy to ensure the robustness and generalization ability of the model.

[0101] Attachment Figure 3 This is a comparison of some of the results of visibility prediction on the ShapeNet dataset between the present invention and existing methods. The present invention shows a significant improvement in the accuracy of visibility prediction, especially in high curvature areas, where the HPR method often produces misjudgments (false negatives). It is worth noting that the method of the present invention can effectively capture the subtle details of complex shapes and provide more reliable visibility predictions in challenging areas. This highlights the strong robustness of the method of the present invention in handling complex geometric shapes, noise and occlusion, and has obvious advantages over the HPR method.

[0102] Attachment Figure 4This figure compares the visibility prediction results of our method with those of existing methods on a real-world scan dataset. Objects in these real-world point clouds exhibit greater complexity than those in the ShapeNet dataset. Despite being trained on ShapeNet, our method is able to produce accurate visibility predictions for these real-world point clouds, demonstrating its strong robustness and generalization capabilities. Compared to the results of the HPR method, our method consistently outperforms the HPR method in areas of high curvature, highlighting its superior performance in handling complex shapes.

[0103] Attachment Figure 5 This figure compares the visibility prediction results of our method with existing methods on noisy point clouds. Our method demonstrates strong robustness to noise, consistently outperforming the HPR method at all noise levels. In contrast, the HPR method is extremely sensitive to noise, and its accuracy decreases significantly with increasing noise levels.

[0104] Attachment Figure 6 This is the visibility prediction result of our method on the residual point cloud. Even if some points are removed, our method can still correctly predict the visibility of the originally occluded points. This shows that our method can effectively extract global geometric information.

[0105] Attachment Figure 7 This is a comparison of the prediction efficiency of our method and existing methods on point clouds of different sizes. When the point cloud size exceeds 130k points, our method significantly outperforms the HPR method, with a speed increase of up to 58 times.

[0106] Attachment Figure 8 This is the result of applying our invention to view-dependent surface reconstruction. After determining the visibility of each point, we project the visible points into view space and apply Delaunay triangulation to establish connectivity. The resulting triangulation is then projected back into 3D space to generate a view-dependent mesh. To ensure the reconstructed surface remains visually coherent, any triangles whose edge length exceeds a preset threshold are discarded.

[0107] Attachment Figure 9 This is the result of applying this invention to point cloud normal vector prediction and surface reconstruction. After view-dependent surface reconstruction, we estimate the normal of each point by averaging the normals of the triangle containing the point. We reconstruct the mesh from six predefined view directions and estimate the normals of the visible points. We then aggregate the normals of each point across these six views to obtain the final estimate. After this process, we can use the predicted normals to reconstruct the surface using Poisson surface reconstruction. For the few points that are not visible from any of these views, we can propagate the normal from the nearest visible point.

[0108] Attachment Figure 10This is the result of our invention applied to optimal viewpoint selection. Our method is differentiable, and the network's dual-channel output represents the probabilities of visibility and invisibility. We can define a differentiable loss function to minimize the invisibility score, thereby determining a viewpoint that maximizes visible detail and minimizes occlusion, resulting in a more comprehensive view of the shape.

[0109] Attachment Figure 11 This is the result of the present invention applied to shadow rendering. For a given point light source, we first use the light source position as the viewpoint position, input it into our system together with the point cloud, and determine the visibility of each point. We then project the visible points into the view space and apply Delaunay triangulation to establish a connection relationship to obtain a two-dimensional triangulation result. Subsequently, we use the depth information of the point cloud and the triangulation result to perform depth difference to obtain a depth map. During rendering, we achieve shadow rendering by comparing the depth relationship between the rendered point and the point light source with the position size corresponding to the depth map.

[0110] The synthetic data-driven end-to-end training strategy of the present invention has multiple advantages. First, through synthetic data generation and restoration, large-scale, high-quality annotated data can be obtained, solving the problem of scarce real data. Second, data augmentation methods such as Gaussian noise significantly improve the robustness and diversity of the model. Finally, the training strategy can adapt to point cloud data of different scales and complexities and has good generalization capabilities. In short, the synthetic data-driven end-to-end training strategy provides an efficient, robust, and practical solution for point cloud processing, which has broad application prospects.

[0111] The above description clearly and completely describes the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. It should be understood that the examples described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.

Claims

1. A method for rapid prediction of 3D point cloud visibility based on neural network, characterized in that: Construct a neural network architecture consisting of a feature extractor and a visibility predictor. The feature extractor uses a 3D U-Net structure to extract view-independent features of the point cloud. The visibility predictor uses a lightweight multi-layer perceptron to obtain the viewing direction of each point in the point cloud based on the input viewpoint position and the extracted features. The viewing direction is positionally encoded and multiplied by the extracted features. The multi-layer perceptron predicts the visibility of each point in the 3D point cloud. The following steps are included: 1) Construct a point cloud neural network based on octree convolution U-Net as a feature extractor for multi-scale feature extraction, and extract the multi-scale geometric feature vectors corresponding to the original point cloud; Convert the point cloud into an octree, where the thinnest octree node contains the average 3D spatial coordinates of each point in the corresponding voxel, which serves as the input signal of the feature extractor; The feature extractor consists of stacked residual blocks, downsampling and upsampling layers, and skip connections. Each residual block contains two octree-based convolutional layers, interspersed with batch normalization and activation functions. The multi-scale geometric features of the point cloud are learned through residual blocks; the spatial resolution is restored through upsampling layers and skip connections, so that the feature map is aligned with the spatial structure of the point cloud; and the octet leaf node features are mapped back to the original point cloud through interpolation to generate viewpoint-independent point-by-point feature vectors. 2) A viewpoint-aware lightweight multilayer perceptron is constructed as a visibility predictor, which converts the point cloud visibility into a binary classification task based on feature vector and viewpoint direction. Multilayer perceptual inference processing is performed on the feature vector of the point cloud and the corresponding viewpoint direction to obtain the visibility of each point in the point cloud.

2. The neural network-based three-dimensional point cloud visibility rapid prediction method according to claim 1, characterized in that: GPU parallel computing is used for multi-viewpoint scenes.

3. The neural network-based three-dimensional point cloud visibility rapid prediction method according to claim 1, characterized in that: For multi-viewpoint scenes, multi-viewpoint parallel reasoning and dynamic feature reuse are performed. For multiple viewpoint positions of the same point cloud, the feature vector of the point cloud is cached, and the multiple view direction vectors are spliced ​​and predicted to obtain the point cloud visibility under multiple viewpoints.

4. The neural network-based three-dimensional point cloud visibility rapid prediction method according to claim 3, characterized in that: Implementation method of dynamic feature reuse and parallel reasoning include: Establish a feature buffer pool to store the viewpoint-independent feature matrix F∈R (N×D) ; When processing K new viewpoints {v1,...,v K }, execute: Parallel calculation of all viewpoint codes γ(v K ); Broadcast feature matrix F to K copies, and multiply each viewpoint code element by element to get F K ; F K Reshape into a (K×N)×D tensor and obtain a K×N×2-dimensional prediction score through a single MLP forward calculation; Taking argmax of the last dimension, we get a K×N dimensional visibility matrix to achieve multi-viewpoint parallel reasoning.

5. The neural network-based three-dimensional point cloud visibility rapid prediction method according to claim 1, characterized in that: Design a synthetic data-driven end-to-end training strategy to perform dataset repair, point cloud sampling, visibility calculation, and data augmentation on the 3D dataset to obtain a synthetic training dataset with real visibility labels for network training; including: Repair meshes in 3D datasets so that they are airtight and manifold; Sampling evenly across the mesh surface and reducing the number of sampling points by sampling at the farthest point; Randomly and uniformly sample viewpoints on the surface of the bounding sphere of the point cloud; connect each point in the point cloud with the viewpoint to form a line segment, and calculate the intersection with all triangular faces; if an intersection occurs, the point is invisible; otherwise, it is visible; repeat this process for all points in the point cloud and generate corresponding visibility labels; Then perform data enhancement on the generated point cloud data.

6. The neural network-based three-dimensional point cloud visibility rapid prediction method according to claim 1, characterized in that: The feature extractor adapts to the processing of point cloud data of different scales and complexities by adjusting the depth of the octree and the number of residual blocks.

7. The neural network-based three-dimensional point cloud visibility rapid prediction method according to claim 1, characterized in that: The feature extractor adopts 3D U-Net architecture, and the processing flow includes: After normalization, the input point cloud is converted into a hierarchical structure containing L layers of nodes through an octree; In the encoder stage, the octree is subjected to layer-by-layer downsampling convolution operations with a convolution kernel size of 3×3. In the decoder stage, the spatial resolution is restored through deconvolution and skip connection, and the output layer feature dimension is D; The octet tree leaf node features are mapped to the original point cloud coordinates to generate a point-by-point feature matrix of dimension N×D, where N is the number of point clouds.

8. The neural network-based three-dimensional point cloud visibility rapid prediction method according to claim 1, characterized in that: In the lightweight multi-layer perceptron predictor, the viewpoint direction is represented as a three-dimensional vector p = (p1, p2, p3), which is mapped to a high-dimensional space through high-frequency position encoding; each component of the viewpoint direction is embedded as follows: c(p i )=[sin(2 0 πp i ),cos(2 0 πp i ),…,sin(2 L-1 πp i ),cos(2 L-1 πp i )] Where L represents the number of frequencies used to encode the viewpoint direction, and p i represents the i-th component of the viewpoint direction p, γ(p i ) is the embedded vector; the three components of the three-dimensional vector in the viewpoint direction are embedded separately and then spliced.

9. The method for rapid prediction of three-dimensional point cloud visibility based on a neural network as claimed in claim 8, characterized in that: The model uses the cross entropy loss function, which is defined as follows: Among them, y i is the true visibility label, is the predicted visibility probability, N is the number of points in the point cloud; Fast 3D point cloud visibility prediction based on neural network is achieved by minimizing cross entropy loss.

10. A three-dimensional point cloud visibility rapid prediction system based on a neural network, used to implement the method of claim 1, characterized in that: include: Feature extraction module, visibility prediction module, multi-viewpoint reasoning module and data generation and training module; among them, The feature extraction module is used to extract multi-scale geometric features from the input 3D point cloud and generate feature vectors; The visibility prediction module is used to predict the visibility state of each point in the 3D point cloud based on the feature vector and viewpoint direction generated by the feature extraction module; The multi-view inference module is used to achieve fast response under multi-view input conditions of a single point cloud, improving efficiency through dynamic feature reuse and parallel computing; The data generation and training module is used to generate synthetic training data with real visibility labels and optimize it through an end-to-end training strategy.

Citation Information

Patent Citations

  • Anchor-frame-free 3D target detection method based on multi-sensor fusion

    CN114118247A

  • Three-dimensional laser radar point cloud semantic segmentation method and device based on deep learning

    CN116229057A

  • Texture generation method and device

    CN117788673A

  • High-precision point cloud completion method based on deep learning and device thereof

    US20230206603A1