Point Cloud Pre-Training Method, System, Device and Medium Based on Neural Rendering

The neural rendering-based approach for point cloud pre-training addresses the limitations of existing methods by projecting three-dimensional point clouds into two-dimensional images, enhancing performance in three-dimensional object detection, segmentation, and reconstruction without additional data annotation.

CN116188894BActive Publication Date: 2025-07-15SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211665153.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2025-07-15
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

The existing point cloud training methods are difficult to achieve high precision and generalization performance when three-dimensional labeling is difficult, and cannot effectively use cheap and easy-to-get image data for training.

Method used

By obtaining color and depth images for three-dimensional backprojection, point cloud features are extracted, three-dimensional feature bodies are constructed, and neural rendering is rendered into images of different perspectives using neural rendering. The neural network is optimized using training loss functions to realize point cloud pre-training.

Benefits of technology

No additional manual annotation is required, and significant performance improvements in multiple downstream tasks can be achieved using only color-deep images, including 3D object detection, 3D semantic segmentation, 3D reconstruction and point cloud rendering, reducing pre-training data requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188894B_ABST
    Figure CN116188894B_ABST
Patent Text Reader

Abstract

Embodiments of this application relate to the field of artificial intelligence technology, and in particular to a point cloud pre-training method, system, device and medium based on neural rendering. The method includes: First, obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; then, extract the features of each point in the three-dimensional point cloud to obtain point cloud features; next, based on the point cloud features, construct a three-dimensional feature volume; then, use neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps; finally, compare the two-dimensional color and depth maps with the color and depth images input by the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function. The pre-training method provided by the embodiments of this application realizes a significant performance improvement in multiple downstream tasks by using neural rendering to project a three-dimensional scene onto a two-dimensional image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a point cloud pre-training method, system, device and medium based on neural rendering. Background Art

[0002] In the field of images, the training method of image neural networks represented by supervised learning has achieved good performance in multiple downstream vision tasks, such as object classification and object detection. However, for the point cloud modality, due to the difficulty of three-dimensional annotation, there are often only very few data annotations. Training a point cloud neural network with such a small amount of data often results in low accuracy and poor generalization performance. Therefore, it is very necessary to design a network training method for point cloud data that does not require a large amount of manual annotation.

[0003] Existing point cloud training methods can be roughly divided into two categories: contrast learning-based methods and point cloud completion-based methods. For contrast learning-based methods, two different data augmentations are performed on the same set of point clouds to obtain two sets of new augmented point clouds. By encouraging the point cloud neural network to obtain as consistent point cloud features as possible for these two sets of augmented new point clouds, pre-training of the point cloud network can be achieved. Another category of point cloud completion-based methods uses point cloud completion as the cloud training task. Such methods first perform a large number of occlusions on a set of point clouds, and then require the point cloud neural network to learn how to complete the complete point cloud from the unoccluded point clouds.

[0004] However, contrast learning-based methods are, firstly, relatively sensitive to the selected data augmentation, and secondly, various techniques are required to avoid model collapse, such as designing an effective positive and negative sample sampling strategy, etc. However, these strategies often require additional design. Point cloud completion-based methods need to solve the difficult problem of point cloud generation. In addition, point cloud completion-based techniques can only rely on three-dimensional point cloud data and cannot use more inexpensive and easily available image data. Therefore, both contrast learning-based methods and point cloud completion-based methods can only solve limited downstream tasks and are mainly effective in three-dimensional point cloud detection and three-dimensional point cloud segmentation. Summary of the Invention

[0005] The embodiments of the present application provide a point cloud pre-training method, system, device and medium based on neural rendering to achieve effective pre-training of the point cloud neural network.

[0006] To solve the above technical problems, in a first aspect, an embodiment of the present application provides a point cloud pre-training method based on neural rendering, including the following steps: First, obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; then, extract the features of each point in the three-dimensional point cloud to obtain point cloud features; next, construct a three-dimensional feature volume based on the point cloud features; then, use neural rendering to render the three-dimensional feature volume into images from different viewpoints to obtain two-dimensional color and depth maps; finally, compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function.

[0007] In some exemplary embodiments, the color and depth images include single-frame or multi-frame color and depth images; the color and depth images are obtained by a depth camera.

[0008] In some exemplary embodiments, performing three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud includes: inputting a plurality of color and depth images and their corresponding camera parameters; and obtaining a three-dimensional point cloud by using the method of point cloud back-projection based on the color and depth images and their corresponding camera parameters.

[0009] In some exemplary embodiments, comparing the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints includes: comparing the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints.

[0010] In some exemplary embodiments, a point cloud editor is used to extract the features of each point in the three-dimensional point cloud.

[0011] In some exemplary embodiments, constructing a three-dimensional feature volume based on the point cloud features includes: performing average pooling processing on the point cloud features, averaging the features of the points in space and distributing them to the three-dimensional grid to obtain a feature volume; and using a three-dimensional convolutional neural network to process the feature volume to obtain a three-dimensional feature volume.

[0012] In some exemplary embodiments, using neural rendering to render the three-dimensional feature volume into images from different viewpoints to obtain two-dimensional color and depth maps includes: setting rendering viewpoints, sampling on the rendering rays to obtain sampling points; obtaining the features of the sampling points from the three-dimensional feature volume through trilinear interpolation; sending the features of the sampling points to a neural network to estimate the color and signed distance function values of the sampling points to obtain estimated values; calculating the color values on the rendering rays by using the integral formula of neural rendering and the estimated values, and obtaining two-dimensional color and depth maps based on the color values.

[0013] Second aspect, the embodiments of the present application further provide a point cloud pre-training system based on neural rendering, including: a three-dimensional point cloud construction module, a three-dimensional feature volume construction module, a neural rendering module, and a data processing and optimization module connected in sequence; the three-dimensional point cloud construction module is used to obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; the three-dimensional feature volume construction module is used to extract the features of each point in the three-dimensional point cloud to obtain point cloud features; and based on the point cloud features, construct a three-dimensional feature volume; the neural rendering module is used to render the three-dimensional feature volume into images from different perspectives by using neural rendering to obtain two-dimensional color and depth maps; the data processing and optimization module is used to compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function.

[0014] In addition, the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned point cloud pre-training method based on neural rendering.

[0015] In addition, the present application further provides a computer-readable storage medium storing a computer program, characterized in that the computer program realizes the above-mentioned point cloud pre-training method when executed by a processor.

[0016] The technical solutions provided by the embodiments of the present application have at least the following advantages:

[0017] The embodiments of the present application provide a point cloud pre-training method, system, device and medium based on neural rendering. The method includes the following steps: First, obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; then, extract the features of each point in the three-dimensional point cloud to obtain point cloud features; next, based on the point cloud features, construct a three-dimensional feature volume; then, use neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps; finally, compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function.

[0018] The point cloud pre-training method based on neural rendering provided by this application does not require the use of additional manual data annotation, but only requires one or more color-depth images as input. The pre-training method proposed in this application projects a three-dimensional scene onto a two-dimensional image by using neural rendering, constructs the relationship between three-dimensional features and two-dimensional images, and can achieve good network pre-training effects without using complex data augmentation and various techniques, achieving significant performance improvements in multiple downstream tasks, including three-dimensional object detection, three-dimensional semantic segmentation, three-dimensional reconstruction, and point cloud rendering. In addition, the point cloud pre-training method proposed in this application uses point cloud rendering as a pre-training task, does not require dealing with complex point cloud completion tasks, and also realizes point cloud pre-training using only images, greatly reducing the requirements for pre-training data. In addition, the pre-training method proposed in this application proves that obvious performance improvements can still be obtained for underlying tasks in three-dimensional scenes, such as three-dimensional reconstruction and point cloud rendering. Brief Description of the Drawings

[0019] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.

[0020] Figure 1 It is a flow diagram of a point cloud pre-training method based on neural rendering provided by an embodiment of this application;

[0021] Figure 2 It is a schematic diagram of the specific process of a point cloud pre-training method based on neural rendering provided by an embodiment of this application;

[0022] Figure 3 It is an effect diagram of the pre-training method provided by an embodiment of this application on two different datasets (ScanNet, SUNRGB-D);

[0023] Figure 4 It is a schematic diagram of the result of the pre-training method provided by an embodiment of this application in three-dimensional segmentation;

[0024] Figure 5 It is a comparison diagram of the effects of different point cloud encoders in the downstream task of three-dimensional reconstruction provided by an embodiment of this application:

[0025] Figure 6 It is a comparison diagram of the convergence speed and convergence accuracy between a pre-trained point cloud rendering model and a non-pre-trained model provided by an embodiment of this application;

[0026] Figure 7 It is a schematic diagram of the result of the pre-training method provided by an embodiment of this application for three-dimensional vision tasks;

[0027] Figure 8Schematic structural diagram of a point cloud pre-training system based on neural rendering provided by an embodiment of the present application;

[0028] Figure 9 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] As can be seen from the background art, in the current prior art, whether it is the method based on contrast learning or the method based on point cloud completion, it can only solve limited downstream tasks, and is mainly effective in 3D point cloud detection and 3D point cloud segmentation.

[0030] For the method based on contrast learning, good point cloud features need to be learned through additional design to achieve pre-training of the point cloud network. First, such methods are relatively sensitive to the selected data augmentation. Selecting inappropriate data augmentation will lead to a significant decline in the effect. Although in the field of images, what kind of data augmentation is used for contrast learning of images has become increasingly clear, what kind of data augmentation to use for contrast learning in point clouds still needs to be explored. Second, the method of contrast learning needs to use various techniques to avoid model collapse, such as designing an effective positive and negative sample sampling strategy, etc. However, these strategies often require additional design.

[0031] For the method based on point cloud completion, the difficult problem of point cloud generation needs to be solved. Due to the characteristics of high sparsity, a large amount of noise, and disorder of point cloud data, it is very difficult to realize the generation and completion of point clouds. Therefore, there are natural difficulties in using point cloud completion as a pre-training task. In addition, the technology based on point cloud completion can only rely on three-dimensional point cloud data, and cannot use more inexpensive and easily available image data.

[0032] To solve the above technical problems, an embodiment of the present application provides a point cloud pre-training method based on neural rendering, including the following steps: First, obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; then, extract the features of each point in the three-dimensional point cloud to obtain point cloud features; next, based on the point cloud features, construct a three-dimensional feature volume; then, use neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps; finally, compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function. On the one hand, the point cloud pre-training method proposed in the present application constructs the relationship between three-dimensional features and two-dimensional images by introducing neural rendering technology, and can achieve good network pre-training effects without using complex data augmentation and various techniques. On the other hand, the point cloud pre-training method proposed in the present application uses point cloud rendering as a pre-training task, without having to process complex point cloud completion tasks, and also realizes point cloud pre-training using only images, greatly reducing the requirements for pre-training data. In addition, the pre-training method proposed in the present application proves that for underlying tasks in three-dimensional scenes, such as three-dimensional reconstruction and point cloud rendering, significant performance improvements can still be obtained.

[0033] The embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are proposed for the purpose of enabling readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.

[0034] See Figure 1 , an embodiment of the present application provides a point cloud pre-training method based on neural rendering, including the following steps:

[0035] Step S1: Obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud.

[0036] Step S2: Extract the features of each point in the three-dimensional point cloud to obtain point cloud features.

[0037] Step S3: Based on the point cloud features, construct a three-dimensional feature volume.

[0038] Step S4: Use neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps.

[0039] Step S5: Compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function.

[0040] This application aims to provide a new unsupervised learning method for point clouds to achieve effective pre-training of point cloud neural networks. Compared with existing point cloud network pre-training methods, this application does not require the use of additional artificial data annotation, but only requires single or multiple color-depth images as input. By using neural rendering to project a three-dimensional scene onto a two-dimensional image, the pre-training method proposed in this application achieves significant performance improvements in multiple downstream tasks, including 3D object detection, 3D semantic segmentation, 3D reconstruction, and point cloud rendering.

[0041] The embodiment of this application provides a point cloud pre-training method based on neural rendering. The method uses neural rendering as the point cloud pre-training task. At the same time, single-frame or multi-frame color-depth is used as pre-training data to achieve pre-training of the point cloud encoder. This method has the advantages of good effect, simple design, and easier data acquisition. At the same time, the point cloud pre-training method provided in this application proves to be effective in multiple downstream tasks, including 3D point cloud detection, 3D point cloud segmentation, 3D reconstruction, and point cloud rendering tasks.

[0042] It should be noted that the two-dimensional color and depth maps in step S5 are two-dimensional images obtained by rendering the three-dimensional feature volume into images from different perspectives using neural rendering in step S4. The corresponding color and depth images input for the viewpoints are obtained by inputting color and depth images before the three-dimensional back-projection of the color and depth images in step S1.

[0043] In some embodiments, the color and depth images in step S1 include single-frame or multi-frame color and depth images; the color and depth images are obtained by a depth camera.

[0044] It should be noted that color and depth images are used in this application. Color and depth images can also be referred to as color and depth maps. It can be understood that this application can also directly use color images or depth images without the need to have both at the same time.

[0045] Figure 2 The flowchart of the point cloud pre-training method proposed in this application is shown. The input of this application is single-frame or multi-frame color and depth images, which can be directly obtained by a depth camera. This application first obtains a three-dimensional point cloud through the three-dimensional back-projection of color and depth images. Then, features at each point in the three-dimensional point cloud are extracted. These features are subsequently processed into a three-dimensional feature volume. By randomly sampling points in the three-dimensional feature volume and performing projection, the three-dimensional feature can be rendered into a two-dimensional color and depth map. These two-dimensional color and depth maps are compared with the corresponding color and depth images input for the viewpoints and used as the training loss function to optimize the entire neural network. After training is completed, the point cloud encoder is used for various downstream tasks. The following will introduce each part in detail.

[0046] In some embodiments, in step S1, three-dimensional back-projection is performed on the color and depth images to obtain a three-dimensional point cloud, including: inputting a plurality of color and depth images and their corresponding camera parameters; and obtaining a three-dimensional point cloud by using a point cloud back-projection method based on the color and depth images and their corresponding camera parameters.

[0047] After obtaining the color and depth images in step S1, a three-dimensional point cloud is constructed from the color and depth images. First, a plurality of color and depth images and their corresponding camera parameters are input, and then a three-dimensional point cloud is obtained by means of point cloud back-projection. Specifically, the image pixels are first back-projected into the camera coordinate space through the camera internal parameters and depth values, and then converted to the unified world coordinate system through the camera external parameters. The point clouds obtained from different images are integrated in this world coordinate system. Different from the prior art methods, the present application not only uses the coordinate information of the point cloud, but also uses the color information of the point cloud as additional point cloud features.

[0048] In some embodiments, in step S5, comparing the two-dimensional color and depth map with the two-dimensional color and depth map input at the corresponding viewing point includes: comparing the two-dimensional color and depth map with the input color and depth images at the corresponding viewing point.

[0049] As mentioned above, in step S5, the two-dimensional color and depth map obtained in step S4 is compared with the input color and depth images at the corresponding viewing point, and the input color and depth images at the corresponding viewing point are obtained by inputting color and depth images before performing three-dimensional back-projection on the color and depth images in step S1. By introducing neural rendering technology, the present application constructs the relationship between three-dimensional features and two-dimensional images, and can achieve good pre-training effects of the network without using complex data augmentation and various techniques.

[0050] In some embodiments, in step S2, a point cloud editor is used to extract the features of each point in the three-dimensional point cloud. The present application uses a point cloud encoder to extract the features at each point. Since there are no additional requirements for the point cloud encoder in the proposed method, theoretically most point cloud encoders can be used in this scheme process. The present application attempts to use the classical point cloud encoders PointNet, PointNet++ and DGCNN. Subsequent experiments prove that good pre-training effects can be obtained by using different point cloud encoders.

[0051] In some embodiments, in step S3, based on the point cloud features, a three-dimensional feature volume is constructed, including: performing average pooling on the point cloud features, averaging the features of the points in the space and distributing them to the three-dimensional grid to obtain a feature volume; and using a three-dimensional convolutional neural network to process the feature volume to obtain a three-dimensional feature volume.

[0052] Specifically, after extracting the point cloud features in step S3, the point cloud features are organized into a three-dimensional feature volume. By way of example, the present application uses the method of average pooling to average the features of the points in space and then distribute them into a three-dimensional grid. Further, the present application uses a three-dimensional convolutional neural network to process the feature volume, and the processed feature volume is a dense three-dimensional feature volume.

[0053] In some embodiments, step S4 uses neural rendering to render the three-dimensional feature volume into images from different perspectives, obtaining two-dimensional color and depth maps, including: setting a rendering viewpoint, sampling on the rendering ray to obtain sampling points; the features of the sampling points are obtained from the three-dimensional feature volume by trilinear interpolation; the features of the sampling points are sent to a neural network to estimate the color and signed distance function (SDF) value of the sampling points, obtaining estimated values; using the integration formula of neural rendering and the estimated values, calculating the color value on the rendering ray, and based on the color value, obtaining the two-dimensional color and depth maps. The pre-training method proposed in the present application projects a three-dimensional scene onto a two-dimensional image by using neural rendering, achieving significant performance improvements in multiple downstream tasks, including three-dimensional object detection, three-dimensional semantic segmentation, three-dimensional reconstruction, and point cloud rendering.

[0054] Specifically, after obtaining the three-dimensional feature volume, the present application uses neural rendering to render the three-dimensional feature volume into images from different perspectives. By way of example, given a rendering viewpoint, the present application first samples on the rendering ray. The features of the sampling points are obtained from the three-dimensional feature volume by trilinear interpolation. The features of the sampling points are then sent to a neural network to estimate the color and signed distance function value (SDF value) of the sampling points. The SDF value represents the distance between the point and the true geometric surface of the scene and is often used as a representation of implicit geometry. In this way, each sampling point on the rendering ray can obtain the corresponding color value and SDF value. Further, using the integration formula of neural rendering, the color value on the rendering ray can be calculated. Rendering each image pixel can obtain the two-dimensional color and depth images corresponding to the viewpoint.

[0055] It should be noted that the neural rendering method used in the present application is not unique. There are multiple implementation methods for neural rendering. The present application uses one of the neural rendering methods, and other alternative neural rendering methods can also be used to achieve similar effects.

[0056] After obtaining the two-dimensional color and depth maps, the two-dimensional color and depth maps and the color and depth images input corresponding to the viewpoints

[0057] Compare to obtain the training loss function of the network, and optimize the neural network based on the training loss function. Specifically, the two-dimensional color and depth maps obtained after image rendering can be compared with the color and depth images input at the corresponding viewpoints.

[0058] Compare them as the training loss function of the network. The point cloud pre-training method proposed in this application takes point cloud rendering as the pre-training task, without having to process complex point cloud completion tasks. At the same time, it also realizes point cloud pre-training using only images, greatly reducing the requirements for pre-training data.

[0059] Preferably, the input and the rendered image should be as similar as possible. The point cloud neural network is required to learn real scene geometry and texture information from sparse point cloud data to achieve network pre-training. In addition, multiple regularization terms are used

[0060] to enhance the training stability of the network.

[0061] Next, the point cloud pre-training method based on neural rendering provided by this application is verified for its effects in downstream tasks of 3D detection, downstream tasks of 3D segmentation, downstream tasks of 3D reconstruction, downstream tasks of point cloud rendering, and direct applications in 3D reconstruction and point cloud rendering.

[0062] 5(1) Effects in downstream task of 3D detection:

[0063] The pre-training method proposed in this application can significantly improve the 3D detection effect of the basic point cloud rendering neural network. As Figure 3 shown, this application has achieved the best results among current algorithms on two different datasets (ScanNet, SUN RGB-D).

[0064] (2) Effects in downstream task of 3D segmentation: As Figure 4 shown, the pre-training algorithm proposed in this application has also achieved the best results in 3D segmentation.

[0065] (3) Effects in downstream task of 3D reconstruction:

[0066] The pre-training method proposed in this application is the first algorithm that proves to be effective for downstream task of 3D reconstruction. As Figure 5 shown, using different point cloud encoders, the pre-training method of this application can achieve improved reconstruction accuracy.

[0067] 5(4) Effects in downstream task of point cloud rendering:

[0068] The pre-training method proposed in this application is also effective for downstream task of point cloud rendering. As Figure 6As shown, the pre-trained point cloud rendering model can achieve a faster convergence speed and better convergence accuracy compared to the model without pre-training.

[0069] (5) Direct application in 3D reconstruction and point cloud rendering:

[0070] The model of this application can not only be used for pre-training downstream tasks, but also be directly used as various 3D vision tasks. The results are as Figure 7 shown. The point cloud pre-training method based on neural rendering provided by this application can achieve good 3D reconstruction and point cloud rendering results.

[0071] Based on this, the embodiments of this application provide a point cloud pre-training method based on neural rendering. Compared with the existing pre-training methods, the advantages of the pre-training method of this application are:

[0072] (1) Compared with the contrastive learning method, the pre-training task adopted in this application does not require designing complex data augmentation and does not require designing special techniques to avoid the common model collapse in contrastive learning.

[0073] (2) Compared with the point cloud completion method, the pre-training task adopted in this application does not require dealing with complex point cloud generation tasks, but only requires using image-level supervision.

[0074] (3) The method proposed in this application only needs to use color-depth images to achieve the pre-training of the point cloud network, without using scanned 3D models as input, greatly reducing the data cost of pre-training and making large-scale point cloud pre-training possible.

[0075] (4) The method proposed in this application utilizes image information, enabling the point cloud network to learn better semantic features from image supervision. The cross-modal training makes the pre-training method proposed in this application achieve better results.

[0076] Refer to Figure 8, an embodiment of the present application also provides a point cloud pre-training system based on neural rendering, including: a three-dimensional point cloud construction module 101, a three-dimensional feature volume construction module 102, a neural rendering module 103, and a data processing and optimization module 104 connected in sequence; the three-dimensional point cloud construction module 101 is used to obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; the three-dimensional feature volume construction module 102 is used to extract the features of each point in the three-dimensional point cloud to obtain point cloud features; and based on the point cloud features, construct a three-dimensional feature volume; the neural rendering module 103 is used to render the three-dimensional feature volume into images from different perspectives by using neural rendering to obtain two-dimensional color and depth maps; the data processing and optimization module 104 is used to compare the two-dimensional color and depth maps with the color and depth images input by the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function.

[0077] Reference Figure 9 , another embodiment of the present application provides an electronic device, including: at least one processor 110; and a memory 111 communicatively connected to the at least one processor; wherein, the memory 111 stores instructions executable by the at least one processor 110, and the instructions are executed by the at least one processor 110 to enable the at least one processor 110 to execute any of the above method embodiments.

[0078] Wherein, the memory 111 and the processor 110 are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 110 and the memory 111 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be one element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 110 is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 110.

[0079] The processor 110 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 111 can be used to store the data used by the processor 110 when performing operations.

[0080] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0081] That is, those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, optical disks, and other various media that can store program codes.

[0082] With the above technical solutions, the embodiments of the present application provide a point cloud pre-training method, system, device, and medium based on neural rendering. The method includes the following steps: First, obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; then, extract the features of each point in the three-dimensional point cloud to obtain point cloud features; next, based on the point cloud features, construct a three-dimensional feature volume; then, use neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps; finally, compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function.

[0083] The point cloud pre-training method based on neural rendering provided by the present application does not require the use of additional artificial data annotation, but only requires single or multiple color-depth images as input. The pre-training method proposed by the present application constructs the relationship between three-dimensional features and two-dimensional images by using neural rendering to project a three-dimensional scene onto a two-dimensional image. It can achieve good network pre-training effects without using complex data augmentation and various techniques, and realizes significant performance improvement in multiple downstream tasks, including three-dimensional object detection, three-dimensional semantic segmentation, three-dimensional reconstruction, and point cloud rendering. In addition, the point cloud pre-training method proposed by the present application takes point cloud rendering as a pre-training task, does not need to process complex point cloud completion tasks, and also realizes point cloud pre-training only using images, greatly reducing the requirements for pre-training data. In addition, the pre-training method proposed by the present application proves that obvious effect improvement can still be obtained for underlying tasks in three-dimensional scenes, such as three-dimensional reconstruction and point cloud rendering.

[0084] Those of ordinary skill in the art can understand that the above various embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.

Claims

1. A point cloud pre-training method based on neural rendering, characterized in that Comprising: Obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; Extract the features of each point in the three-dimensional point cloud to obtain point cloud features; Construct a three-dimensional feature volume based on the point cloud features; Use neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps; Compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function; Construct a three-dimensional feature volume based on the point cloud features, including: Perform average pooling on the point cloud features, average the features of points in space and distribute them to a three-dimensional grid to obtain a feature volume; Use a three-dimensional convolutional neural network to process the feature volume to obtain a three-dimensional feature volume; The step of using neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps includes: Set the rendering viewpoints, sample on the rendering rays to obtain sampling points; the features of the sampling points are obtained from the three-dimensional feature volume by trilinear interpolation; Send the features of the sampling points to a neural network to estimate the color and signed distance function values of the sampling points to obtain estimated values; Use the integral formula of neural rendering and the estimated values to calculate the color values on the rendering rays, and obtain two-dimensional color and depth maps based on the color values.

2. The method for pre-training point cloud based on neural rendering according to claim 1, wherein The color and depth images include single-frame or multi-frame color and depth images; The color and depth images are obtained by a depth camera.

3. The method for pre-training point cloud based on neural rendering according to claim 1, wherein Perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud, including: Input a number of color and depth images and their corresponding camera parameters; Based on the color and depth images and their corresponding camera parameters, use the method of point cloud back-projection to obtain a three-dimensional point cloud.

4. The method for pre-training point cloud based on neural rendering according to claim 3, wherein Compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints, including: Compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints.

5. The method for pre-training point cloud based on neural rendering according to claim 1, wherein Use a point cloud editor to extract the features of each point in the three-dimensional point cloud.

6. A point cloud pre-training system based on neural rendering, characterized in that, Comprising: A three-dimensional point cloud construction module, a three-dimensional feature volume construction module, a neural rendering module, and a data processing and optimization module connected in sequence; The three-dimensional point cloud construction module is used to obtain color and depth images, and perform three-dimensional back-projection on the color and depth images to obtain a three-dimensional point cloud; The three-dimensional feature volume construction module is used to extract the features of each point in the three-dimensional point cloud to obtain point cloud features; and construct a three-dimensional feature volume based on the point cloud features; The neural rendering module is used to use neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps; The data processing and optimization module is used to compare the two-dimensional color and depth maps with the color and depth images input at the corresponding viewpoints to obtain the training loss function of the network, and optimize the neural network based on the training loss function; Construct a three-dimensional feature volume based on the point cloud features, including: Perform average pooling on the point cloud features, average the features of the points in space and distribute them to the three-dimensional grid to obtain a feature volume; Use a three-dimensional convolutional neural network to process the feature volume to obtain a three-dimensional feature volume; The method of using neural rendering to render the three-dimensional feature volume into images from different perspectives to obtain two-dimensional color and depth maps includes: Set the rendering viewpoints, sample on the rendering rays to obtain sampling points; the features of the sampling points are obtained from the three-dimensional feature volume through the trilinear interpolation method; Send the features of the sampling points to the neural network to estimate the color and signed distance function values of the sampling points to obtain estimated values; Use the integral formula of neural rendering and the estimated values to calculate the color values on the rendering rays, and obtain two-dimensional color and depth maps based on the color values.

7. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the neural rendering-based point cloud pre-training method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the neural rendering-based point cloud pre-training method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Sparse sampling-based method and system for generating images from shot images to any viewpoint images

    CN114820945A

  • Three-dimensional object recognition method based on image pre-training model prompt learning

    CN115294296A