Rapid laser radar point cloud scene completion method and device based on diffusion model
Through variational fraction distillation and structural loss function optimization student model, the problem of slow completion speed of diffusion model is solved, and efficient and high-quality LiDAR scenario completion is achieved, suitable for autonomous vehicles.
Patent Information
- Application Number
- CN202510145238.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing LiDAR scenario completion method based on diffusion model is slow and cannot meet the demand for fast environment perception of autonomous vehicles, and the existing methods may lead to a decline in complementary quality.
Using the Variable Fraction Distillation (VSD) method, the multi-step diffusion model is distilled into a student model with few steps, and a structural loss function is introduced, including global loss and key point loss, and the student model is trained to capture the global information and local structure of the 3D LiDAR scene.
Significantly improve the completion speed, maintain or exceed the completion quality of the original model, ensure the overall and detailed accuracy of the completion scene, and is suitable for fast environmental perception of autonomous vehicles.
Smart Images

Figure CN120278904A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of lidar scanning scenarios, and particularly relates to a fast lidar point cloud scene completion method and device based on a diffusion model. Background Art
[0002] Accurately and effectively identifying the surrounding environment using vehicle-mounted sensors is crucial for the safe operation of autonomous vehicles. Among different types of sensors, 3D lidar (3D LiDAR) has become one of the most widely used sensors due to its wide detection range and accurate detection accuracy. However, the 3D point cloud collected by lidar is usually sparse, especially in occluded areas. This sparsity leads to a decline in the ability of autonomous vehicles to understand 3D scenes. Therefore, it is necessary to reason about and complete sparse 3D LiDAR scenes. Due to the advantages of strong training stability and high generation quality, diffusion models have achieved excellent results in the 3D LiDAR scene completion task. However, diffusion models often require multiple network iterations to obtain a dense, complete, and high-quality LiDAR scene, which is very time-consuming. Autonomous vehicles need to quickly and effectively perceive and identify the surrounding environment, so the slow sampling speed limits the practical application of diffusion models. Therefore, it is necessary to effectively accelerate existing LiDAR scene completion methods based on diffusion models to achieve efficient and high-quality LiDAR scene completion.
[0003] Chinese Patent No. CN118644405A discloses a method for completing lidar point cloud data, including the following steps: (1) A dual-light sensor collects lidar images and visible light images in the same time and space and performs preprocessing; (2) Construct an adversarial network, which includes a generator and a discriminator; (3) Input the preprocessed lidar images and visible light images into the adversarial network for training, and train through multiple iterations until optimal; (4) The output result of the generator of the adversarial network trained to optimality is the final completed image of the lidar point cloud data. This patent is based on obtaining the target contour from the visible light image collected in the dual-light coaxial system to adversarially generate the data required to complete the point cloud data obtained by lidar. While maintaining the accuracy of the data structure, it effectively increases the sampling point density in the edge area and improves the data quality and sensor performance. However, models based on generative adversarial networks are prone to problems such as unstable training and mode collapse, which affect the completion quality of the final model.
[0004] Chinese Patent with Publication No. CN111553859A discloses a method and system for complementing the reflection intensity of lidar point clouds. The specific steps are as follows: (1) Using a calibrated vehicle-mounted camera and lidar, obtain the grayscale image and the original point cloud of the same road surface; (2) Use a preset edge extraction strategy to extract the edge information of the grayscale image to obtain the edge image of the grayscale image; (3) Preprocess the original point cloud to obtain the original point cloud reflection intensity projection image and the interpolated and complemented point cloud reflection intensity projection image; (4) Input the grayscale image, the edge image of the grayscale image, the original point cloud reflection intensity projection image, and the interpolated and complemented point cloud reflection intensity projection image into a pre-trained point cloud reflection intensity complementation model to output the complemented point cloud reflection intensity projection image. The point cloud reflection intensity complementation method provided by this patent can make full use of the potential correlation between the point cloud and the image, thus effectively and accurately complementing the reflection intensity image of the lidar point cloud. However, the complemented result of this patent is the reflection image of the point cloud projected onto a two-dimensional plane, rather than the original 3D LiDAR point cloud.
[0005] LiDiff is a 3D lidar point cloud complementation method based on the diffusion model. LiDiff is based on the DDPM model, expands the process of the original diffusion model, and performs point-level noise addition and denoising on each point in the lidar point cloud scene. The specific steps are as follows: (1) Construct a DDPM model that takes the noise-added point cloud and the time step t as inputs, with the initial LiDAR scan as a condition; (2) Perform point-level noise addition on the lidar point cloud and train the DDPM model to learn the noise in the noise addition process; (3) Given a single radar scan By connecting its points K times to increase the size to obtain Then perform noise addition on to obtain the initial noise-added point cloud (4) Input into the trained diffusion model for denoising to obtain the complemented point cloud. LiDIff can complement sparse 3D LiDAR scans into a relatively complete point cloud scene. However, due to the inherent characteristics of the diffusion model, its complementation speed is slow, and it often takes 30 seconds to complement a scene, which limits its application in autonomous vehicles. Summary of the Invention
[0006] The present invention provides a fast lidar point cloud scene complementation method based on the diffusion model, which can quickly and accurately complement the lidar point cloud scene.
[0007] The present invention provides a fast lidar point cloud scene complementation method based on the diffusion model, including:
[0008] Use the sparse 3D LiDAR scan scene as training samples, and multiple training samples are used to construct a training sample set;
[0009] Construct a training model, which includes a frozen teacher diffusion model, a trainable student diffusion model and an auxiliary diffusion model. Denoise the training samples through the student diffusion model to obtain a completed scene, add noise to the completed scene, and input the noised completed scene and the training samples into the teacher diffusion model and the auxiliary diffusion model respectively to obtain the teacher diffusion model predicted noise and the auxiliary diffusion model predicted noise;
[0010] Construct a total loss function and a backpropagation loss function. The total loss function includes a KL divergence loss function, a global loss function and a key point loss function. The KL divergence loss function is constructed based on the difference between the teacher diffusion model predicted noise and the auxiliary diffusion model predicted noise. The global loss function is constructed based on the distances between points in the completed scene and the real scene. The key point loss function is constructed based on the difference between the distance matrix between key points in the real scene and the distance matrix between the points in the completed scene that are closest to the key points in the real scene;
[0011] Train the auxiliary diffusion model based on the generated completed scene through the backpropagation loss function, and train the student diffusion model based on the training sample set through the total loss function to obtain a point cloud scene completion model;
[0012] When applied, input the sparse 3D LiDAR scan scene into the point cloud scene completion model to obtain a completed LiDAR scan scene.
[0013] Preferably, denoising the training samples through the student diffusion model to obtain a completed scene includes:
[0014] Copy the training samples multiple times through the student diffusion model, add noise to each point in the copied training samples for the maximum number of levels, and denoise the noised training samples for a set number of denoising times to obtain a completed scene.
[0015] Preferably, the global loss function L scene is:
[0016]
[0017] where E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the real scene, is the completed scene, and this loss calculates the mean square error between each point x in the generated scene and the closest corresponding point y in the real scene .
[0018] Preferably, the key point loss function Lpoint is:
[0019]
[0020] wherein, E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the Euclidean distance matrix between the key points of the real scenario, is the Euclidean distance matrix between the points closest to the key points of the real scenario in the completed scenario.
[0021] Preferably, the method for obtaining the key points of the real scenario includes randomly selecting multiple points from the real scenario, calculating the curvature of the selected points, and taking the points with the top n largest curvatures as the key points.
[0022] Preferably, the method for calculating the curvature of the selected points includes:
[0023] Using the K-nearest neighbor method to calculate the set of neighboring points of the selected point i and obtaining the center of the set of neighboring points;
[0024] Based on the set of neighboring points, calculate the covariance matrix of the center, perform eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues, and obtain the curvature of the selected i-th point according to the multiple eigenvalues as:
[0025]
[0026] where, λ1 < λ2 < … < λ m , m is the number of eigenvalues, j is the index of the eigenvalue, and λ1 is the smallest eigenvalue.
[0027] Preferably, the KL divergence loss function L KL is:
[0028]
[0029] where, is the noise sample obtained by adding t levels of noise to the completed scenario, is the training sample, ∈ θ (·) is the teacher diffusion model, is the noise predicted by the teacher diffusion model, ∈ φ (·) is the auxiliary diffusion model, is the noise predicted by the auxiliary diffusion model, E t,∈ [·] is the joint mathematical expectation of variables t and ∈.
[0030] Preferably, the backpropagation loss function L φ is:
[0031]
[0032] Among them, E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the predicted noise of the auxiliary diffusion model, and ∈ is the noise added to the completed scene.
[0033] Preferably, initialize the teacher diffusion model, the student diffusion model, and the auxiliary diffusion model based on the LiDiff model, and set the denoising steps of the teacher diffusion model and the student diffusion model respectively.
[0034] The present invention also provides a fast LiDAR point cloud scene completion device based on a diffusion model, including: a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the fast LiDAR point cloud scene completion method based on the diffusion model.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] The present invention can distill a multi-step 3D LiDAR scene completion diffusion model into a few-step student model, so that the scene completion speed of the student model is increased exponentially, and the completion quality equivalent to or even better than that of the original teacher model is maintained.
[0037] Since the 3D LiDAR point cloud scene contains complex geometric structure information, directly using the original distillation method will cause the loss of details in the completed scene and result in a decrease in the completion quality. Therefore, in order to ensure the completion quality of the student model during the distillation process, this patent further introduces a structural loss to help the student model better learn the distribution of the 3D LiDAR point cloud scene. The structural loss includes two parts: the scene-based loss and the point-based loss. The scene-based loss will constrain the student model to capture the global information of the 3D LiDAR scene and ensure the overall quality of the completed scene. The point-based loss will constrain the student model to capture the relative structure between different point clouds and ensure the detail quality in the completed scene. This patent can achieve effective and high-quality completion of 3D LiDAR sparse scenes. Description of the Drawings
[0038] Figure 1 It is a flowchart of the fast LiDAR point cloud scene completion method based on the diffusion model provided by a specific embodiment of the present invention;
[0039] Figure 2 It is a detail diagram of the ScoreLiDAR completed scene and LiDiff provided by a specific embodiment of the present invention. Detailed Embodiments
[0040] To address the problem of the slow completion speed of the LiDiff model, a specific embodiment of the present invention provides a fast LiDAR point cloud scene completion method based on a diffusion model, called ScoreLiDAR. This is a new distillation method tailored for 3D LiDAR scene completion diffusion models, which can achieve efficient and high-quality scene completion. Based on variational score distillation (VSD), a specific embodiment of the present invention can distill a multi-step 3D LiDAR scene completion diffusion model into a few-step student model, enabling the scene completion speed of the student model to increase exponentially while maintaining a comparable or even better completion quality than the original teacher model. Since the 3D LiDAR point cloud scene contains complex geometric structure information, directly using the original distillation method may lead to the loss of details in the completed scene, resulting in a decline in completion quality. Therefore, to ensure the completion quality of the student model during the distillation process, this patent further introduces a structural loss to help the student model better learn the distribution of 3D LiDAR point cloud scenes. The structural loss consists of two parts: a global loss based on the scene and a key point loss based on key points. The global loss based on the scene constrains the student model to capture the global information of the 3D LiDAR scene, ensuring the overall quality of the completed scene. The key point loss based on points, on the other hand, constrains the student model to capture the relative structure between different point clouds, ensuring the detail quality in the completed scene. A specific embodiment of the present invention can achieve efficient and high-quality completion of 3D LiDAR sparse scenes.
[0041] A specific embodiment of the present invention provides a fast LiDAR point cloud scene completion method based on a diffusion model, as Figure 1 shown, including:
[0042] (1) Using the obtained sparse 3D LiDAR scan scene as a training sample, multiple training samples are used to construct a training sample set. Initialize the teacher diffusion model, student diffusion model, and auxiliary diffusion model using the LiDiff model. During training, the teacher diffusion model is from the existing 3D LiDAR scene completion model LiDiff with the best completion quality. The student diffusion model and the auxiliary diffusion model are both initialized with the same model as the teacher diffusion model, and the denoising steps of the student diffusion model are set. In one embodiment, the denoising steps of the student diffusion model are 8 steps, and the denoising steps of the teacher diffusion model, which is the LiDiff model, are 50 steps.
[0043] (2) Construct a training model, which includes a frozen teacher diffusion model, a trainable student diffusion model, and an auxiliary diffusion model. The student diffusion model is used to denoise the training samples to obtain a completed scene. Randomly sample the noise level \(t\in(0,1000]\), add noise to the completed scene, and input the noised completed scene and the training samples into the teacher diffusion model, so that the teacher diffusion model uses the training samples as a reference standard to complete the noised completed scene to obtain the predicted noise of the teacher diffusion model at the current moment. Similarly, input the noised completed scene and the training samples into the auxiliary diffusion model to obtain the predicted noise of the auxiliary diffusion model at the current moment.
[0044] In a specific embodiment, denoising the training samples by the student diffusion model to obtain a completed scene includes:
[0045] The student diffusion model copies the training samples multiple times, adds noise to each point in the copied training samples at the maximum level \(T\), so as to realize the noise addition of the scanned point cloud, and denoise the noised training samples for a set number of times to obtain a completed scene.
[0046] (3) Construct a total loss function and a backpropagation loss function. The total loss function provided in the specific embodiment of the present invention includes a KL divergence loss function and a structural loss function. The structural loss function includes a global loss function and a key point loss function. The KL divergence loss function is constructed based on the difference between the predicted noise of the teacher diffusion model and the predicted noise of the auxiliary diffusion model. The global loss function is constructed based on the distances between the points of the completed scene and the real scene. The key point loss function is constructed based on the difference between the distance matrix between the key points of the real scene and the distance matrix between the points in the completed scene that are closest to the key points of the real scene.
[0047] In a specific embodiment, the KL divergence loss function \(L\) provided in this embodiment KL is:
[0048]
[0049] where is the noise sample obtained by adding \(t\) levels of noise to the completed scene, is the training sample, \(\in\) θ \((\cdot)\) is the teacher diffusion model, is the predicted noise of the teacher diffusion model, \(\in\) φ \((\cdot)\) is the auxiliary diffusion model, is the predicted noise of the auxiliary diffusion model, \(E\) t,∈[·] is the joint mathematical expectation of variables t and ∈. In a specific embodiment of the present invention, the predicted noise of the auxiliary diffusion model is approximated to the predicted noise of the teacher diffusion model, so as to indirectly make the completed scenarios output by the student diffusion model approach those output by the teacher diffusion model, and achieve that the quality of the completed scenarios output by the student diffusion model reaches the level of the teacher diffusion model.
[0050] In a specific embodiment, the structural loss function provided in this embodiment includes a global loss function and a key-point loss function. The global loss function based on the scenario will constrain the student diffusion model to capture the global information of the 3D LiDAR scenario, ensuring the overall quality of the completed scenario. And the key-point loss function based on points will constrain the student diffusion model to capture the relative structure between different point clouds, ensuring the detail quality in the completed scenario.
[0051] During the distillation process, when t >> 0, the gradient corresponding to the KL divergence loss function is well-defined, but when t is very small, the gradient becomes unreliable. This is because due to the complexity of the point cloud data, the student diffusion model often produces substandard results in the early stage. The generated noise samples are easily outside the training distribution of the teacher model, resulting in unreliable network predictions of the teacher model. Therefore, a specific embodiment of the present invention introduces a global loss function at the scene level to solve this problem and constrain the student diffusion model to capture the overall features of the completed scenario.
[0052] The global loss function L provided by a specific embodiment of the present invention scene is:
[0053]
[0054] where E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the real scenario, is the completed scenario, and this loss calculates the mean square error between each point x in the generated scenario and the closest corresponding point y in the real scenario . By minimizing the distance between the real scenario and the completed scenario through the global loss function, it helps the student diffusion model capture the overall structure, prevent the optimization direction from deviating in the early stage, and enhance the training stability. The scene loss also makes the generated scenario globally closer to the real scenario, thereby improving the completion quality and fidelity.
[0055] During the distillation process, only the overall distribution of the completed scene is constrained, and the relative position structure information between different points is ignored. Directly using the gradient in the KL divergence loss function to optimize the student model will result in the loss of local details. Therefore, in a specific embodiment of the present invention, a key-point loss function based on key points is introduced to capture the relative structure information between different points in the 3D LiDAR scene. The point-based key-point loss function calculates the difference between the point-to-point distance matrices of the completed scene and the real scene. Since there are a large number of points in the scene, the computational cost of calculating the distance matrix for all points is very high. Therefore, in a specific embodiment of the present invention, n key points are selected to calculate the distance matrix, and according to the local geometric features of each point, the key points that are crucial for representing the 3D LiDAR scene structure are selected. For each point in the completed scene, key points are selected based on the magnitude of its curvature. The larger the curvature, the greater the local shape change. These points with large local changes are usually located at corners, edges, or endpoints, and these points tend to form the main structure of the scene. Therefore, the top n points with the highest curvature values are selected as key points.
[0056] In a specific embodiment, obtaining the Euclidean matrix between the key points of the real scene includes:
[0057] Thus, randomly select q points in the real scene, and construct a q×q distance matrix based on the Euclidean distances between each pair of the q points represents the Euclidean distance between the s-th point and the t-th point, and then from the completed scene select the q points that are closest to the key points in as the corresponding key points, and obtain the corresponding distance matrix The key-point loss function L provided by a specific embodiment of the present invention point is:
[0058]
[0059] where E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the Euclidean distance matrix between the key points of the real scene, is the Euclidean distance matrix between the points closest to the key points of the real scene in the completed scene. The point-wise loss implemented based on the key-point loss function can help the student diffusion model capture the relative structure information of the key points, further improving the geometric accuracy and detail retention of the completed scene. The above operations can ensure that key objects such as cars, road studs, and signs are better completed, which is crucial for autonomous vehicles to accurately identify the surrounding environment.
[0060] In a specific embodiment, the method for calculating the curvature of the selected points includes:
[0061] Calculate the set of neighboring points of the selected points using the K-nearest neighbor method, and obtain the center of the set of neighboring points;
[0062] Calculate the covariance matrix of the center based on the set of neighboring points, perform eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues, and obtain the curvature of the i-th selected point according to the multiple eigenvalues as:
[0063]
[0064] where λ1 < λ2 < … < λ m , m is the number of eigenvalues, j is the index of the eigenvalue, and λ1 is the smallest eigenvalue.
[0065] The total loss function L provided by the specific embodiment of the present invention is:
[0066] L = L KL + λ scene L scene + λ point L point
[0067] where λ scene and λ point are the weights of the global loss function and the key point loss function respectively, and the point cloud scene completion model is trained by the total loss function based on the training sample set.
[0068] In a specific embodiment, the backpropagation loss function L φ provided by this embodiment is:
[0069]
[0070] where E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the noise predicted by the auxiliary diffusion model, ∈ is the noise added to the completed scene, and the auxiliary diffusion model is trained by the backpropagation loss function based on the generated completed scene.
[0071] (4) During application, input the sparse 3D LiDAR scan scene into the point cloud scene completion model to obtain the completed LiDAR scan scene.
[0072] In an embodiment, the fast LiDAR point cloud scene completion method based on the diffusion model provided by this embodiment includes:
[0073] S1 Initialize the teacher model ∈ φ , the student model G and the auxiliary diffusion model ∈ φ based on the pre-trained model of LiDiff.
[0074] S2 uses the KL divergence loss L KL , the scene-based loss L scene and the point-based loss L point to distill the student model G.
[0075] S3 uses the trained student model G to quickly and highly-quality complete any 3D LiDAR scene.
[0076] Specific steps:
[0077] S1 reads the pre-trained LiDiff model and directly regards it as the teacher model ∈ θ . Uses deepcopy to copy the teacher model ∈ θ twice. One copy is regarded as the student diffusion model G, and the other copy is regarded as the auxiliary diffusion model ∈ φ .
[0078] S2 uses the student diffusion model to generate the completed scene at the specified denoising steps (8 steps) Based on alternately optimizes the student model G and the auxiliary diffusion model ∈ φ until the student model G converges.
[0079] S2-1 Based on the sparse LiDAR scan increases the size by connecting its points K times to obtain Then is noise-added to obtain the initial noise-added point cloud
[0080] S2-2 Uses the student model to denoise and sets the denoising steps to 8 steps (the original teacher model is 50 steps). Obtains the completed scene
[0081] S2-3 Randomly samples the noise level t ∈ (0, 1000], and noise-adds the completed scene
[0082] S2-4 Based on the obtained in S2-3, calculates the KL divergence loss L KL .
[0083] S2-5 Based on the real scene and the obtained in S2-3, calculates the scene-based loss L scene .
[0084] S2-6 Based on the real scene Randomly select 1 / 10 of the points from all points and calculate the curvature of each point: First, use K-nearest neighbors to calculate the set of neighboring points for each point Calculate the center of Based on For each Calculate the covariance matrix For the covariance matrix Perform eigenvalue decomposition to obtain m eigenvalues λ1 < λ2 < … < λ m ; Calculate the curvature based on the eigenvalues Finally, arrange according to the curvature magnitude and select 1 / 3 of the points with the largest curvature as key points to calculate the point distance matrix Similarly, in the completed scene select the points closest to the points in to construct the point distance matrix Calculate the point-based loss L point .
[0085] S2-7 Calculate the total loss L = L KL + λ scene L scene + λ point L point And update the student model G. Where λ scene = 0.5, λ point = 0.01. The optimizer for updating the student model G is the SGD optimizer with a learning rate of 3e-5.
[0086] After updating the student model G in S2-8, use the student model G to generate a new completed scene and update the auxiliary diffusion model ∈ φ . Update the auxiliary diffusion model ∈ φ The optimizer for updating the auxiliary diffusion model ∈ is the SGD optimizer with a learning rate of 3e-5.
[0087] S2-9 Repeat S2-1 - S2-8 until the student diffusion model G converges, and the update ratio of the student diffusion model G and the auxiliary diffusion model ∈ φ is 1:1, that is, after updating the student model G once, the auxiliary diffusion model ∈ will be updated once φ . The total number of times to repeat S2-1 - S2-8 is set to 50 times. Run on a single A40 GPU.
[0088] S3 Use the trained student model G to quickly and high-quality complete any 3D LiDAR scene.
[0089] S3-1 Randomly select a sparse 3D LiDAR scan scene to be completed First, copy the point cloud it contains K times to obtain a dense point cloud
[0090] S3-2 pairs Add noise. The scale of adding noise is set to the maximum value T of the scale. The input maximum noise scale T means starting from the maximum noise scale for denoising to obtain the initial noisy point cloud.
[0091] S3-3 Input and T into the trained student model G, and perform denoising according to the specified number of steps. Given the sample to be denoised, the method of denoising is as follows: Perform denoising once to obtain as follows:
[0092]
[0093] Here, α t is a series of defined variables used to schedule the denoising process. is random noise. After denoising according to the specified number of steps (8 steps), a complete and dense 3D LiDAR scene is finally obtained.
[0094] A specific embodiment of the present invention also provides a fast LiDAR point cloud scene completion device (ScoreLiDAR) based on a diffusion model, including: a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the fast LiDAR point cloud scene completion method based on the diffusion model.
[0095] Performance analysis:
[0096] Compared with the existing optimal 3D LiDAR scene completion diffusion model, the fast LiDAR point cloud scene completion method provided by the specific embodiment of the present invention improves the completion speed by more than 5 times and achieves better completion quality.
[0097] Secondly, the specific embodiment of the present invention can provide a good data basis for autonomous driving vehicles to recognize and perceive the surrounding driving environment, and has practical significance and good application prospects.
[0098] Table 1 Test results on the SemanticKITTI dataset
[0099]
[0100] Table 2 Test results on the KITTI-360 dataset
[0101]
[0102] Table 1 and Table 2 show the completion effects of this patent on different datasets. The metrics are Chamfer Distance (CD) and Jensen-Shannon Divergence (JSD). Both of these metrics can calculate the gap between the completed scene and the real scene. The smaller the value of the metric, the closer the completed scene is to the real scene. It can be seen from the results that the ScoreLiDAR model proposed in this patent achieves the optimal completion quality compared with the existing models, and the completion speed is increased by more than 5 times compared with LiDiff.
[0103] As Figure 2 shown in (a), (b), (c), and (d) of [], the comparison of the completion results between the ScoreLiDAR method provided by the specific embodiment of the present invention and the current optimal 3D LiDAR scene completion method. It can be seen that compared with the LiDiff method, the method proposed in this patent can better complete the details such as vehicles and road piles in the scene, which is very important for the safe driving of autonomous vehicles.
Claims
1. A fast lidar point cloud scene completion method based on a diffusion model, characterized in that, Including: Taking sparse 3D LiDAR scanning scenes as training samples, and constructing a training sample set with multiple training samples; Constructing a training model, the training model includes a frozen teacher diffusion model, a trainable student diffusion model and an auxiliary diffusion model. The student diffusion model is used to denoise the training samples to obtain a completed scene, the completed scene is added noise, and the noise-added completed scene and the training samples are respectively input into the teacher diffusion model and the auxiliary diffusion model to obtain the teacher diffusion model predicted noise and the auxiliary diffusion model predicted noise; Constructing a total loss function and a backpropagation loss function, the total loss function includes a KL divergence loss function, a global loss function and a key point loss function. The KL divergence loss function is constructed based on the difference between the teacher diffusion model predicted noise and the auxiliary diffusion model predicted noise, the global loss function is constructed based on the distances between points of the completed scene and the real scene, and the key point loss function is constructed based on the difference between the distance matrix of key points in the real scene and the distance matrix of points closest to the key points in the real scene in the completed scene; Based on the generated completed scene, the auxiliary diffusion model is trained through the backpropagation loss function, and the student diffusion model is trained through the total loss function based on the training sample set to obtain a point cloud scene completion model; During application, the sparse 3D LiDAR scanning scene is input into the point cloud scene completion model to obtain a completed LiDAR scanning scene.
2. The method for quickly completing a lidar point cloud scene based on a diffusion model according to claim 1, wherein Using the student diffusion model to denoise the training samples to obtain a completed scene, including: The student diffusion model is used to copy the training samples multiple times, each point in the training samples copied multiple times is added noise at the maximum number of levels, and the noise-added training samples are denoised for a set number of denoising steps to obtain a completed scene.
3. The method for quickly completing a lidar point cloud scene based on a diffusion model according to claim 1, wherein The global loss function L scene is as follows: Among them, E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the real scenario, is the completed scenario, and this loss calculates the mean square error between each point x in the generated scenario and the closest corresponding point y in the real scenario .
4. The method for quickly completing a lidar point cloud scene based on a diffusion model according to claim 1, wherein Key point loss function L point is as follows: Among them, E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the Euclidean distance matrix between key points in the real scenario, is the Euclidean distance matrix between points closest to the key points in the real scenario in the completed scenario.
5. The method for rapidly completing a lidar point cloud scene based on a diffusion model according to claim 1 or 4, characterized in that, A method for obtaining key points of a real scene, including randomly selecting multiple points from the real scene, calculating the curvature of the selected points, and taking the points with the top n largest curvatures as key points.
6. The method for fast lidar point cloud scene completion based on the diffusion model according to claim 5, characterized in that A method for calculating the curvature of the selected points, including: Using the K-nearest neighbor method to calculate the set of neighboring points of the selected point i, and obtaining the center of the set of neighboring points; Based on the set of neighboring points, calculating the covariance matrix of the center, performing eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues, and obtaining the curvature of the selected i-th point as: Among them, λ1 < λ2 < … < λ m , m is the number of eigenvalues, j is the index of the eigenvalue, and λ1 is the smallest eigenvalue.
7. The method for quickly completing a lidar point cloud scene based on a diffusion model according to claim 1, wherein The KL divergence loss function L KL is as follows: Among them, is the noise sample obtained by adding t levels of noise to the completed scene, is the training sample, ∈ θ (·) is the teacher diffusion model, is the noise predicted by the teacher diffusion model, ∈ φ (·) is the auxiliary diffusion model, is the noise predicted by the auxiliary diffusion model, E t,∈ [·] is the joint mathematical expectation of variables t and ∈.
8. The method for quickly completing a lidar point cloud scene based on a diffusion model according to claim 1, wherein The backpropagation loss function L φ is as follows: Among them, E t,∈ [·] is the joint mathematical expectation of variables t and ∈, is the noise predicted by the auxiliary diffusion model, and ∈ is the noise added to the completed scene.
9. The method for quickly completing a lidar point cloud scene based on a diffusion model according to claim 1, wherein, Initializing the teacher diffusion model, the student diffusion model and the auxiliary diffusion model based on the LiDiff model, and respectively setting the denoising steps of the teacher diffusion model and the student diffusion model.
10. A fast lidar point cloud scene completion device based on a diffusion model, characterized in that, Including: Including a memory and one or more processors, where executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the method for fast lidar point cloud scene completion based on a diffusion model according to any one of claims 1-9.
Citation Information
Patent Citations
Laser radar point cloud reflection intensity completion method and system
CN111553859A
Complementation method for laser radar point cloud data
CN118644405A