Road side scene point cloud completion method and system based on diffusion model
By using a point cloud completion method based on a diffusion model for roadside scenes, the pose of occluded objects is corrected using roadside visual information, pseudo-ground value supervision signals are generated, and a point cloud completion diffusion model is trained. This solves the contradiction between global consistency and local refinement in roadside scenes and achieves high-precision environmental perception.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGXIA UNIVERSITY
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing point cloud completion methods struggle to balance global consistency and local refinement in roadside scenarios, resulting in insufficient environmental perception accuracy for vehicles in autonomous driving and intelligent transportation systems.
A point cloud completion method based on a diffusion model is adopted for roadside scenes. By establishing a global optimization objective, constructing a geometric-semantic joint learning framework and designing a region-aware reconstruction loss function, the pose of occluded objects is corrected using roadside visual information, pseudo-ground value supervision signals are generated, and a point cloud completion diffusion model is trained for completion.
While maintaining global consistency of the roadside scene, it restores the local details of traffic targets in the scene, provides reliable global perception capabilities, and improves the environmental perception accuracy of autonomous driving and intelligent transportation systems.
Smart Images

Figure CN121883786A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of global perception technology for autonomous driving and intelligent transportation systems, specifically to a method and system for roadside scene point cloud completion based on a diffusion model. Background Technology
[0002] LiDAR (Light Detection and Ranging) is a sensor that uses laser beams for scanning and ranging. It acquires three-dimensional information about the surrounding environment, including distance, height, and shape, by emitting a laser beam and measuring the time it takes for it to reflect back. LiDAR-based 3D environmental perception has become one of the key technologies for high-precision environmental understanding in autonomous driving and intelligent transportation systems. However, existing research mainly focuses on the vehicle-mounted perspective, such as 3D object detection, trajectory prediction, and semantic segmentation. Limited by the physical limitations of vehicle-mounted sensors in terms of perception range and blind spots, single-vehicle intelligent perception struggles to achieve full-area perception capabilities. To overcome this physical limitation, collaborative perception based on roadside LiDAR is gradually becoming a highly promising solution. Compared to vehicle-mounted sensors, roadside LiDAR provides a more comprehensive and accurate environmental characterization through its unique top-down perspective and beyond-line-of-sight coverage. However, the vast spatial span of roadside scenes leads to more significant sparsity and density decay characteristics in LiDAR point clouds. Furthermore, the long-term or complete occlusion between roadside traffic participants (such as a large truck completely occluding a small car) exacerbates the incompleteness of roadside scene point clouds. This incompleteness severely impacts the accuracy of downstream 3D perception tasks. Therefore, a point cloud completion method is needed that can recover the geometry of occluded targets while maintaining global consistency and detail integrity.
[0003] Existing point cloud completion methods can be broadly categorized into object-level completion methods and scene-level completion methods based on the vehicle's perspective. Object-level completion methods learn shape priors from large-scale datasets to complete single, isolated objects. When these methods are directly applied to large-scale scenes, they struggle to effectively handle the severe non-uniformity of point cloud coverage scale within a single frame caused by LiDAR, leading to a sharp decline in performance. Scene completion methods based on the vehicle's perspective guide the model to complete the scene by aggregating information from the entire scene. While these methods achieve excellent performance in scene completion from the vehicle's perspective and can effectively handle the severe non-uniformity of point cloud coverage scale within a single frame caused by LiDAR, when applied to roadside scenes, the pose uncertainty caused by long-term or complete occlusion makes it difficult for these methods to maintain the local fine details of scene objects while ensuring global consistency in scene completion. This results in difficulties in providing accurate environmental perception for vehicles in autonomous driving and intelligent transportation systems. Summary of the Invention
[0004] In view of this, the present invention provides a roadside scene point cloud completion method and system based on a diffusion model to solve the technical problem that existing point cloud completion methods, when applied to roadside scenes, are difficult to balance the contradiction between global consistency and local refinement, resulting in difficulty in providing accurate environmental perception for vehicles in autonomous driving and intelligent transportation systems.
[0005] The technical solution adopted by this invention to solve its technical problem is:
[0006] A point cloud completion method for roadside scenes based on a diffusion model includes the following steps:
[0007] S1. Establish a global optimization objective, and determine the probabilistic attitude difference objective function based on the global optimization objective;
[0008] S2. Correct the pose of occluded objects in the roadside scene based on the probability pose difference objective function and generate a pseudo-true value supervision signal.
[0009] S3. Construct a joint learning framework for geometric and semantic knowledge;
[0010] S4. Design a region-aware reconstruction loss function based on the aforementioned geometric-semantic joint learning framework;
[0011] S5. Establish a point cloud completion diffusion model, and train the point cloud completion diffusion model based on the pseudo-true value supervision signal and the region perception reconstruction loss function.
[0012] S6. Perform point cloud completion for roadside scenes based on the trained point cloud completion diffusion model.
[0013] Preferably, in step S1, the global optimization objective is as follows:
[0014] ;
[0015] in, Represents sparse point clouds for lidar, This represents the erroneous estimated pose from the roadside visual images and the output of the visual detector. Representing a true and complete roadside scene point cloud, Represents the true posture of an object; This indicates the visibility sampling of the lidar. This indicates scene composition based on object pose and pseudo-realistic templates. Given the true pose Visual likelihood of visual evidence C It is a priori attitude.
[0016] Preferably, in step S2, the pose of the occluded object in the roadside scene is corrected based on the probabilistic pose difference objective function to generate a pseudo-true value supervision signal, specifically including the following steps:
[0017] S21. Calculate the objective function that minimizes the probability pose difference, integrate external visual evidence and lidar geometric likelihood to correct the pose of occluded objects in the roadside scene, and obtain the corrected pose and confidence level for each object.
[0018] S22. Generate a pseudo-true value point cloud with confidence based on the corrected attitude and confidence level, as a pseudo-true value supervision signal.
[0019] Preferably, in step S3, the geometric-semantic joint learning framework is constructed, which specifically includes the following steps:
[0020] S31. Based on the corrected posture, construct a binary dynamic scene mask matrix to distinguish between dynamic variable regions and static background regions, and provide global spatial geometric priors;
[0021] S32. Parallel convolutional branching is used to extract geometric features at multiple granularities;
[0022] S33. Model local structural relationships based on anchor point grouping and output the final features.
[0023] Preferably, in step S4, the region-aware reconstruction loss function consists of three parts: mean squared error, dynamic scene mask weighted loss, and regularization term.
[0024] Preferably, the dynamic scene mask weighted loss is:
[0025] ;
[0026] in , and These are the weight hyperparameters for the dynamically changing local region and the static background region, respectively. For dynamic scene mask matrix;
[0027] The regularization loss is:
[0028] ;
[0029] The total loss is:
[0030] .
[0031] Preferably, in step S5, the specific training process is as follows: using the original sparse point cloud and pseudo-true value supervision signal as conditions, Gaussian noise is gradually added using the forward process of the diffusion model, and training optimization is performed using the region-aware reconstruction loss function.
[0032] Preferably, in step S6, the roadside scene point cloud is completed based on the trained point cloud completion diffusion model, specifically including: using the reverse inference denoising of the point cloud completion diffusion model to complete the single frame roadside scene point cloud.
[0033] Preferably, during the completion process, a classification-free guidance strategy is adopted for the input single-frame roadside scene point cloud, as follows:
[0034] ;
[0035] in, This represents the final noise prediction after adjustment without a classifier, used to guide the denoising process; This represents the diffusion model θ given the current diffusion state. Time step t, sparse point cloud Y, and pseudo-real values The noise predicted below; Indicates the intensity of unclassified guidance. This indicates the condition for removing false truth values.
[0036] This invention also provides a roadside scene point cloud completion system based on a diffusion model, applied to the roadside scene point cloud completion method based on a diffusion model as described above. The system includes a perception data acquisition module, an edge computing module, and a storage module. The perception data acquisition module includes a roadside LiDAR and a roadside camera, used to acquire the original sparse point cloud and visual images of the roadside scene. The edge computing module includes an industrial control computer and a graphics processing unit. The industrial control computer is communicatively connected to the perception data acquisition module and is used to process the acquired data and run the diffusion model algorithm to perform attitude correction, feature extraction, and point cloud completion tasks. The storage module stores the roadside visual detector model, the roadside scene point cloud completion model based on the diffusion model, a preset object template library, and a trained diffusion model weight file.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] This invention first establishes a global optimization objective based on the technical problem to be solved, determines a probabilistic pose difference objective function based on the global optimization objective, corrects the roadside visual detector error based on the probabilistic pose difference objective function, and provides reliable estimated poses of occluded objects in the roadside scene. Then, it generates pseudo-ground values for the occluded objects and provides pseudo-ground value supervision signals for subsequent diffusion model training to complete the scene. Next, it constructs a geometric-semantic joint learning framework to learn the structured geometric representations of the roadside scene layout and dynamic traffic participants, achieving both global and refined point cloud completion of the roadside scene. Based on the geometric-semantic joint learning framework, it designs a region-aware reconstruction loss function; establishes a point cloud completion diffusion model and trains it based on the pseudo-ground value supervision signals and the region-aware reconstruction loss function; finally, it performs point cloud completion of the roadside scene based on the trained point cloud completion diffusion model. Therefore, this invention uses roadside visual information as compensation to estimate the pose of occluded objects and guides the diffusion model to generate point clouds of the entire scene. This allows for the restoration of local details of traffic targets in the scene while maintaining the global consistency of the roadside scene, providing reliable global perception capabilities for autonomous driving and intelligent transportation systems, so as to better optimize traffic flow and improve the driving safety of autonomous vehicles. Attached Figure Description
[0039] Figure 1 This is a flowchart of a point cloud completion method for roadside scenes based on a diffusion model.
[0040] Figure 2 It is a picture of the original sparse roadside scene.
[0041] Figure 3 This is the completed roadside scene image.
[0042] Figure 4 This is the result image of the visual detector.
[0043] Figure 5 This is the image after correction by the visual detector. Detailed Implementation
[0044] The technical solutions and effects of the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0045] Please refer to Figure 1 A point cloud completion method for roadside scenes based on a diffusion model includes the following steps:
[0046] S1. Establish a global optimization objective, and determine the probabilistic attitude difference objective function based on the global optimization objective;
[0047] S2. Correct the pose of occluded objects in the roadside scene based on the probability pose difference objective function and generate a pseudo-true value supervision signal.
[0048] S3. Construct a joint learning framework for geometric and semantic knowledge;
[0049] S4. Design a region-aware reconstruction loss function based on the aforementioned geometric-semantic joint learning framework;
[0050] S5. Establish a point cloud completion diffusion model, and train the point cloud completion diffusion model based on the pseudo-true value supervision signal and the region perception reconstruction loss function.
[0051] S6. Perform point cloud completion for roadside scenes based on the trained point cloud completion diffusion model.
[0052] This invention first establishes a global optimization objective based on the technical problem to be solved, determines a probabilistic pose difference objective function based on the global optimization objective, corrects the roadside visual detector error based on the probabilistic pose difference objective function, and provides reliable estimated poses of occluded objects in the roadside scene. This then generates pseudo-ground values for the occluded objects and provides pseudo-ground value supervision signals for subsequent diffusion model training to complete the scene. Next, a geometric-semantic joint learning framework is constructed to learn the structured geometric representations of the roadside scene layout and dynamic traffic participants, achieving both global and refined point cloud completion of the roadside scene. A region-aware reconstruction loss function is designed based on the geometric-semantic joint learning framework. A point cloud completion diffusion model is established and trained based on the pseudo-ground value supervision signals and the region-aware reconstruction loss function. Finally, point cloud completion of the roadside scene is performed based on the trained point cloud completion diffusion model. Thus, this invention uses roadside visual information as compensation to estimate the pose of occluded objects, guiding the diffusion model to generate the point cloud of the entire scene, thereby restoring the local details of traffic targets in the scene while maintaining the global consistency of the roadside scene. To provide reliable all-domain perception capabilities for autonomous driving and intelligent transportation systems, so as to better optimize traffic flow and improve the driving safety of autonomous vehicles.
[0053] This invention uses the V2X-Seq dataset as an example. For a given dataset, 70% of the data in each scene is used for training, 10% for validation, and 20% for testing to perform roadside scene point cloud completion. Due to the fixed deployment of roadside LiDAR, there is long-term persistent occlusion, resulting in uncertainty in the pose (length, width, height, x, y, z coordinates, heading angle) of occluded objects. However, these occluded objects are partially visible in the roadside visual images captured by roadside cameras. Therefore, this invention uses the state-of-the-art roadside visual detector BEVHeight to estimate the pose of objects in the roadside scene. However, visual detectors based on the roadside view have detection errors and missed detections, and directly using the noisy pose estimated by this detector is insufficient to meet the global consistency and local refinement requirements of roadside scene point cloud completion. Therefore, to achieve global consistency and local refinement requirements for roadside scene point cloud completion, further, in step S1, the global optimization objective is as follows:
[0054] This invention achieves the following complete optimization objectives for the roadside scene point cloud completion method:
[0055] (1)
[0056] in Represents sparse point clouds for lidar, This represents the erroneous estimated pose from the roadside visual images and the output of the visual detector. Representing a true and complete roadside scene point cloud, This represents the object's true pose. In the above factorization, This invention addresses the problem of point cloud completion in roadside scenes, which is to be modeled in this invention. This indicates the visibility sampling of the lidar. This indicates scene composition based on object pose and pseudo-realistic templates. Given the true pose Visual likelihood of visual evidence C (image and detector output) It is a priori attitude.
[0057] The objective of this invention is to provide a given observation and Inferring the potential complete scenario and the object's true pose This is equivalent to calculating the posterior distribution. However, this posterior requires the computation of high-dimensional and complex integrals. This makes the posterior distribution difficult to calculate directly. Therefore, this invention introduces a variational approximation. And maximize its lower bound on evidence (ELBO):
[0058] (2)
[0059] Traditional methods use alternating optimization to calculate Formula 2; however, this strategy is heavily reliant on... and This leads to error accumulation. Therefore, this invention uses a mean-field approximation. Substitute into Formula 2 to separate and Alternating optimization is performed as follows:
[0060] (3)
[0061] in This represents the observation likelihood term, which measures the likelihood of a given complete scene. Below, sparse point clouds observed. The likelihood of . This represents the scene composition likelihood term, which measures the likelihood of a given object's pose. Next, synthesize the complete scene. The likelihood of . This represents the likelihood term for visual evidence, which measures the likelihood of a given object's pose. Below, visual evidence The likelihood of . This represents the attitude prior, which measures the attitude of an object. The prior probability. Represents the variational distribution of the scene entropy, Attitude variational distribution The entropy.
[0062] Furthermore, for calculating Formula 3, this invention reformulates this maximization problem as a staged minimization problem of the difference between two core probabilities. This formulation approximates... and accomplish.
[0063] Firstly, in order to define the pose target, this invention addresses the difficulty in handling generator terms. Approximation of tractable geometric likelihood This approximation will capture the observed sparse point cloud. With the object's posture Related, without needing a complete scene Under this approximation, when When minimizing the negative log-posterior, Equation 3 depends on item The probability pose difference is maximized. At this point, the optimization objective can be derived:
[0064] (5)
[0065] By minimizing this probabilistic pose difference objective function, the roadside visual detector error can be corrected, providing a reliable estimated pose of occluded objects in the roadside scene. This generates pseudo-real values of occluded objects, providing explicit supervision signals for subsequent diffusion model training to complete the scene.
[0066] Secondly, in order to define the scene target, this invention will use the corrected pose. As the actual pose input of the occluded objects in the roadside scene, it is about to Fixed in the middle Substituting into Formula 3 depends on The term is then used to obtain the final optimization objective for roadside scene point cloud completion. Maximizing this objective is equivalent to minimizing its negative term, which is defined in this invention as minimizing scene differences:
[0067] (6)
[0068] This invention designs an attitude probability difference function and a scene difference function, thereby minimizing these two functions to achieve point cloud completion for roadside scenes, and has been deployed in real-world perception closed-loop scenarios in complex traffic intersections.
[0069] Due to the detection errors and missed detections inherent in roadside visual detectors, directly using the noisy pose estimated by the detector to generate pseudo-ground truth supervision signals is insufficient to meet the requirements of global consistency and local refinement in roadside scene point cloud completion. Therefore, this invention corrects the errors and missed detections of roadside visual detectors by minimizing Equation 5, thereby generating pseudo-ground truth supervision signals with confidence. Specifically, in step S2, the pose of occluded objects in the roadside scene is corrected based on the probabilistic pose difference objective function to generate pseudo-ground truth supervision signals, which includes the following steps:
[0070] S21. This invention integrates external visual evidence and lidar geometric likelihood to calculate and minimize the probabilistic pose difference objective function to correct the pose of occluded objects in roadside scenes. Calculating and minimizing the probabilistic pose difference is an object-by-object optimization process, which minimizes the probabilistic pose difference objective function defined in Equation 5, and outputs the results from each noisy detector. Refined into high-precision attitude estimation .
[0071] Subsequently, based on the corrected posture Constructing pseudo-true point clouds It also provides a confidence score for each corrected pose. .
[0072] Specifically, this invention transforms the probabilistic pose difference objective function in Equation 5 into a maximum a posteriori estimation problem, as follows:
[0073] (7)
[0074] in, It is the log-posterior probability of the pose of a single object, representing the probability of a given observed sparse point cloud. and visual evidence Under the condition of, the pose of a single object The logarithmic probability. Visual evidence likelihood represents the likelihood of a given single object pose. At that time, the visual evidence observed The logarithmic probability of (image and detector output). It is the geometric likelihood, which acts as a pose corrector for each object, representing the geometric likelihood at a given pose. At that time, sparse point clouds were observed. The logarithmic probability. It is a pose prior, representing the pose of a single object. The prior log probability is used to impose regularization constraints on the pose.
[0075] Since the maximum a posteriori estimation function in Equation 7 is a separable optimization, this means that the pose of each object can be calculated independently. Therefore, this invention formulates the per-object energy by minimizing the negative log-likelihood term in Equation 7. Thus, the objective function for probabilistic attitude difference in Formula 5 is redefined:
[0076] (8)
[0077] The core of Formula 8 is that it doesn't treat all errors equally, but rather weights them according to their uncertainty, thus allowing more reliable cues to have a stronger impact. Specifically, the first term is the visual prior, corresponding to... It punishes Image-based detector estimation The bias; its intensity is determined by the inverse detector covariance. Control. The second term is geometric likelihood. It acts as a per-object pose corrector. It detects object categories. Select a canonical template from the ShapeNet dataset Then the converted template High-precision observation of sparse point clouds Alignment refines the pose. This factor is derived from the sensor noise variance. Modulation. Finally. This represents the prior-based pose regularization term. Therefore, the probabilistic pose difference function, which is minimized by Equation 8, corrects the pose by integrating external visual evidence and LiDAR geometric likelihood.
[0078] Due to mapping Nonlinearly dependent on Direct calculation of Formula 8 yields no closed-form solution. Therefore, this invention employs the Gauss-Newton iteration method for calculation, initialized at... This method is currently used for estimation. Template transformation Perform a first-order Taylor expansion:
[0079] (9)
[0080] Substituting linearization formula 9 into formula 8, the nonlinear minimization of formula 8 can be transformed into an iterative calculation of a linear system:
[0081] (10)
[0082] in Indicates the update increment. It is the currently estimated point-to-point residual, where This is the nearest LiDAR point. It is calculated iteratively using formula 10 and updated. Until convergence, the final object-by-object corrected pose is obtained. and confidence level Among them, confidence level It is the detector confidence level Alignment residuals and approximate posterior covariance The results of the collaborative calculations: (11)
[0083] S22. Generate a pseudo-true value point cloud with confidence based on the corrected pose and confidence level, as a pseudo-true value supervision signal. Specifically, based on the corrected pose with confidence level obtained in step S21... Based on the pose, this invention accurately removes the occluded object point cloud from the original scene by correcting the pose metadata of each object, thus achieving semantic segmentation and point cloud cleaning of the occluded target. Subsequently, based on the object template library extracted from ShapeNet and standardized through processes such as minimum bounding box correction, skewness-guided flipping, and axial reordering, the optimal matching template is dynamically selected according to the criterion of minimizing the three-dimensional scale error (|Δℓ|+|Δw|+|Δh|) after scaling the actual target length, width, and height to each template proportionally. Finally, through rigid body rotation (yaw angle alignment), two-dimensional translation (vehicle center positioning), and equal-quantity random sampling and closed-mesh resampling strategies, the standardized template is accurately restored to the size, orientation, and position of the real object, generating a high-density, geometrically consistent pseudo-ground value point cloud that is closely attached to the scene ground. This pseudo-ground value point cloud is used as a pseudo-ground value supervision signal.
[0084] This invention designs a geometric-semantic joint learning framework. This framework provides full-scene spatial geometric priors and refined semantic information of occluded objects through spatial alignment collaboration and feature-level collaboration, achieving point cloud completion for roadside scenes that balances global and refined aspects. Therefore, further, in step S3, the geometric-semantic joint learning framework is constructed, specifically including the following steps:
[0085] S31. Based on the corrected attitude, a binary dynamic scene mask matrix is constructed to distinguish between dynamically variable regions and static background regions, providing global spatial geometric priors. Roadside lidar sensors are fixedly deployed at high altitudes, thus providing spatiotemporally stable static background information, such as road surfaces and buildings. In contrast, dynamic traffic participants on the road surface exhibit localized dynamic changes, and these dynamic points are not randomly distributed but show significant temporal continuity and inherent clustering. Based on this characteristic, this invention constructs a binary dynamic scene mask matrix based on the corrected attitude obtained in step two. This matrix is used to distinguish between dynamically variable regions and static background regions, and provides global spatial geometric priors.
[0086] Specifically, The markers belong to the points of dynamic traffic participants, while This represents a static background point. The mask matrix provides crucial spatial geometric priors, guiding the diffusion model to learn the geometry of the occluded object while suppressing redundant background information. Specifically, Points belonging to dynamic traffic participants are marked, and their spatial extent, occupied volume, and external contour are explicitly defined in the current frame. This structured geometric representation provides a spatial prior to occluded objects in the roadside scene. This represents static background points. The spatial distribution and geometric contours of these points reflect the static scene along the roadside. This mask matrix provides crucial spatial geometric priors, guiding the diffusion model to learn the geometry of occluded objects while suppressing redundant background information.
[0087] S32. Parallel convolutional branches are used to extract geometric features at multiple granularities. To fully extract global features of the roadside scene point cloud and achieve global consistency in roadside scene point cloud completion, this invention designs a cross-level feature aggregation module. This module uses parallel convolutional branches to extract geometric features at multiple granularities. Specifically:
[0088] This invention designs four convolutional kernel sizes: 1×1, 3×1, 5×1, and 7×1, to output detailed features respectively. Local features Intermediate features With global contour features Multi-scale features are concatenated into Then, the learningable projection layer and batch normalization are used for fusion:
[0089] (12)
[0090] in and The parameters to be learned enable adaptive integration of cross-scale information, effectively improving the representation ability of object geometric features.
[0091] The specific network architecture is as follows:
[0092] 1) Pass the MinkUNet output data x to the first convolutional layer, which has 96 input channels, 24 output channels, and a kernel size of 1.
[0093] 2) Pass the MinkUNet output data x to the second convolutional layer, which has 96 input channels, 24 output channels, a kernel size of 3, and padding of 1.
[0094] 3) Pass the MinkUNet output data x to the third convolutional layer, which has 96 input channels, 24 output channels, a kernel size of 5, and padding of 2.
[0095] 4) Pass the MinkUNet output data x to the fourth convolutional layer, which has 96 input channels, 24 output channels, a kernel size of 7, and padding of 3.
[0096] 5) Fuse the features from steps (1) to (4) and apply batch normalization and ReLU activation function. The final output channel count is 96.
[0097] S33. Local structural relationship modeling based on anchor point grouping, outputting the final features. To enhance the contextual relevance of dynamically variable regions, this paper proposes local structural relationship modeling based on anchor point grouping. For input features... First, the query, key, and value are obtained through projection:
[0098] (13)
[0099] Then, a subset of points are randomly selected as the anchor point set. For each point Calculate its characteristic space Euclidean distance to all anchor points, assign each point to the nearest anchor point group, and within the same group, calculate the distance for each point. Select k neighbors from the group to form a local neighborhood set. ,in .
[0100] Next point With each of his neighbors , obtain attention score To highlight the information of dynamically variable regions, dynamic variable region enhancement is introduced. Then, after Softmax normalization, we get (14)
[0101] Aggregate value vectors based on normalized weights. Combine the residuals with the original features and linearly map them back to the original dimensions. .
[0102] Finally, a mask-based weighted and nonlinear transformation is applied again to perform a second enhancement on the dynamic traffic participant region, and the final features are output:
[0103] (15)
[0104] This invention uses the geometric-semantic joint learning framework in step S3 to minimize the scene difference optimization objective of formula 6. By training a diffusion model and using the loss function designed in this invention, it achieves the global consistency and local refinement requirements of roadside scene point cloud completion. Specifically, in step S4, the region-aware reconstruction loss function consists of three parts: mean squared error, dynamic scene mask weighted loss, and regularization term. The dynamic scene mask weighted loss is:
[0105] (16)
[0106] in , For dynamic scene mask matrix, and These are the weight hyperparameters for dynamically changing local regions and static background regions, respectively, allowing flexible control over the loss contribution of different semantic regions. To make the predicted noise closer to the real data, this invention introduces regularization loss.
[0107] (17)
[0108] The final total loss is And train the model using the diffusion model based on the loss function.
[0109] Further, in step S5, the specific training process is as follows: using the original sparse point cloud and pseudo-ground value supervision signal as conditions, Gaussian noise is progressively added using a diffusion model forward pass, and training optimization is performed using a region-aware reconstruction loss function. Specifically, this invention trains on two A6000 images for 50 epochs. For the optimizer, the Adam optimizer is used, and the learning rate is set to 10⁻⁻⁶. 4 The decay rate is halved every 5 epochs, with a decay rate of 10⁻ 4 The batch size is set to 4. For the diffusion model parameters, this paper uses linear scheduling with 1000 diffusion steps. and The background region loss weight is 3.0, and the dynamic variable region loss weight is 8.0. The voxel size is 0.05m. For the number of diffusion points, since the number of occluded objects in each frame of the point cloud is different, this invention performs dynamic point diffusion. During training, using the original sparse point cloud and pseudo-ground value supervision signal as conditions, Gaussian noise is progressively added to the forward pass of the diffusion model. The loss function in step S4 is used for training optimization, and validation is performed every 5 epochs to check the performance of the diffusion model. After training, the final training weight file is saved.
[0110] Further, in step S6, point cloud completion for the roadside scene is performed based on the trained point cloud completion diffusion model. Specifically, this includes: using the back-inference denoising of the point cloud completion diffusion model to complete the point cloud of a single frame of the roadside scene. This invention uses the V2X-Seq roadside scene dataset as a benchmark and uses the back-inference denoising of the diffusion model to complete the point cloud of a single frame of the scene, obtaining the following... Figure 3 The image shown is the completed roadside scene. During the completion process, a classification-free guidance strategy is employed on the input single-frame sparse point cloud to enhance the guidance strength against false ground truth values, as detailed below:
[0111] (18)
[0112] in This represents the final noise prediction after adjustment without a classifier, used to guide the denoising process. This represents the diffusion model θ given the current diffusion state. Time step t, sparse point cloud Y, and pseudo-real values The noise predicted below. Indicates the intensity of unclassified guidance. This indicates the removal of false truth conditions. Furthermore, by minimizing the probability pose difference, this invention assesses the guidance strength. Dynamic adjustments are made. Specifically, in high-confidence regions, the guiding strength s is increased to make the model more strictly adhere to the conditions. In the low confidence region, the value is reduced. This encourages the model to rely more on unconditionally generated paths, thereby improving overall robustness. For the completed results, this invention uses chamfer distance, Jensen-Shannon divergence in bird's-eye view and 3D space, and voxel cross-sectional area ratios at 0.5m and 0.2m as evaluation metrics to assess the results.
[0113] This invention also provides a roadside scene point cloud completion system based on a diffusion model, characterized by comprising a perception data acquisition module, an edge computing module, and a storage module; the perception data acquisition module includes a roadside 128-line LiDAR and an 8-megapixel roadside camera, used to acquire the original sparse point cloud and visual image of the roadside scene; the edge computing module includes an industrial control computer and a 24GB graphics processing unit (GPU), the industrial control computer and the perception data acquisition module are connected via gigabit Ethernet, used to process the acquired data and run the diffusion model algorithm to perform attitude correction, feature extraction, and point cloud completion tasks; the storage module is used to store the roadside visual detector model, the roadside scene point cloud completion model based on the diffusion model, a preset object template library, and a trained diffusion model weight file.
[0114] The above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the invention. Those skilled in the art will understand that implementing all or part of the above-described embodiments and making equivalent changes in accordance with the claims of the present invention are still within the scope of the invention.
Claims
1. A point cloud completion method for roadside scenes based on a diffusion model, characterized in that, Includes the following steps: S1. Establish a global optimization objective, and determine the probabilistic attitude difference objective function based on the global optimization objective; S2. Correct the pose of occluded objects in the roadside scene based on the probability pose difference objective function and generate a pseudo-true value supervision signal. S3. Construct a joint learning framework for geometric and semantic knowledge; S4. Design a region-aware reconstruction loss function based on the aforementioned geometric-semantic joint learning framework; S5. Establish a point cloud completion diffusion model, and train the point cloud completion diffusion model based on the pseudo-true value supervision signal and the region perception reconstruction loss function. S6. Perform point cloud completion for roadside scenes based on the trained point cloud completion diffusion model.
2. The method for roadside scene point cloud completion based on a diffusion model according to claim 1, characterized in that, In step S1, the global optimization objective is as follows: ; in, Represents sparse point clouds for lidar, This represents the erroneous estimated pose from the roadside visual images and the output of the visual detector. Representing a true and complete roadside scene point cloud, Represents the true posture of an object; This indicates the visibility sampling of the lidar. This indicates scene composition based on object pose and pseudo-realistic templates. Given the true pose Visual likelihood of visual evidence C It is a priori attitude.
3. The method for roadside scene point cloud completion based on a diffusion model according to claim 1, characterized in that, In step S2, the pose of occluded objects in the roadside scene is corrected based on the probabilistic pose difference objective function to generate a pseudo-ground value supervision signal, which specifically includes the following steps: S21. Calculate the objective function that minimizes the probability pose difference, integrate external visual evidence and LiDAR geometric likelihood to correct the pose of occluded objects in the roadside scene, and obtain the corrected pose and confidence level for each object. S22. Generate a pseudo-true value point cloud with confidence based on the corrected attitude and confidence level, as a pseudo-true value supervision signal.
4. The method for roadside scene point cloud completion based on a diffusion model according to claim 3, characterized in that, In step S3, a joint learning framework for geometry and semantics is constructed, which specifically includes the following steps: S31. Based on the corrected posture, construct a binary dynamic scene mask matrix to distinguish between dynamic variable regions and static background regions, and provide global spatial geometric priors; S32. Parallel convolutional branching is used to extract geometric features at multiple granularities; S33. Model local structural relationships based on anchor point grouping and output the final features.
5. The method for roadside scene point cloud completion based on a diffusion model according to claim 4, characterized in that, In step S4, the region-aware reconstruction loss function consists of three parts: mean squared error, dynamic scene mask weighted loss, and regularization term.
6. The method for roadside scene point cloud completion based on a diffusion model according to claim 5, characterized in that, The dynamic scene mask weighted loss is: ; in , and These are the weight hyperparameters for the dynamically changing local region and the static background region, respectively. It is a dynamic scene mask matrix; The regularization loss is: ; The total loss is: 。 7. The method for roadside scene point cloud completion based on a diffusion model according to claim 1, characterized in that, In step S5, the specific training process is as follows: using the original sparse point cloud and pseudo-true value supervision signal as conditions, Gaussian noise is gradually added using the forward process of the diffusion model, and training optimization is performed using the region-aware reconstruction loss function.
8. The method for roadside scene point cloud completion based on a diffusion model according to claim 1, characterized in that, In step S6, point cloud completion for the roadside scene is performed based on the trained point cloud completion diffusion model. Specifically, this includes using the reverse inference denoising of the point cloud completion diffusion model to complete the point cloud of a single frame of the roadside scene.
9. The method for roadside scene point cloud completion based on a diffusion model according to claim 8, characterized in that, During the completion process, a classification-free guidance strategy is adopted for the input single-frame roadside scene point cloud, as follows: ; in, This represents the final noise prediction after adjustment without a classifier, used to guide the denoising process; This represents the diffusion model θ given the current diffusion state. Time step t, sparse point cloud Y, and pseudo-real values The noise predicted below; Indicates the intensity of unclassified guidance. This indicates the condition for removing false truth values.
10. A roadside scene point cloud completion system based on a diffusion model, applied to the roadside scene point cloud completion method based on a diffusion model as described in any one of claims 1-9, characterized in that, It includes a perception data acquisition module, an edge computing module, and a storage module. The perception data acquisition module includes a roadside lidar and a roadside camera, used to acquire the original sparse point cloud and visual images of the roadside scene. The edge computing module includes an industrial control computer and a graphics processing unit. The industrial control computer is communicatively connected to the perception data acquisition module and is used to process the acquired data and run a diffusion model algorithm to perform attitude correction, feature extraction, and point cloud completion tasks. The storage module is used to store the roadside visual detector model, the roadside scene point cloud completion model based on the diffusion model, the preset object template library, and the trained diffusion model weight file.