Kinetics-based intelligent driving LiDAR point cloud generation method
By combining LiDAR point cloud generation methods with dynamics, utilizing velocity scalars and heading angles for modeling, and combining Maxwell–Boltzmann distribution and Langevin dynamics, the problem of lack of physical constraints in existing methods is solved, achieving more realistic and continuous point cloud generation, adapting to complex traffic scenarios, and improving the stability of intelligent driving simulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XI'AN POLYTECHNIC UNIVERSITY
- Filing Date
- 2026-04-14
- Publication Date
- 2026-05-29
AI Technical Summary
Existing 3D point cloud generation methods lack physical motion constraints, have insufficient modeling of directional changes, and have fixed diffusion intensity that cannot reflect dynamic changes in the scene, resulting in insufficient realism of the generated point cloud data in complex dynamic scenes.
The dynamics-based intelligent driving LiDAR point cloud generation method weakens the target motion into a velocity scalar and heading angle, combines Maxwell-Boltzmann equilibrium distribution and Langevin dynamics to construct a diffusion generation process, and uses a lightweight denoising network to achieve point cloud generation with arbitrary number of points, fusing dynamic and static point sets.
It improves the realism and continuity of the generated results, enhances the model's adaptability to complex traffic scenarios, avoids inter-frame jitter and sudden changes in direction, and improves the stability of intelligent driving simulation and data generation.
Smart Images

Figure CN122115740A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of autonomous driving and 3D perception technology, and relates to a dynamics-based intelligent driving LiDAR point cloud generation method. Background Technology
[0002] Currently, in autonomous driving scenarios, LiDAR 3D point cloud data is the core information source for environmental perception, target detection, and path planning. With the increasing demands for autonomous driving simulation training and data augmentation, generating high-quality 3D point cloud data that conforms to real physical characteristics has become a research hotspot. Existing point cloud generation and completion methods mainly include geometric interpolation-based methods, physical modeling-based methods, and data-driven deep learning methods.
[0003] Traditional geometric interpolation or rule-based modeling methods typically reconstruct missing point clouds through neighborhood interpolation, surface fitting, or rule-based point completion. However, these methods heavily rely on local geometric assumptions, making it difficult to characterize dynamic targets and occlusion relationships in complex scenes. The generated results lack realistic radar statistical characteristics and are not applicable to continuous time series scenarios.
[0004] Physical modeling-based methods typically simulate point clouds by constructing vehicle kinematics models, scene geometry models, or ray propagation models. While these methods can reflect the motion patterns and spatial relationships of objects to some extent, their drawbacks include numerous model parameters, complex modeling, poor adaptability, and difficulty in depicting multi-objective dynamic interactions in complex traffic scenarios. They are particularly inadequate in portraying motion processes with significant directional changes, such as vehicle turning, lane changing, and U-turns. Furthermore, they struggle to simulate the random noise and statistical perturbation characteristics generated during real LiDAR scanning.
[0005] In recent years, data-driven deep learning methods have been widely applied to point cloud generation and completion tasks. Generative models, especially diffusion models and generative adversarial networks (GANs), have achieved good results in 3D shape generation and completion tasks by progressively generating point cloud data from Gaussian noise. However, most existing diffusion models employ pre-defined fixed noise scheduling strategies, with diffusion intensity depending solely on the number of time steps, without considering the physical motion state and energy changes in the scene, resulting in a lack of physical constraints in the generation process.
[0006] Meanwhile, existing diffusion-based video or sequence generation methods typically generate continuous point cloud frames using frame-by-frame independent sampling or weak temporal constraints, neglecting the continuous dynamic evolution of vehicles, pedestrians, and environmental structures in the real physical world. In particular, they fail to explicitly consider target heading changes, turning rate evolution, and curved path motion characteristics, easily leading to problems such as inter-frame flickering, target drift, direction jumps, or motion discontinuities. Furthermore, most models treat static environments and dynamic targets uniformly, failing to distinguish the differences in motion mechanisms between static objects such as roads and trees and dynamic objects such as vehicles and pedestrians, resulting in physical inconsistencies in the generated results over long time sequences.
[0007] On the other hand, in real-world autonomous driving scenarios, the vehicle's motion state, changes in target speed and direction, and the system's motion energy level directly affect the distribution characteristics of LiDAR point clouds. For example, changes in point cloud density, local morphology, and scanning distortion become more pronounced during high-speed driving, continuous turning, or U-turns. However, existing methods fail to integrate the dynamic models from statistical physics with the diffusion generation process, lacking a mechanism for adaptively adjusting the generation intensity based on translational motion energy and directional change energy. This results in insufficient realism of the generated point cloud data in complex dynamic scenarios.
[0008] To address the above problems, this invention proposes a 3D point cloud generation method that simultaneously combines vehicle motion dynamics, continuous target motion modeling, target orientation change modeling, and statistical physical diffusion mechanisms. This method improves the stability, directional continuity, and realism of continuous-time scene generation while ensuring physical consistency. Summary of the Invention
[0009] The purpose of this invention is to provide a dynamic-based intelligent driving LiDAR point cloud generation method to solve the problems of existing 3D point cloud generation methods, such as lack of physical motion constraints, insufficient modeling of direction changes, and fixed diffusion intensity that cannot reflect dynamic changes in the scene.
[0010] The technical solution adopted in this invention is a dynamics-based intelligent driving LiDAR point cloud generation method. This method acquires a continuous LiDAR point cloud sequence and performs standardization processing, weakening the target motion from a three-dimensional velocity vector to a low-dimensional motion state variable dominated by a velocity scalar. It also introduces heading angle and turning rate to characterize directional changes. A Maxwell-Boltzmann equilibrium distribution of velocity and statistical constraints on turning rate are constructed. Based on Langevin dynamics, the motion state is continuously evolved, yielding translational kinetic energy from velocity and turning energy from turning rate. These two are then jointly mapped to diffusion intensity. ; build The diffusion generation process is controlled, and a lightweight denoising network is used to generate point clouds with arbitrary point counts. Finally, a continuous sequence is generated, and the dynamic point set and static point set are fused to output a continuous LiDAR point cloud sequence. The specific operation steps are as follows:
[0011] Step 1: Acquire a series of consecutive LiDAR 3D point cloud sequences collected by the intelligent driving vehicle in the road environment, and perform standardization processing; Step 2: Reduce the target motion from a three-dimensional velocity vector to a low-dimensional motion state variable dominated by a velocity scalar, and establish motion intensity and direction change variables; Step 3: Construct the Maxwell–Boltzmann equilibrium distribution of the velocity, establish statistical constraints on the turning rate, and sample the initial motion state variables; Step 4: Based on Langevin dynamics, continuously evolve and update the motion state variables, obtain translational kinetic energy from velocity, obtain turning energy from turning rate, and jointly adjust diffusion intensity; Step 5: Construct a diffusion generation process regulated by diffusion intensity; Step 6: Construct a lightweight denoising network and an arbitrary point generation mechanism; Step 7: Continuous sequence generation and dynamic / static fusion output.
[0012] The invention is further characterized in that, In step 1, the acquired LiDAR continuously acquired 3D point cloud sequence data category c includes multiple categories of original point cloud information such as vehicles, pedestrians, roads, and trees. First, the original point cloud is subjected to coordinate normalization and scale standardization processing, and the continuous multi-frame point cloud is aligned to a unified reference coordinate system according to the vehicle pose. The point cloud is divided into dynamic point sets and static point sets, which are used for subsequent differentiated motion modeling and diffusion generation modeling, respectively. The vehicle pose is the vehicle's own position and orientation information at the sampling time, which is used to transform the point cloud collected at different times to the same reference coordinate system. Step 1.1: Acquire a continuous multi-frame sequence of 3D point cloud data collected by the intelligent driving vehicle in the road environment. Let the first frame be the... k The frame point cloud is: (1) In the formula, For the first k Frame point cloud 3D coordinates Points; Let the first k The frame of the vehicle position pose is T k Then historical frames k - l Point transformation to frame k Coordinate system: (2) , They represent the first kFrame, First k The rigid body transformation matrix corresponding to the vehicle's pose in frame l. Indicates the first k The inverse of the frame pose matrix; This indicates that the historical frames are... k The point in l is transformed to the current point. k The result after the frame coordinate system is restored; Indicates to the first k Points seen in the l-frame are first transformed to a unified world coordinate system or reference coordinate system, and then multiplied by... This means transforming a point that was already in the common coordinate system to the next coordinate system. k In the local coordinate system of the frame; Rigid body transformation matrix , No. k The rotation matrix of a frame represents the change in the vehicle's orientation at that moment. The translation vector of the k-th frame represents the change in the vehicle's position at that moment; Step 1.3: Normalization process: (3) In the formula, For the first k The mean vector of the 3D coordinates of the frame point cloud. ,in Let be the standard deviation vector of the 3D coordinates of the point cloud in the k-th frame; Step 1.4: Delineation of Dynamic and Static Elements (4) in, It is a dynamic point set, including vehicles and pedestrians. It is a static point set, including roads, buildings, trees, etc.
[0013] Step 2 is as follows: With rate Describe the intensity of the target's translational motion in terms of heading angle. With steering ratio Describe the trend of change in the target's direction; When the target center is available for consecutive frames, the velocity observation value of the dynamic target is defined as: (5) in, Indicates the first k The first frame j The central location of a dynamic target Indicates the first k -1 frame jThe central location of a dynamic target; Indicates the time interval between adjacent frames; Indicates the first k The first frame j Rate observations of a dynamic target; The heading angle observation value of a dynamic target is defined as: (6) in, Indicates the first j The dynamic objective is from the first k frame l to the k The displacement component of the frame in the y-direction is denoted as ; Indicates the first j The dynamic objective is from the first k frame l to the k The displacement component of the frame in the x-direction is denoted as ; It is the heading angle, indicating the target's direction of travel. k l frame to k Which direction to move between frames; using Can be directly based on The sign is used to determine the direction, thus obtaining the angle in the correct quadrant. The output range is: ; The observed turning rate of a dynamic target is defined as: (7) Static targets are directly defined as follows: .
[0014] Step 3 is as follows: Step 3.1: Set the equivalent mass for dynamic or static target category c. With temperature parameters Define rate Maxwell–Boltzmann distribution: (8) In the formula, Boltzmann's constant, Both temperature parameters are used to characterize the statistical scale of the intensity of this type of motion and to introduce class differences into the dynamic system; The initial statistical distribution of the turning rate can be set as a zero-mean Gaussian distribution: (9) The steering rate is a signed state variable describing the change in target direction, with positive and negative values representing different steering directions. A mean of zero avoids introducing prior biases for left or right turns during the initial sampling phase. For category c The turning rate scale parameter is used to characterize the statistical strength of the change in the direction of targets of this category; Step 3.2: Initial Rate Sampling For each dynamic objective The initial motion state is sampled from equations (14) and (15): (10) For static categories, simply let: .
[0015] Step 3.3: Express the Maxwell distribution in potential energy form, and define the joint potential energy function corresponding to the velocity and turning rate: (11) In the formula, To prevent small constants from becoming numerically unstable; The velocity balance distribution and turning rate statistical constraints established in step 3 enable the dynamic target to simultaneously possess translational motion continuity and directional change continuity during subsequent continuous generation processes.
[0016] In step 4, using the joint potential energy function obtained in step 3 as the target equilibrium state, the motion state variables are discretized and updated frame by frame using overdamped Langevin dynamics, and a non-negativity constraint is applied to the rate, as follows: Step 4.1: Using the joint potential energy function obtained in Step 3 as the target equilibrium state, the motion state variables are discretized and updated frame by frame using overdamped Langevin dynamics to update the dynamic target. The frame-by-frame discretization update form of its rate and turning rate is as follows: (12) (13) in, and These are the diffusion coefficients for velocity and turning rate, respectively. For a long walk; and It is a zero-mean Gaussian random variable.
[0017] Due to the rate Perform a nonnegative projection on the output of equation (12): The target heading angle is continuously updated according to the turning rate. ; Note: u represents the scalar magnitude of the target's translational motion intensity, hence the velocity. The Langevin discrete update corresponding to equation (12) contains a random perturbation term, and intermediate results with values less than zero may occur after the update. To ensure that the motion state variables satisfy the physical feasibility constraints, a non-negative projection is applied to the output of equation (12), i.e.: That is, the rate at the next moment calculated by equation (12) One more revision: If , then it remains unchanged; If , then it is truncated to 0.
[0018] Step 4.2: Obtain the translational kinetic energy from the velocity, and the turning energy from the turning rate. For a dynamic target j of category c, calculate the translational kinetic energy based on the velocity scalar: (14) Then, the steering energy is calculated based on the steering ratio: (15) In the formula, For category The equivalent moment of inertia parameter is used to characterize the inertial scale of the target's orientation change; The total kinetic energy of the target is defined as: (16) in, and is a non-negative weighting coefficient used to adjust the relative contribution of translational and directional motion to the diffusion intensity; for static targets, directly set: .
[0019] Step 4.3: Adjust the diffusion intensity based on the total kinetic energy
[0020] Use a single per target or per frame And obtained by the monotonic mapping of total kinetic energy: (17) In the formula, It is a monotonically increasing mapping function, which makes targets with high total kinetic energy have a greater diffusion intensity, and targets with low motion or static energy have a smaller diffusion intensity; (18) in, The energy scale constant is
[0021] The diffusion intensity of the entire frame is taken as a weighted average of the dynamic target diffusion intensity: (19) in, Indicates the first k Number of dynamic targets per frame Take the target point weight or category weight; in a static scene, It automatically approaches zero.
[0022] In step 5, the diffusion step number S is set to the single diffusion obtained in step 4. The basic noise schedule is modulated, a forward diffusion process for the point cloud is defined for training, and a corresponding reverse denoising generation process is executed during the inference phase to generate the point cloud; specifically as follows: Step 5.1: β Controlled diffusion noise scheduling Let the basic scheduling be A single energy obtained from physical energy Used to modulate the noise intensity at each step: (20) in, Indicates the reference diffusion intensity. This represents the sensitivity coefficient, and its value range is... ; Represents the clipping function; Indicates the upper and lower boundaries of the clipping; Step 5.2: Let For the first k The frame-normalized point cloud tensor, forward diffused, is as follows: (twenty one) Its closed form is: (twenty two) in, The mean of forward diffusion is The variance is ; Step 5.3: Reverse Denoising (twenty three) Where θ is the denoising network; the mean of the inverse denoising is The variance is .
[0023] Step 6 constructs a lightweight denoising network for the reverse diffusion model to predict noise terms and progressively generate 3D point clouds. The denoising network employs a reverse diffusion network structure composed of a nine-layer MLP variant, and uses diffusion condition vectors to gate and modulate the feature updates of each layer, thereby achieving stable denoising generation while maintaining network lightweightness. An arbitrary number of points can be specified during inference. N kThe point cloud is generated by performing S-step reverse denoising after initialization with Gaussian noise.
[0024] The denoising network uses the point cloud state at the current diffusion step. The input is the noise term prediction. This is used for reverse updates; the input of the first layer of the network is the 3D coordinates of the point cloud, and a conditional vector C is introduced to conditionally modulate the network; from the lth layer to the th... l The update format for +1 layer is:
[0025] in, and These represent the network input and output features, respectively, and ⊙ denotes element-wise multiplication. Weight parameters; C Diffusion condition vector The diffusion condition vector C Build as ; in, c These are conditional latent features, extracted from partial point clouds, historical frame point clouds, or other prior information by an encoder. For diffusion intensity embedding, It extends a single β to a triaxially consistent embedding form. .
[0026] The preferred layer width of the nine-layer MLP is... LeakyReLU is preferred for inter-layer nonlinearity, and BatchNorm is preferred for the first eight layers to stabilize training and improve convergence speed. The network's final output is a point-level 3D prediction, which is used to construct noise term predictions and perform inverse denoising updates.
[0027] Compared with the prior art, the beneficial effects of the present invention are: (1) Introducing Langevin dynamics and Maxwell–Boltzmann statistical physics model into the LiDAR point cloud generation process to realize scene modeling under physical constraints and improve the realism and continuity of the generation results.
[0028] (2) Construct a diffusion intensity adjustment mechanism driven by translational kinetic energy and steering energy to enable dynamic and static targets to be modeled differently under a unified framework, thereby enhancing the model’s adaptability to complex traffic scenarios, especially turning, lane changing and U-turn scenarios.
[0029] (3) A dynamic propagation method for continuous time series is proposed to achieve consistent generation of motion trends and direction change trends between multiple frames of point clouds, effectively avoiding the jitter, jump, sudden change of direction and time discontinuity problems that occur in the traditional frame-by-frame generation method, and improving the stability and reliability of intelligent driving simulation and data generation. Attached Figure Description
[0030] Figure 1 This is the overall flowchart of the dynamics-based intelligent driving LiDAR point cloud generation method of the present invention; Figure 2 This is a schematic diagram of a lightweight denoising network structure and conditional gating modulation; Figure 3 This is a flowchart of the diffusion intensity generation process based on joint modeling of motion states and modulation of total motion energy. Detailed Implementation
[0031] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0032] Example 1 The present invention provides a dynamic-based LiDAR point cloud generation method for intelligent driving, such as... Figure 1 As shown, the details are as follows: Step 1: Acquire a series of consecutive LiDAR 3D point cloud sequences collected by the intelligent driving vehicle in the road environment, and perform standardization processing; Step 2: Reduce the target motion from a three-dimensional velocity vector to a low-dimensional motion state variable dominated by a velocity scalar, and establish motion intensity and direction change variables; Step 3: Construct the Maxwell–Boltzmann equilibrium distribution of the velocity, establish statistical constraints on the turning rate, and sample the initial motion state variables; Step 4: Based on Langevin dynamics, continuously evolve and update the motion state variables, obtain translational kinetic energy from velocity, obtain turning energy from turning rate, and jointly adjust diffusion intensity; Step 5: Construct a diffusion generation process regulated by diffusion intensity; Step 6: Construct a lightweight denoising network and an arbitrary point generation mechanism; Step 7: Continuous sequence generation and dynamic / static fusion output.
[0033] Example 2 Based on Example 1, Step 1 is as follows: Step 1: Acquire a continuous multi-frame LiDAR 3D point cloud sequence collected by the intelligent driving vehicle in the road environment. Align the historical frame vehicle point clouds to a unified reference coordinate system using the vehicle's pose and perform mean-variance normalization on the point cloud coordinates. Further divide the point cloud into dynamic point sets (vehicles, pedestrians, etc.) and static point sets (roads, buildings, trees, etc.) for subsequent differential modeling of rate, direction changes, and diffusion generation.
[0034] (1) Data acquisition: Acquire a continuous multi-frame sequence of 3D point clouds collected by the intelligent driving vehicle in the road environment. Let the point cloud of the k-th frame be: (1) In the formula, In three-dimensional coordinates, The number of points in the point cloud at frame k.
[0035] (2) Coordinate Alignment: To ensure that the static structure remains approximately fixed in the sequence and the dynamic target exhibits continuous motion, it is preferable to align the point clouds of multiple frames to the same reference coordinate system using the vehicle's pose. Let the vehicle pose of the k-th frame be T. k The formula for transforming the points of historical frame kl to the coordinate system of the kth frame is: (2) Coordinate alignment means unifying point clouds captured from frames at different times into the same reference coordinate system. . Representing the k-th frame and the k-th frame respectively. The rigid body transformation matrix corresponding to the vehicle's pose in frame l. It represents the inverse matrix of the pose matrix in the k-th frame. This indicates that the historical frame k is... The result of transforming a point in frame l to the coordinate system of the current k-th frame.
[0036] This means first put the k-th... Points seen in frame l are first transformed to a unified world coordinate system or reference coordinate system, and then multiplied by... This means transforming the points that were already in the common coordinate system back to the local coordinate system of the k-th frame. This ensures that the point clouds of all historical frames before the k-th frame are aligned with the coordinate system of the k-th frame.
[0037] Rigid body transformation matrices are usually written as: , The rotation matrix of the k-th frame represents the change in the vehicle's orientation at that moment. The translation vector of the k-th frame represents the change in the vehicle's position at that moment.
[0038] (3) Normalization process: (3) In the formula, Let be the mean vector of the 3D coordinates of the point cloud in the k-th frame. in Let be the standard deviation vector of the 3D coordinates of the point cloud in the k-th frame. Normalization can make the data distribution across different frames more stable, making the network easier to train and reducing the impact of scale differences between different scenes.
[0039] (4) Division of movement and stillness: (4) in, It is a dynamic point set (vehicles, pedestrians, etc.). This is a static point set (roads, buildings, trees, etc.). The static / dynamic separation is used to distinguish dynamic targets from static backgrounds, enabling dynamic targets to be modeled for motion states, while static targets maintain structural stability. This allows for subsequent differentiated generation and improves the physical consistency and temporal stability of continuous point cloud sequences.
[0040] Note: Step 1 is mainly used to statistically analyze the spatiotemporal characteristics of real LiDAR during the training phase; real point clouds are not required as input during the inference generation phase.
[0041] Example 3 Based on Example 2, step 2 is as follows: Step 2: To adapt to lightweight models and fast inference, the target motion is simplified from a three-dimensional velocity vector to a low-dimensional motion state variable mainly composed of a one-dimensional "velocity scalar". Furthermore, heading angle and turning rate are introduced to jointly characterize the intensity and direction change characteristics of the target's translational motion. During the training phase, the velocity, heading angle, and turning rate observations of the dynamic target can be obtained from the statistical analysis of the target center displacement in consecutive frames. The velocity and turning rate of the static structure are set to zero to ensure the stability of the static structure in the generated sequence.
[0042] (1) Rate statistics during training phase When the training data contains the target centers of consecutive frames, the velocity observation of a dynamic target is defined as: (5) in, This indicates the center position of the j-th moving target in the k-th frame. This indicates the time interval between adjacent frames.
[0043] Furthermore, the heading angle observation value of the dynamic target is defined as: (6) in, The j-th dynamic objective starts from the k-th... The displacement component in the y-direction from frame l to frame k is denoted as... ; Indicates that the j-th dynamic target starts from the k-th target. The displacement component in the x-direction from frame l to frame k is denoted as... ; This is the heading angle, indicating the direction the target moved between these two frames. (Using...) Can be directly based on The sign of the pointer determines the direction, yielding the angle in the correct quadrant, making it more suitable for defining the heading angle. Its typical output range is: .
[0044] (2) The observed turning rate of the dynamic target is defined as: (7) The turning rate represents the rate of change of the heading angle over time, which is the speed of turning.
[0045] Static targets are directly defined as follows: .
[0046] Note: In step 2, at the rate Describe the intensity of the target's translational motion in terms of heading angle. With steering ratio It describes the changing trend of the target's direction, so that subsequent models can not only characterize the target's "speed of movement", but also the target's directional evolution behavior such as "turning, turning and U-turning".
[0047] Example 4 Based on Example 3, step 3 is as follows: Step 3: To ensure that the velocity modeling of dynamic targets conforms to real traffic scenarios and avoids distortion of the network learning target and deviation of the generated results from the real traffic speed distribution due to arbitrary sampling of initial velocity values in the absence of physical statistical priors, this invention performs differentiated modeling of targets according to categories. Equivalent mass parameters and temperature parameters are set for vehicle, pedestrian, and static categories respectively, establishing a category-conditional velocity Maxwell-Boltzmann equilibrium distribution. The temperature parameter can be obtained by fitting the velocity statistical results during the training phase or set by engineering priors. For dynamic categories, during the inference generation phase, the initial velocity is sampled from the corresponding category's Maxwell-Boltzmann equilibrium distribution, and the initial turning rate is sampled from the turning rate statistical constraint distribution. For static categories, both velocity and turning rate are set to zero. Furthermore, the Maxwell-Boltzmann equilibrium distribution is written in potential energy form and used as the initial sampling basis for subsequent Langevin dynamic evolution and the target equilibrium distribution throughout the evolution process, thus ensuring that the generated motion state satisfies both category differences and remains consistent with the statistical laws of real traffic motion.
[0048] (1) Maxwell–Boltzmann rate distribution Set equivalent mass for category c (e.g., vehicles / pedestrians / static). With temperature parameters Define rate Maxwell–Boltzmann distribution: (8) In the formula, This is Boltzmann's constant. Due to the influence of vehicles, pedestrians, and static objects... They are all different, and the function of these two parameters is to characterize the statistical scale of a certain type of motion intensity, which naturally introduces "class difference" into the dynamic system, making the regulation of β naturally class-aware.
[0049] The turning rate is a signed state variable describing the change in target direction, with positive and negative signs representing different turning directions. A mean of zero avoids introducing prior biases for left or right turns during the initial sampling phase. The initial statistical distribution of the turning rate can be set to a zero-mean Gaussian distribution. (9) in, is the turning rate scale parameter for category c, used to characterize the statistical intensity of the change in the target direction for that category.
[0050] (2) Initial rate sampling (used during inference generation) For each dynamic objective Category The initial motion state is sampled from equations (14) and (15): (10) Static categories can be directly defined as follows: .
[0051] (3) Write the Maxwell distribution in the form of potential energy (to prepare for Langevin evolution) Define the joint potential energy function corresponding to the velocity and turning rate (ignoring constant terms independent of the variables) as follows: (11) In the formula, To prevent small constants from being numerically unstable.
[0052] The purpose of writing the Maxwell–Boltzmann distribution in potential energy form is to transform the statistical distribution constraints of velocity and turning rate into a target potential function that can be directly used for Langevin dynamics evolution, so that the motion state variables evolve towards the target equilibrium state along the potential energy gradient under random perturbation; at the same time, by constructing a joint potential energy function corresponding to velocity and turning rate, the translational motion and orientation change of the target can be jointly constrained under a unified framework, thereby improving the smoothness, stability and physical consistency of motion evolution between consecutive frames.
[0053] By establishing the velocity balance distribution and turning rate statistical constraints in step 3, the dynamic target can simultaneously possess translational motion continuity and directional change continuity during subsequent continuous generation processes.
[0054] Example 5 Based on Example 4, step 4 is as follows: Step 4: Using the combined potential energy obtained in Step 3 as the target equilibrium state, the motion state variables are discretized and updated frame by frame using overdamped Langevin dynamics, and a non-negativity constraint is applied to the velocity. The dynamic target translational kinetic energy is calculated from the velocity, and the turning energy is calculated from the turning rate. Then, the translational kinetic energy and the turning energy are jointly mapped to the diffusion intensity. And preferably for dynamic targets across the entire frame. Weighted aggregation is performed to obtain "the entire frame as a single unit". ", in order to further reduce the conditional dimension.
[0055] (1) The joint dynamics evolves towards the equilibrium state (with the Maxwell distribution as the target equilibrium state) like Figure 3 As shown, to transform the static statistical distribution established in step 3 into a propagable motion trajectory between consecutive frames, and to ensure that the target's velocity and orientation changes evolve smoothly over time and continuously approach the statistical equilibrium state, this invention employs overdamped Langevin dynamics based on the potential energy function. The Langevin dynamics only act on the low-dimensional motion state space composed of velocity u and turning rate ω. By continuously updating u and ω over time, and recursively deriving the heading angle θ from ω, continuous propagation of the translation and orientation states from the k-th frame to subsequent frames is achieved. Subsequently, the corresponding energy is calculated from the low-dimensional state, and the diffusion generation process is modulated, rather than directly updating all points in the high-dimensional point cloud coordinate space. This reduces computational complexity, improves numerical stability, and enhances the temporal smoothness, physical rationality, and engineering feasibility of continuous LiDAR point cloud generation.
[0056] For dynamic targets Its rate and turning rate, in continuous form, are expressed as follows:
[0057]
[0058] In the formula, and These are the diffusion coefficients for velocity and turning rate, respectively. and These are mutually independent one-dimensional Wiener processes.
[0059] Let's take a walk. Then the frame-by-frame discretization update can be written as: (12) (13) Among them, the drift term: It will bring the excessively high speed and turning rate back, preventing the system from drifting indefinitely. Diffusion term: It will introduce random disturbances. and It is a zero-mean Gaussian random variable. Equations (12) and (13) show that the motion state is not rigidly determined, and there will be perturbations, fluctuations and randomness between frames.
[0060] Note: u represents the scalar magnitude of the target's translational motion intensity, hence the velocity. The Langevin discrete update corresponding to equation (12) contains a random perturbation term, and intermediate results with values less than zero may occur after the update. To ensure that the motion state variables satisfy the physical feasibility constraints, a non-negative projection is applied to the output of equation (12), i.e.: That is, the rate at the next moment calculated by equation (12) One more revision: If , then it remains unchanged; If , then it is truncated to 0.
[0061] Furthermore, the target heading angle is continuously updated according to the turning rate: .
[0062] Note: The speed and orientation states of a dynamic target evolve continuously over time, thus ensuring good temporal continuity for the vehicle in scenarios such as straight-line driving, turning, lane changing, and U-turns.
[0063] (2) The translational kinetic energy is obtained from the velocity, and the turning energy is obtained from the turning rate. For a dynamic target j of category c, calculate the translational kinetic energy based on the velocity scalar: (14) Then, the steering energy is calculated based on the steering ratio: (15) In equation (15), For category The equivalent moment of inertia parameter, , is an inertial scale used to characterize changes in the orientation of a target.
[0064] The total kinetic energy of the target is defined as: (16) in, and These are non-negative weighting coefficients. Used to adjust the relative contributions of translational and directional motion to the diffusion intensity.
[0065] Static categories can be directly defined as follows: .
[0066] Explanation: Through equations (14) to (16), the translational motion component and the direction change component of the target are uniformly expressed as the total motion energy; and combined with the monotonic mapping relationship defined by equation (17), the diffusion intensity parameter β can be adaptively adjusted with the change of the target's motion state.
[0067] (3) Adjust the diffusion intensity based on the total kinetic energy
[0068] To match lightweight models with fast inference, this invention uses a single [model / mechanism] per instance or per frame. And obtained by the monotonic mapping of total kinetic energy: (17) In equation (25), It is a monotonically increasing mapping function, which makes targets with high total kinematic energy have a greater diffusion intensity, and targets with low motion or static energy have a smaller diffusion intensity.
[0069] In a preferred embodiment, the following may be taken: (18) in, The energy scale constant is .
[0070] To further reduce the dimensionality of conditional inputs and adapt to lightweight models, it is preferable to take the weighted average of the diffusion intensity of dynamic targets across the entire frame: (19) in, This indicates the number of dynamic targets in the k-th frame. Instance point weights or category weights can be used. In static scenarios, It automatically approaches zero.
[0071] Explanation: Through equations (17) to (19), the diffusion intensity is affected not only by the target's translational motion but also by changes in direction, thus achieving diffusion modulation that is more in line with physical principles in scenarios such as turning, lane changing, and U-turns. Due to the mapping function in equation (17) Since the diffusion intensity β is a monotonically increasing function, the higher the total kinetic energy, the greater the diffusion intensity β; conversely, the lower the total kinetic energy, the smaller the diffusion intensity β. A larger β implies stronger random perturbation during the forward diffusion stage, thus allowing for greater adjustment and generation flexibility during the reverse recovery stage, making it suitable for high-dynamic scenarios such as high-speed driving, sharp turns, lane changes, and U-turns. A smaller β implies less disruption to the original point cloud structure during the forward diffusion stage, with the reverse denoising process tending to perform stable recovery near the original structure, making it suitable for low-dynamic or static scenarios. Through this method, the diffusion intensity can adaptively change with the target's kinetic energy, thereby balancing generation flexibility in high-dynamic scenarios with structural stability in low-dynamic scenarios.
[0072] Step 5: Set the diffusion step count S=100, based on the single diffusion obtained in Step 4. The basic noise schedule is modulated, a forward diffusion process for the point cloud is defined for training, and a corresponding reverse denoising generation process is executed during the inference phase to generate the point cloud. Wherein, the... The diffusion intensity is determined by the combined translational kinetic energy and the turning energy of the target.
[0073] Note: In a preferred embodiment, the number of diffusion steps is S=100. This is because if the number of diffusion steps is too small, the reverse denoising process is insufficient, which can easily affect the structural integrity, detail recovery capability, and inter-frame continuity of the generated point cloud. While too many diffusion steps can increase the number of denoising iterations to some extent, they significantly increase inference latency and computational overhead, which is detrimental to the rapid generation of lightweight models and deployment on vehicle or simulation ends. Therefore, this invention preferably sets S=100 to achieve a better balance between point cloud generation quality, temporal stability, and inference efficiency, ensuring that the diffusion generation process can meet the structural recovery requirements of continuous LiDAR point cloud scenes while also considering rapid inference and engineering feasibility.
[0074] (1) β Controlled diffusion noise scheduling Let the basic scheduling be A single energy obtained from physical energy Used to modulate the noise intensity at each step: (20) in, This represents a reference diffusion strength, typically the average value obtained from the statistics of the entire training set. This represents the sensitivity coefficient, and its value range is: . This represents the clipping function. This indicates the upper and lower bounds of the clipping. It is equivalent to multiplying the original spread step scheduling table by a coefficient determined by the target motion energy of the current frame.
[0075] (2) Definition of forward diffusion (for training purposes) set up For the normalized point cloud tensor of the k-th frame (containing Forward diffusion is: (twenty one) Its closed form is: (twenty two) in, The mean of forward diffusion is The variance is .
[0076] (3) Reverse denoising generation (for inference) (twenty three) Where θ represents the denoising network. The mean value of the inverse denoising is... The variance is .
[0077] Example 6 Based on Example 5, step 6 is as follows: Step 6: Construct a lightweight denoising network for the reverse process of the diffusion model to predict the noise term. The process gradually generates a 3D point cloud. The denoising network preferably employs a back-diffusion network structure composed of a nine-layer MLP variant, and the feature updates of each layer are gated and modulated using a diffusion condition vector, thereby achieving stable denoising generation while maintaining a lightweight network. During inference, any number of points N can be specified. k The point cloud is generated by initializing with Gaussian noise and then performing S=100 steps of inverse denoising. The optimal model size is approximately 17.8MB to facilitate rapid deployment on the vehicle or simulation end.
[0078] (1) Network prediction objectives and network structure The denoising network is based on the point cloud state at the current diffusion step. The input is the noise term prediction. This is used for reverse updates. The input to the first layer of the network is the three-dimensional coordinates of the point cloud ((x, y, z) for each point), and a conditional vector C is introduced to conditionally modulate the network. The preferred update form from the 1st layer to the (1+1)th layer is: (twenty four) in, and These represent the network input and output features, respectively, and ⊙ denotes element-wise multiplication. These are trainable parameters. The conditional vector C is preferably constructed as follows: (25) Wherein, c is an optional conditional latent feature: in one embodiment, c can be extracted from partial point clouds, historical frame point clouds, or other prior information by an encoder to achieve conditional generation. In another embodiment, when no external conditional constraints are required, c can be set as a zero vector or a fixed vector to achieve unconditional generation.
[0079] illustrate: For diffusion intensity embedding. To maintain consistency with the "single β energy condition" of this invention, in a preferred embodiment, the single β obtained in step 4 is extended into a triaxially consistent embedding form. This ensures that the network's conditional input is still determined by a single β, while simultaneously through splicing... This enhances the ability to express conditions.
[0080] like Figure 2 As shown, the preferred layer width of the nine-layer MLP is... LeakyReLU is preferred for inter-layer nonlinearity, and BatchNorm is preferred for the first eight layers to stabilize training and improve convergence speed. The network's final output is a point-level 3D prediction, which is used to construct noise term predictions and perform inverse denoising updates.
[0081] (2) Support for any number of points The denoising network employs a point-level shared parameter MLP structure, performing feature updates and noise prediction independently for each point. Therefore, it is insensitive to the order of point arrangement, and the number of output points is consistent with the number of input points. During inference, any number of points can be specified as needed. and obtained from Gaussian noise initialization Subsequently, reverse denoising updates are iteratively performed within S=100 diffusion steps to obtain the generated point cloud. In a preferred embodiment, take The final output point cloud is preferably composed of 8192 points.
[0082] (3) Model size In a preferred embodiment, the denoising network model file is approximately 17.8 MB in size, suitable for rapid deployment on the vehicle / simulation end.
[0083] Step 7: This step uses a continuous multi-frame point cloud sequence as the training target. In the generation stage, it also generates the data frame by frame in sequence, and merges the dynamic point set and the static point set in each frame to output a continuous LiDAR point cloud sequence dataset that meets the needs of autonomous driving simulation.
[0084] (1) Initial frame generation For the initial frame k=0, specify the number of points. The initial noise is sampled using equation (26) and denoised in 100 steps to obtain the result. .
[0085] (2) Continuous generation driven by dynamic state For subsequent frames k≥1, first generate or update the velocity, turning rate, and heading angle according to steps 3-4, then calculate the translational kinetic energy and turning energy and obtain... Then, perform 100 denoising generation steps according to steps 5 and 6 to ensure that the motion activity, direction change trend and diffusion intensity between frames are consistent, and achieve stable generation of continuous sequences.
[0086] (3) Dynamic and static integrated output Each frame's output point cloud is composed of a static point set and a dynamic point set: (26) It can output the final LiDAR point cloud frames according to rules such as sensor field of view and distance clipping, forming a point cloud sequence dataset that can be used for autonomous driving simulation.
Claims
1. A dynamics-based LiDAR point cloud generation method for intelligent driving, characterized in that, A continuous LiDAR point cloud sequence is acquired and normalized to weaken the target motion from a three-dimensional velocity vector to a low-dimensional motion state variable dominated by a velocity scalar. Heading angle and turning rate are introduced to characterize directional changes. A Maxwell-Boltzmann equilibrium distribution of velocity and statistical constraints on turning rate are constructed. Based on Langevin dynamics, the motion state is continuously evolved. Translational kinetic energy is obtained from velocity, and turning energy is obtained from turning rate. These two are jointly mapped to diffusion intensity. ; build The diffusion generation process is controlled, and a lightweight denoising network is used to generate point clouds with any number of points. Finally, a continuous sequence is generated, and the dynamic point set and the static point set are fused to output a continuous point cloud sequence.
2. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 1, characterized in that, The specific steps are as follows: Step 1: Acquire a series of consecutive LiDAR 3D point cloud sequences collected by the intelligent driving vehicle in the road environment, and perform standardization processing; Step 2: Reduce the target motion from a three-dimensional velocity vector to a low-dimensional motion state variable dominated by a velocity scalar, and establish motion intensity and direction change variables; Step 3: Construct the Maxwell–Boltzmann equilibrium distribution of the velocity, establish statistical constraints on the turning rate, and sample the initial motion state variables; Step 4: Based on Langevin dynamics, continuously evolve and update the motion state variables, obtain translational kinetic energy from velocity, obtain turning energy from turning rate, and jointly adjust diffusion intensity; Step 5: Construct a diffusion generation process regulated by diffusion intensity; Step 6: Construct a lightweight denoising network and an arbitrary point generation mechanism; Step 7: Continuous sequence generation and dynamic / static fusion output.
3. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 2, characterized in that, In step 1, the acquired LiDAR continuously acquired 3D point cloud sequence data category c includes multiple categories of original point cloud information such as vehicles, pedestrians, roads, and trees. First, the original point cloud is subjected to coordinate normalization and scale standardization processing, and the continuous multi-frame point cloud is aligned to a unified reference coordinate system according to the vehicle pose. The point cloud is divided into dynamic point sets and static point sets, which are used for subsequent differentiated motion modeling and diffusion generation modeling, respectively. The vehicle pose refers to the vehicle's own position and orientation information at the sampling time, which is used to uniformly transform point clouds collected at different times to the same reference coordinate system.
4. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 2, characterized in that, Step 1.1: Acquire a continuous multi-frame sequence of 3D point cloud data collected by the intelligent driving vehicle in the road environment. Let the i-th frame be... k The frame point cloud is: (1) In the formula, For the first k Frame point cloud 3D coordinates, Points; Step 1.2: Align the point clouds from multiple frames to the same reference coordinate system using the vehicle's pose. Let the first k The frame is the vehicle position pose. T k Then historical frames k - l Point transformation to frame k Coordinate system: (2) , They represent the first k Frame, First k The rigid body transformation matrix corresponding to the vehicle's pose in frame l. Indicates the first k The inverse of the frame pose matrix; This indicates that the historical frames are... k The point in l is transformed to the current point. k The result after the frame coordinate system is restored; Indicates to the first k Points seen in the l-frame are first transformed to a unified world coordinate system or reference coordinate system, and then multiplied by... This means transforming a point that was already in the common coordinate system to the next coordinate system. k In the local coordinate system of the frame; Rigid body transformation matrix , No. k The rotation matrix of a frame represents the change in the vehicle's orientation at that moment. The translation vector of the k-th frame represents the change in the vehicle's position at that moment; Step 1.3: Normalization process: (3) In the formula, For the first k The mean vector of the 3D coordinates of the frame point cloud. ,in Let be the standard deviation vector of the 3D coordinates of the point cloud in the k-th frame; Step 1.4: Delineation of Dynamic and Static Elements (4) in, It is a dynamic point set, including vehicles and pedestrians. It is a static point set, including roads, buildings, and trees.
5. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 4, characterized in that, Step 2 is as follows: With rate The intensity of the target's translational motion is described, while the heading angle and turning rate are used to describe the trend of the target's directional change. When the target center is available for consecutive frames, the velocity observation value of the dynamic target is defined as: (5) in, Indicates the first k The first frame j The central location of a dynamic target Indicates the first k -1 frame j The central location of a dynamic target; Indicates the time interval between adjacent frames; Indicates the first k The first frame j Rate observations of a dynamic target; The heading angle observation value of a dynamic target is defined as: (6) in, Indicates the first j The dynamic objective is from the first k frame l to the k The displacement component of the frame in the y-direction is denoted as ; Indicates the first j The dynamic objective is from the first k frame l to the k The displacement component of the frame in the x-direction is denoted as ; It is the heading angle, indicating the target's direction of travel. k l frame to k Movement direction between frames; using Can be directly based on The sign is used to determine the direction, thus obtaining the angle in the correct quadrant. The output range is: ; The observed turning rate of a dynamic target is defined as: (7) Static targets are directly defined as follows: .
6. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 5, characterized in that, Step 3 is as follows: Step 3.1: Categorize dynamic or static targets c Set equivalent quality With temperature parameters Define rate Maxwell–Boltzmann distribution: (8) In the formula, Boltzmann's constant, For temperature parameters; The initial statistical distribution of the turning rate can be set as a zero-mean Gaussian distribution: (9) in, For category c The turning rate scale parameter is used to characterize the statistical strength of the change in the direction of targets of this category; Step 3.2: Initial Rate Sampling For each dynamic objective The initial motion state is sampled from equations (8) and (9): (10) For static categories, simply let: ; Step 3.3: Express the Maxwell distribution in potential energy form, and define the joint potential energy function corresponding to the velocity and turning rate: (11) In the formula, To prevent small constants from becoming numerically unstable; The velocity balance distribution and turning rate statistical constraints established in step 3 enable the dynamic target to simultaneously possess translational motion continuity and directional change continuity during subsequent continuous generation processes.
7. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 6, characterized in that, In step 4, using the joint potential energy function obtained in step 3 as the target equilibrium state, the motion state variables are discretized and updated frame by frame using overdamped Langevin dynamics, and a non-negativity constraint is applied to the rate, as follows: Step 4.1: Using the joint potential energy function obtained in Step 3 as the target equilibrium state, the motion state variables are discretized and updated frame by frame using overdamped Langevin dynamics to update the dynamic target. The frame-by-frame discretization update form of its rate and turning rate is as follows: (12) (13) in, and These are the diffusion coefficients for velocity and turning rate, respectively. For a long walk; and It is a zero-mean Gaussian random variable; Due to the rate Apply a nonnegative projection to the output of equation (12): The target heading angle is continuously updated according to the turning rate. ; Step 4.2: Obtain the translational kinetic energy from the velocity, and the turning energy from the turning rate. For dynamic targets of category c j Calculate the translational kinetic energy based on the velocity scalar: (14) Then, the steering energy is calculated based on the steering ratio: (15) In the formula, For category The equivalent moment of inertia parameter is used to characterize the inertial scale of the target's orientation change; The total kinetic energy of the target is defined as: (16) in, and These are non-negative weighting coefficients. Used to adjust the relative contributions of translational and directional motion to the diffusion intensity; Static targets are directly defined as follows: ; Step 4.3: Adjust the diffusion intensity based on the total kinetic energy Use a single per target or per frame And obtained by the monotonic mapping of total kinetic energy: (17) In the formula, It is a monotonically increasing mapping function, which makes targets with high total kinetic energy have a greater diffusion intensity, and targets with low motion or static energy have a smaller diffusion intensity; (18) in, The energy scale constant is ; The diffusion intensity of the entire frame is taken as a weighted average of the dynamic target diffusion intensity: (19) in, Indicates the first k Number of dynamic targets per frame Take the target point weight or category weight; in a static scene, It automatically approaches zero.
8. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 6, characterized in that, In step 5, the diffusion step number S is set to the single diffusion obtained in step 4. The basic noise schedule is modulated, a forward diffusion process of the point cloud is defined for training, and a corresponding reverse denoising generation process is executed during the inference phase to generate the point cloud.
9. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 8, characterized in that, Step 5 is detailed below: Step 5.1: β Controlled diffusion noise scheduling Let the basic scheduling be A single energy obtained from physical energy Used to modulate the noise intensity at each step: (20) in, Indicates the reference diffusion intensity. This represents the sensitivity coefficient, and its value range is... ; Represents the clipping function; Indicates the upper and lower boundaries of the clipping; Step 5.2: Let For the first k The frame-normalized point cloud tensor, forward diffused, is as follows: (21) Its closed form is: (22) in, The mean of forward diffusion is The variance is ; Step 5.3: Reverse Denoising (23) Where θ is the denoising network; the mean of the inverse denoising is The variance is .
10. The dynamics-based intelligent driving LiDAR point cloud generation method according to claim 8, characterized in that, Step 6 constructs a lightweight denoising network for the reverse diffusion model to predict noise terms and progressively generate 3D point clouds. The denoising network employs a reverse diffusion network structure composed of a nine-layer MLP variant, and uses diffusion condition vectors to gate and modulate the feature updates of each layer, thereby achieving stable denoising generation while maintaining network lightweightness. An arbitrary number of points can be specified during inference. N k The point cloud is generated by performing S-step reverse denoising after initialization with Gaussian noise.