Automatic driving-oriented kinematics priori guided vehicle trajectory generation method

By constructing a vehicle trajectory generation method guided by kinematic priors, using a dynamic bicycle model and a dual-stream mechanism to obtain the vehicle's physical state, and combining radar and camera data to generate future trajectories that conform to physical characteristics, the problem of poor physical consistency between the vehicle and the trajectory in existing methods is solved, and the prediction and decision-making capabilities of autonomous driving are improved.

CN120686825APending Publication Date: 2025-09-23NINGXIA UNIVERSITY
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510825602.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing autonomous driving trajectory generation methods lack physical feasibility constraints, resulting in poor physical consistency between the vehicle and the trajectory, making it difficult to generate a trajectory that conforms to the vehicle's kinematic characteristics in complex environments.

Method used

A vehicle trajectory generation method guided by kinematic priors is constructed. The vehicle's explicit physical state is obtained through a dynamic bicycle model and a two-stream mechanism. Radar and camera data are combined for feature extraction. An anchored Gaussian distribution and a conditional diffusion model are used to generate noisy trajectory candidate samples. The candidate samples are iteratively trained in a diffusion decoder to guide future trajectory generation.

Benefits of technology

The consistency between vehicle trajectory and physical characteristics is improved, and the generated future trajectory conforms to the actual physical characteristics of the vehicle, enhancing the prediction and decision-making capabilities in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005457972770000085
    Figure BDA0005457972770000085
  • Figure BDA0005457972770000092
    Figure BDA0005457972770000092
  • Figure BDA0005457972770000095
    Figure BDA0005457972770000095
Patent Text Reader

Abstract

The invention provides an automatic driving-oriented kinematics priori guided vehicle trajectory generation method, and belongs to the technical field of trajectory generation and motion control in automatic driving. Comprising the following steps: constructing a vehicle kinematics differential equation, performing nonlinear compensation and control correction based on a double-flow mechanism to obtain a vehicle explicit physical state, and converting the vehicle explicit physical state into a vehicle implicit physical state; feature extraction, target detection and multi-mode fusion are carried out on the collected point cloud data and the vehicle surrounding image data, and environment information of a scene where the vehicle is located is provided; anchoring Gaussian distribution to simulate a feasible noise track of a vehicle in a current scene, generating a noise track candidate sample through sampling and noise adding, performing reverse denoising reasoning on the noise track candidate sample, generating a track anchor point, and generating a reasoning noise track around the track anchor point; environment information of a scene where the vehicle is located, a reasoning noise track and an implicit physical state are input into a diffusion decoder for iterative training, and kinematic prior guides generation of a future track of the vehicle in the denoising process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of trajectory generation and motion control in autonomous driving, and in particular to a vehicle trajectory generation method guided by kinematic priors for autonomous driving. Background Art

[0002] Trajectory generation is the core of autonomous driving decision-making and planning, directly affecting the safety and interactive capabilities of the vehicle. However, existing trajectory generation methods cannot reflect the actual physical motion characteristics of the vehicle while capturing the diverse distribution of trajectories. Based solely on fixed rules or simple feature combinations, the generated trajectories may lack physical rationality in the actual environment. Physical models only focus on the intrinsic properties of the vehicle and ignore the interactions between traffic participants in the environment. Deep learning-based models focus on interactions but ignore the limitations of physical characteristics on the generated results. Currently, trajectory generation methods for autonomous driving are mainly divided into four categories:

[0003] Physics-based methods, such as single trajectory and Kalman filtering, design inputs based on the intended task, impose numerical upper and lower limits, and update vehicle states at discrete time steps. These methods can provide highly interpretable trajectory predictions in simple driving environments. However, they struggle to fully capture the dynamic behavior of complex interactions between vehicles, and their ability to model multi-vehicle interactions or unstructured environments is limited, presenting significant limitations.

[0004] Classic machine learning methods, such as Gaussian processes, dynamic Bayesian networks, hidden Markov models, and support vector machines, improve the ability to handle complex scenarios by learning patterns from data. This alleviates the over-reliance on physical models to some extent. However, mathematical models constructed from hand-crafted features or traditional algorithms are more complex to adjust parameters, and their ability to process time series data and model complex scenarios remains limited.

[0005] Deep learning methods: such as sequential networks, generative models, and graph neural networks. Unlike traditional machine learning models, trajectory prediction based on deep learning uses attention mechanisms and diffusion model architectures to effectively handle long sequence dependencies, solving the problem of limited explicit relationship modeling capabilities. As a "black box", it captures the uncertainty of driving decisions and the complexity of the environment. By combining environmental and interactive information to consider a variety of possible reasonable trajectories in the future, it greatly improves the generalization ability of the model. However, due to the neglect of kinematic mechanisms, the physical motion characteristics of the generated instances are inaccurate, making it difficult to support downstream prediction, decision-making, planning, and other tasks in the physical world;

[0006] Reinforcement learning methods, such as deep inverse reinforcement learning, inverse reinforcement learning, and generative adversarial imitation learning, learn through interaction with the environment, adjusting strategies based on reward signals from the environment to achieve the optimal solution. Using a large amount of interaction data to train the model, visually coherent and geometrically accurate results are achieved. However, the lack of explicit modeling of physical constraints may generate theoretically feasible trajectories that do not conform to kinematic laws and physical reality. This still poses the problem of not being able to accurately characterize the vehicle's kinematic characteristics in the physical world.

[0007] Therefore, due to the lack of physical feasibility constraints and limited modeling capabilities, the predicted vehicle trajectory generated by the current mainstream method is difficult to be highly consistent with the vehicle kinematic theory, and there is a problem of poor physical consistency between the vehicle and the trajectory. Summary of the Invention

[0008] In view of this, the present invention provides a vehicle trajectory generation method guided by kinematic priors for autonomous driving. By constructing a kinematic prior, the model is guided by the kinematic prior when generating future trajectories. The generated future trajectories conform to the actual physical characteristics of the vehicle, solving the problem of poor physical consistency between the vehicle and the trajectory caused by the lack of physical feasibility constraints in existing methods.

[0009] The technical solution adopted by the embodiment of the present invention to solve the technical problem is:

[0010] A vehicle trajectory generation method guided by kinematic priors for autonomous driving, comprising:

[0011] Step S1, constructing the vehicle kinematic differential equation based on the dynamic bicycle model, performing nonlinear compensation and control correction based on the dual-flow mechanism, and obtaining the vehicle explicit physical state;

[0012] Step S2, converting the vehicle explicit physical state into the vehicle implicit physical state;

[0013] Step S3: The perception module performs feature extraction, target detection, and multimodal fusion on the point cloud data collected by the radar and camera and the image data around the vehicle to provide environmental information of the scene in which the vehicle is located;

[0014] Step S4: Anchoring the Gaussian distribution to simulate feasible noise trajectories of the vehicle in the current scene, sampling and denoising the noise trajectories based on the conditional diffusion model to generate candidate noise trajectory samples, performing reverse denoising reasoning on the candidate noise trajectory samples to generate trajectory anchor points, and generating inferred noise trajectories around the trajectory anchor points;

[0015] In step S5, the environmental information of the scene in which the vehicle is located, the inferred noise trajectory, and the implicit physical state are input into a diffusion decoder for iterative training, and the kinematic prior guides the generation of the vehicle's future trajectory during the denoising process.

[0016] Preferably, the step S1 includes:

[0017] Step S11, dividing the physical state variables of the vehicle into attributes, state variables, and control variables, constructing differential equations based on a dynamic bicycle model, and calculating the instantaneous rates of change of the vehicle's position, speed, heading angle, and yaw angle respectively;

[0018] Step S12: constructing a continuous-time differential equation based on the dynamic bicycle model, and using the explicit Euler method to perform full-state discretization update on the instantaneous rate of change of the vehicle position, velocity, heading angle, and yaw angle with a fixed step size Δt, to obtain the kinematic equation for calculating the vehicle's physical state at time t+1;

[0019] Step S13: The forward physical state generation flow of the dual-flow mechanism performs nonlinear compensation on the vehicle physical state at time t+1 to obtain the vehicle explicit physical state at time t+1.

[0020] Step S14, reverse control of the dual flow mechanism modifies the flow to control variable C t The error generated at time step △t is dynamically corrected, and the corrected control variable is used as the initial value of the control variable in the next time step.

[0021] Preferably, the step S2 includes:

[0022] Step S21: The vehicle's explicit physical state is transformed into Convert to The weight matrix is Bias Then the first feature representation unit is stabilized Feature representation;

[0023] Step S22, the implicit state encoder uses a multi-head attention mechanism to Split into Q, K, V vectors, where the V vector carries the physical feature representation The Q, K, V vectors are passed through the implicit state encoder to generate a temporal implicit physical feature sequence {Z t};

[0024] Step S23, aggregate the temporal implicit physical feature sequence {Z t}, forming the implicit physical state of the vehicle

[0025] Step S24, the second fully connected layer linearly expands Convert to The weight matrix is Bias Then the second feature representation unit is stabilized feature representation.

[0026] Preferably, the step S3 includes:

[0027] Step S31: Using radar and cameras to collect raw data around the vehicle, the raw data around the vehicle includes point cloud data and image data around the vehicle. The perception module performs feature extraction and target detection on the raw data around the vehicle to generate information about the intelligent agent around the vehicle.

[0028] In step S32, the perception module performs feature extraction and multimodal fusion on the original data around the vehicle to generate environmental information of the scene where the vehicle is located, including a high-precision map and a bird's-eye view.

[0029] Preferably, the step S4 includes:

[0030] Step S41: Determine the starting point of the distribution simulation based on the vehicle's position, determine the road topology and the lane the vehicle is in based on the environmental information of the scene in which the vehicle is located, analyze the vehicle's historical trajectory data, simulate multiple possible future trajectory distributions of the vehicle in the current scene by anchoring the Gaussian distribution, sample and add noise to the multiple possible future trajectory distributions of the vehicle in the current scene based on the conditional diffusion model, and generate noise trajectory candidate samples.

[0031] Step S42: perform reverse denoising reasoning on the candidate noise trajectory samples based on the conditional diffusion model, extract the historical driving data of the vehicle according to the K-means clustering algorithm, generate trajectory anchor points, and generate inference noise trajectory N around the trajectory anchor points. infer ;

[0032] Step S43: Minimizing the noise prediction error using the loss function to optimize the denoising reasoning capability of the conditional diffusion model.

[0033] Preferably, according to the data output direction, the structure of the diffusion decoder includes a spatial cross attention layer, an agent cross attention layer, a physical state cross attention layer, a feedforward network layer, a time step modulation layer, and a multi-layer perceptron. The diffusion decoder adopts the implicit physical state H phy Acts as a physical prior to guide model trajectory generation;

[0034] The step S5 comprises:

[0035] Step S51: the implicit physical state H phy , the inference noise trajectory N infer , the agent information, the high-precision map and the bird's-eye view are used as inputs of the diffusion decoder, and the inference noise trajectory N inferDefine the sampling space of the diffusion decoder, the spatial cross attention layer makes the inference noise trajectory N infer performing feature interactions with the bird's-eye view;

[0036] Step S52: the agent cross attention layer enables the output of the spatial cross attention layer to perform feature interaction with the agent information and the high-precision map;

[0037] Step S53: the physical state cross attention layer makes the output of the agent cross attention layer and the implicit physical state H phy Perform feature interaction;

[0038] Step S54, the feedforward network layer performs nonlinear abstract processing on the output of the physical state cross attention layer to extract deep features;

[0039] Step S55, the time step modulation layer synchronizes time step information;

[0040] Step S56: the multi-layer perceptron performs nonlinear mapping on the modulated features, and the fully connected layer further extracts trajectory features;

[0041] Step S57, repeating steps S51-S57 to iterate, and using the trajectory features obtained in S56 to define the sampling space of the diffusion decoder during the iterative training process, wherein the implicit physical state H phy As a kinematic prior, it guides the generation of vehicle trajectories during the denoising process of the diffusion decoder iterative training. After the model converges, iterative training stops and the future trajectory of the vehicle is output.

[0042] From the above technical solution, it can be seen that the embodiment of the present invention provides a vehicle trajectory generation method guided by kinematic prior for autonomous driving. First, the vehicle kinematic differential equation is constructed based on a dynamic bicycle model. Nonlinear compensation and control correction are performed based on the dual-flow mechanism to obtain the vehicle's explicit physical state; the vehicle's explicit physical state is converted into the vehicle's implicit physical state; the perception module extracts features, detects targets, and integrates multimodal data from the point cloud data collected by the radar and camera and the vehicle's surrounding image data to provide environmental information of the vehicle's scene; the anchored Gaussian distribution is used to simulate the feasible noise trajectory of the vehicle in the current scene, and the noise trajectory is sampled and denoised based on the conditional diffusion model to generate candidate noise trajectory samples. The candidate noise trajectory samples are subjected to reverse denoising reasoning to generate trajectory anchor points, and inference noise trajectory is generated around the trajectory anchor points; the environmental information of the vehicle's scene, the inference noise trajectory, and the implicit physical state are input into the diffusion decoder for iterative training. The kinematic prior guides the generation of the vehicle's future trajectory during the denoising process. By constructing a kinematic prior, the present invention enables the model to be guided by the kinematic prior when generating future trajectories. The generated future trajectory conforms to the actual physical characteristics of the vehicle, improving the physical consistency between the vehicle and the trajectory. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Attachment Figure 1 This is a flow chart of a vehicle trajectory generation method guided by kinematic priors for autonomous driving according to the present invention;

[0044] Attachment Figure 2 This is the model architecture diagram of the dual-stream mechanism;

[0045] Attachment Figure 3 This is the implicit physical state generation model architecture diagram;

[0046] Attachment Figure 4 This is the diffusion decoder model architecture diagram;

[0047] Attachment Figure 5 It is an example graph of trajectory generation guided by kinematic prior knowledge; DETAILED DESCRIPTION

[0048] The technical solutions and technical effects of the present invention are further described in detail below with reference to the accompanying drawings of the present invention.

[0049] Despite significant progress in trajectory generation for autonomous vehicles, maintaining physical consistency between the vehicle and its trajectory remains a significant challenge. The core of this issue lies in the fact that coupled kinematics are even more complex in complex, multi-interactive driving environments, requiring the trajectory to strictly adhere to vehicle kinematics and physical constraints. However, existing methods often overlook the critical role of vehicle kinematics, and their physical characteristics still make it difficult to support downstream tasks such as prediction, decision-making, and planning. To this end, this paper proposes a vehicle trajectory generation model guided by kinematic priors for autonomous driving scenarios. By constructing a powerful kinematic prior, the model is guided by the kinematic prior when generating future trajectories. The generated future trajectories conform to the actual physical characteristics of the vehicle, improving the physical consistency between the vehicle and its trajectory.

[0050] The present invention provides a vehicle trajectory generation method guided by kinematic priori for autonomous driving. Figure 1 The flow chart shown further illustrates the idea of ​​the present invention:

[0051] This paper uses the nuPlan dataset as an example. nuPlan is a large-scale, real-world benchmark platform for autonomous driving decision-making and planning tasks, covering a variety of urban traffic scenarios, including complex driving tasks such as intersections, roundabouts, and lane changes. We evaluate our model on its official validation set (Val14), test set (Test14), and a high-difficulty test subset (Test14-hard). All experiments were conducted in closed-loop control (non-reactive) and reactive modes. The final score is the average of all test scenarios, ranging from 0 to 100, with higher scores indicating better performance.

[0052] Step S1, constructing a vehicle kinematic differential equation based on a dynamic bicycle model, performing nonlinear compensation and control correction based on a dual-flow mechanism, and obtaining an explicit physical state of the vehicle;

[0053] Step S11, dividing the physical state variables of the vehicle into attributes, state variables, and control variables, constructing differential equations based on a dynamic bicycle model, and calculating the instantaneous rates of change of the vehicle's position, speed, heading angle, and yaw angle respectively;

[0054] Step S12: constructing a continuous-time differential equation based on the dynamic bicycle model, and using the explicit Euler method to perform full-state discretization update on the instantaneous rate of change of the vehicle position, velocity, heading angle, and yaw angle with a fixed step size Δt, to obtain the kinematic equation for calculating the vehicle's physical state at time t+1;

[0055] Step S13: The forward physical state generation flow of the dual-flow mechanism performs nonlinear compensation on the vehicle physical state at time t+1 to obtain the vehicle explicit physical state at time t+1.

[0056] Step S14, reverse control of the dual flow mechanism modifies the flow to control variable C t The error generated at time step △t is dynamically corrected, and the corrected control variable is used as the initial value of the control variable in the next time step.

[0057] In step S11, the attributes are the inherent attributes of the vehicle, including vehicle mass m, moment of inertia l z , the distance from the vehicle's center of mass to the front axle l f and the distance from the center of mass to the rear axle l r , and the steering stiffness k of the front wheels f and the steering stiffness of the rear wheels k r The state variables change with time, namely the x and y coordinates of the vehicle position, the velocity components v of the vehicle in the x and y directions x and v y , heading angle The control variables are the external commands required to drive the model, including the front wheel steering angle δ and the acceleration command α. Based on the dynamic bicycle model, the vehicle kinematic differential equation Y is constructed to calculate the x and y coordinates of the vehicle position and the velocity components v in the x and y directions. x and v y , heading angle and the instantaneous rate of change of the yaw angle ω and The details are as follows:

[0058]

[0059] In step S12, in order to achieve numerical solution in discrete time steps, the explicit Euler method is used to discretize and update the vehicle kinematic differential equations with a fixed time step Δt. and Combined with the vehicle state at the current time t, the kinematic equation of the vehicle's physical state at time t+1 is constructed. The details are as follows:

[0060]

[0061] In step S13, since the kinematic equation of the vehicle physical state at time t+1 can only describe the motion under ideal conditions and cannot cover complex nonlinear factors, the dual-flow mechanism is used to calculate the vehicle physical state X at time t+1. t+1 For nonlinear compensation, the vehicle kinematic differential equation Υ provides the interpretability backbone, and the neural network compensation term ΔXNN is introduced as a correction. A 5-step sliding window is selected to convert the physical state (X) at the current time t and the past 4 time periods into a t-4 ,…,X t-1 ,X) and control variables {Ct-4 ,…,C t-1 ,C t} is considered as an optimizable variable and the neural compensation term ΔXNN is optimized and generated. The vehicle’s explicit physical state at time t+1 The details are as follows:

[0062]

[0063] The loss function of the sliding window is as follows, where k is the prediction step size:

[0064]

[0065] In step S14, the vehicle state at the next moment is actually observed. and the current state X t By inversely solving the vehicle kinematic differential equation Υ, the theoretical control variables are deduced Introducing latent variable z t (including vehicle status, environmental information and historical residuals) to correct it. Will be used as the control variable C for the next time step t+1 The initial value of is used to form a closed-loop feedback loop, ensuring that the control variable can be dynamically adjusted according to the actual observation, avoiding the prediction offset caused by error accumulation. The formula for inverse control correction is as follows:

[0066]

[0067] In formula (5), in low uncertainty scenarios (such as constant speed driving), σ(z t )≈1, which is completely dependent on the result of inverse solution of the vehicle kinematic differential equation Υ; in high uncertainty scenarios (such as emergency braking), σ(z t )≈0, allowing the model to compensate for the deficiencies in the physical equations through neural networks.

[0068] Step S2: converting the vehicle's explicit physical state into the vehicle's implicit physical state.

[0069] Step S21: The vehicle’s physical state is represented by the fully connected layer. Convert to The weight matrix is Bias Then the feature representation is stabilized by the feature representation unit;

[0070] Step S22, the implicit state encoder uses a multi-head attention mechanism to Split into Q, K, V vectors, where the V vector carries the physical feature representation The Q, K, V vectors are passed through the implicit state encoder to generate a temporal implicit physical feature sequence {Zt};

[0071] Step S23, aggregate the temporal implicit physical feature sequence {Z t}, forming the implicit physical state of the vehicle

[0072] Step S24, linearly expand the fully connected layer Convert to Feature representation unit where the weight matrix is Bias The feature representation is then stabilized by the feature representation unit.

[0073] In step S21, the initial input explicit physical state variables As input, linear projection is performed through the fully connected layer, which is essentially a d-dimensional mathematical variable. Mapped to 128-dimensional latent space. The weight matrix of the fully connected layer is set to Bias The formula of the fully connected layer is:

[0074]

[0075] Then, the multi-layer perceptron in the feature representation unit performs nonlinear transformation and normalization to obtain the attention weight matrix. Layer normalization standardizes the feature dimension and stabilizes the vehicle's explicit physical state. feature representation.

[0076] In step S22, in the implicit state encoder, the multi-head attention mechanism is used to Split into Q, K, and V vectors, where the V vector carries the physical characteristics of the vehicle Direct linear mapping is performed to ensure that key physical information is not lost. The Q vector carries the original characteristics of the vehicle. K vector carries the historical characteristics of the vehicle The Q and K vectors are implicitly physically guided and their features are transformed. They are then used as the input of the multi-layer perceptron for nonlinear transformation. After the dot product operation is performed to calculate the correlation between features and to suppress numerical fluctuations, they are linearly mapped with the V vector at the same time. The Q, K, and V vectors are normalized together to obtain the attention weight matrix. Finally, the dot product operation is performed again to calculate the feature relationship between the Q, K, and V vectors to form a temporal implicit physical feature sequence {Z t}.

[0077] In step S23, the output of the implicit state encoder is used as the input of the implicit state decoder, and position encoding is first performed to obtain the temporal implicit physical feature sequence {Z t}Inject temporal position information, then perform layer normalization to standardize its feature dimension, after nonlinear mapping of multi-layer perceptron, finally capture the temporal implicit physical feature sequence through multiple attention heads {Z t}, aggregate them, and finally form a 128-dimensional implicit physical state

[0078] In step S24, the 128-dimensional implicit physical state As input, the 128-dimensional implicit physical state H is linearly projected through the fully connected layer. phy Mapped to 256 dimensions, the fully connected layer weight matrix is ​​set to Bias The formula of the fully connected layer is:

[0079] h 256 =W T H phy +b (7)

[0080] Then, the multi-layer perceptron in the feature representation unit performs nonlinear transformation and normalization to obtain the attention weight matrix. Layer normalization standardizes the feature dimension and stabilizes the implicit physical state of the vehicle. The feature representation forms the final implicit physical state of the vehicle

[0081] In step S3, the perception module performs feature extraction, target detection, and multimodal fusion on the point cloud data collected by the radar and camera and the surrounding image data of the vehicle to provide environmental information of the scene in which the vehicle is located.

[0082] Step S31: Using radar and cameras to collect raw data around the vehicle, the raw data includes point cloud data and image data of the vehicle's surroundings. The perception module performs feature extraction and target detection on the raw data around the vehicle to generate information about the intelligent agent around the vehicle.

[0083] In step S32, the perception module performs feature extraction and multimodal fusion on the original data around the vehicle to generate environmental information of the scene in which the vehicle is located, including high-precision maps and bird's-eye views.

[0084] In step S31, the lidar acquires precise three-dimensional spatial position information, the millimeter-wave radar measures the distance to the target object, and the camera captures images of the vehicle's surroundings from different perspectives. The perception module uses a CNN to extract features and generate point clouds and image features. YOLO-V8 is used to detect objects and identify surrounding vehicles and pedestrians. The cloud and image features are then integrated with the surrounding vehicle and pedestrian information to generate information about the intelligent entities surrounding the vehicle.

[0085] In step S32, the perception module uses CNN to extract features from the point cloud and image multimodal information obtained by the radar and camera, constructs spatial geometric structure features and semantic features, and then fuses their feature information to generate environmental information of the vehicle scene through scene modeling, including high-precision maps and bird's-eye views.

[0086] Step S4: Anchor the Gaussian distribution to simulate the feasible noise trajectory of the vehicle in the current scene, sample and add noise based on the conditional diffusion model to generate candidate noise trajectory samples. Perform reverse denoising reasoning on it, generate trajectory anchor points, and generate reasoning noise trajectory N around the trajectory anchor points. infer .

[0087] Step S41: Determine the starting point of the distribution simulation based on the vehicle's position, determine the road topology and the lane the vehicle is in based on the environmental information of the scene in which the vehicle is located, analyze the vehicle's historical trajectory data, simulate multiple possible future trajectory distributions of the vehicle in the current scene by anchoring the Gaussian distribution, sample and add noise to the multiple possible future trajectory distributions of the vehicle in the current scene based on the conditional diffusion model, and generate noise trajectory candidate samples.

[0088] Step S42: perform reverse denoising reasoning on the noise trajectory candidate samples based on the conditional diffusion model, extract the vehicle's historical driving data according to the K-means clustering algorithm, generate trajectory anchor points, and generate inference noise trajectory N around the trajectory anchor points. infer ;

[0089] Step S43: Minimize the noise prediction error using the loss function to optimize the denoising reasoning capability of the conditional diffusion model.

[0090] In step S41, the global navigation satellite system GNSS obtains the precise vehicle position, determines the starting point of the anchored Gaussian distribution simulation, and uses the environmental information of the scene in which the vehicle is located to determine the road topology and the lane in which the vehicle is located. By integrating the vehicle's position, the information of the intelligent entities surrounding the vehicle, and the environmental information of the scene in which the vehicle is located, an analysis is performed based on the vehicle's historical trajectory data. The multiple future trajectory distributions of the vehicle in the current scene are simulated by anchoring the Gaussian distribution. Based on the conditional diffusion model, the multiple future trajectory distributions of the vehicle in the current scene are sampled and noised to obtain noise trajectory candidate samples. Although the noise trajectory candidate samples have a certain degree of randomness, they fundamentally conform to the physical characteristics of the vehicle and the logic of the autonomous driving scene. Noise trajectory candidate samples The generation formula is as follows:

[0091]

[0092] Among them, i represents the i-th sampling, k represents the k-th noise trajectory, α-i is the time step coefficient, which controls the weight distribution of basic trajectory information and noise at each time step, ψ k It is the basic trajectory information obtained by analyzing the historical trajectory data of the vehicle and extracting its features, ∈ i is from the standard normal distribution The random noise vector obtained by sampling has a standard normal distribution with a mean of 0 and a variance of 1, which ensures the randomness and uniformity of the noise while simulating the possible changes and interference factors in the actual scene.

[0093] In step S42, reverse denoising reasoning is performed on the noise trajectory candidate samples based on the conditional diffusion model, and the historical driving data of the vehicle is extracted according to the K-means clustering algorithm to generate a multimodal trajectory anchor point N for each noise trajectory. anchor Track anchor point N anchor represents the possible driving situations, that is, only focusing on those parts that are most likely to produce useful driving actions, such as going straight, changing lanes, turning, etc. Compared with the randomness of traditional Gaussian noise, the anchored Gaussian distribution reduces the physical inaction in trajectory generation, such as sharp turns with large amplitudes, instant lane changes, etc., around the trajectory anchor point N anchor Generate inference noise trajectory N infer It can effectively limit the initial samples of trajectory generation to the physically feasible range, making the generated inference noise trajectory N infer To a certain extent, it has a reasonable structure and trend, which reduces unnecessary exploration and calculation, and provides a variety of trajectory initial samples with strong physical feasibility for the subsequent steps of trajectory generation. The noise trajectory with high feasibility is selected from the noise trajectory candidate samples to generate a noise trajectory composed of trajectory anchor points. The formula is as follows:

[0094]

[0095] In formula (9), the noise trajectory candidate sample output by formula (8) is As input, is a candidate sample of noise track with track anchor points, z is conditional information, including the position of the vehicle, information about the intelligent agents around the vehicle, and environmental information of the scene where the vehicle is located. Starting from i=1, k=1, the candidate sample of noise track with track anchor points is And conditional information z is input into the conditional diffusion model f θ Perform reverse denoising reasoning, select noise trajectories with high feasibility, and generate noise trajectories composed of trajectory anchor points in, represents the state prediction at the trajectory anchor point, Represents the direction prediction at the trajectory anchor point. The trajectory anchor points in are connected in sequence to form the inference noise trajectory N infer .

[0096] In step S43, in order to evaluate the inference noise trajectory N infer The difference between the real trajectory and the state prediction at the trajectory anchor point and the direction prediction at the trajectory anchor point are close to the actual accurate value. Combining the loss of trajectory and anchor point state, the inference noise trajectory N is adjusted by minimizing the loss function. infer The generation of the loss value is continuously reduced during the training process, making the inference noise trajectory N infer The generation performance of is gradually improved. The formula of the loss function is:

[0097]

[0098] In formula (10) pass reconstruction loss, is the direction prediction at the trajectory anchor point, is the true trajectory direction, y k It is used as a weight coefficient to weight the direction prediction loss of the trajectory anchor point. The binary cross entropy loss of the state prediction at the trajectory anchor point is calculated by BCE. is the state prediction at the trajectory anchor point, y k is the true state of the vehicle, and λ is the hyperparameter of the equilibrium state loss. The prediction and reality are combined, and the difference between the two is calculated to minimize the loss, thereby optimizing the denoising ability of the conditional diffusion model.

[0099] Step S5: The environmental information of the scene where the vehicle is located, the inferred noise trajectory and the implicit physical state are input into the diffusion decoder for iterative training. The inferred noise trajectory N infer The sampling space of the diffusion decoder is defined, and the kinematic prior guides the generation of the vehicle's future trajectory during the denoising process.

[0100] Step S51: The implicit physical state H phy , inference noise trajectory N infer , agent information, high-precision map and bird's-eye view are used as inputs of the diffusion decoder to infer the noisy trajectory N infer Define the sampling space of the diffusion decoder, and the spatial cross attention layer makes the inference noise trajectory N infer Feature interaction with Bird's Eye View;

[0101] Step S52: The agent cross attention layer enables the output of the spatial cross attention layer to interact with the agent information and the high-precision map;

[0102] Step S53: The implicit physical state Hphy As a kinematic prior, the physical state cross attention layer makes the output of the agent cross attention layer consistent with the implicit physical state H phy Perform feature interaction;

[0103] Step S54: the feedforward network layer performs nonlinear abstract processing on the output of the physical state cross attention layer to extract deep features;

[0104] Step S55, the time step modulation layer synchronizes the time step information;

[0105] Step S56: The multi-layer perceptron performs nonlinear mapping on the modulated features, and the fully connected layer further extracts trajectory features;

[0106] Step S57, repeat steps S51-S57 to iterate, and use the trajectory features obtained in S56 to replace the inferred noise trajectory N in step S51 during the iterative training process. infer Redefine the sampling space of the diffusion decoder, where the implicit physical state H phy As a kinematic prior, the vehicle trajectory is generated during the denoising process of the diffusion decoder during N iterative training. By continuously redefining the sampling space, the iterative training is stopped after the model converges, and the future trajectory of the vehicle is finally generated.

[0107] In step S51, the noise trajectory N is inferred infer Define the sampling space of the diffusion decoder so that it samples and adds noise in the defined sampling space. The spatial cross attention layer focuses on the inference noise trajectory N infer The distribution features in space, the features at different spatial locations are given different weights to highlight the information of key spatial locations. infer The bird's-eye view information is introduced at the same time as the feature, and the noise trajectory N is inferred infer The features of the image are interacted with the features of the bird's-eye view based on trajectory coordinates to establish feature association between the two.

[0108] In step S52, the agent cross-attention layer incorporates agent information and the high-precision map, extracting features from each, then interacting with the output of the cross-attention layer. The agent information is used to determine the agent's position based on the high-precision map, allowing the trajectory generation process to be adjusted and optimized based on actual driving conditions.

[0109] In step S53, the physical state cross attention layer introduces the implicit physical state H phy , which is used as a kinematic prior to interact with the output of the cross-attention layer of the agent to guide the trajectory generation process. phy The formula as a physical prior is as follows:

[0110]

[0111] In formula (11), To infer the interactive features of the bird’s-eye view of noisy trajectories, F agent,map is the feature of the agent information and high-precision map, W Q ,W K ,W V is a learnable linear transformation matrix, is a scaling factor used to control the stability of the value and gradient, normalized by Softmax to generate a weight distribution, and then compared with H phy W V Multiply them together to get the trajectory feature F guided by physical prior traj .

[0112] In step S54, the trajectory feature F after interaction traj Information with different scales and levels is fused using the feedforward network layer. Low-level high-resolution, detail-rich features are combined with high-level low-resolution, semantically rich features through top-down lateral connections, enabling the diffusion decoder to simultaneously utilize information at different levels and improve the comprehensive expression ability of trajectory features.

[0113] In step S55, the features fused by the feedforward network layer may have the problem of time step asynchrony. In order to encode the diffusion time step information, the features fused by the feedforward network layer are adjusted through the time modulation layer. The fused features are scaled, offset, and other operations are performed according to the current time step and the internal state of the diffusion decoder. This helps the model to adaptively adjust the feature representation and synchronize the time step in different denoising stages, and better adapt to the needs of the denoising task.

[0114] In step S56, a multi-layer perceptron is used to perform nonlinear transformation on the features modulated by the time modulation layer, and the trajectory features and the offset of the initial noise trajectory coordinates are further extracted and abstracted through the calculation of the fully connected layer.

[0115] In step S57, the first output may still have some noise or the trajectory is not accurate enough, so the output of the fully connected layer is used as the input of the spatial cross attention layer, and steps S51-S56 are repeated for iteration, where the implicit physical state H phy It serves as a kinematic prior to guide the generation of vehicle trajectories during the denoising process of the diffusion decoder during N iterative training until the model converges and finally generates the future trajectory of the vehicle.

[0116] This invention combines the robust interpretability of physics-based approaches with the powerful generalization capabilities of multimodal deep learning methods to fine-tune the kinematics of vehicles in the physical world within existing coarse-grained generative models. This model effectively captures the dynamic motion of vehicles during complex interactions at the physical level, guiding and constraining them based on kinematics to generate future trajectories that are more consistent with physical properties and motion reality.

[0117] The above disclosure is only a preferred embodiment of the present invention, and it is certainly not intended to limit the scope of the present invention. A person skilled in the art can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A vehicle trajectory generation method guided by kinematic priors for autonomous driving, characterized by: include: Step S1, constructing the vehicle kinematic differential equation based on the dynamic bicycle model, performing nonlinear compensation and control correction based on the dual-flow mechanism, and obtaining the vehicle explicit physical state; Step S2, converting the vehicle explicit physical state into the vehicle implicit physical state; Step S3: The perception module performs feature extraction, target detection, and multimodal fusion on the point cloud data collected by the radar and camera and the image data around the vehicle to provide environmental information of the scene in which the vehicle is located; Step S4: Anchoring the Gaussian distribution to simulate feasible noise trajectories of the vehicle in the current scene, sampling and denoising the noise trajectories based on the conditional diffusion model to generate candidate noise trajectory samples, performing reverse denoising reasoning on the candidate noise trajectory samples to generate trajectory anchor points, and generating inferred noise trajectories around the trajectory anchor points; In step S5, the environmental information of the scene in which the vehicle is located, the inferred noise trajectory, and the implicit physical state are input into a diffusion decoder for iterative training, and the kinematic prior guides the generation of the vehicle's future trajectory during the denoising process.

2. The method for generating vehicle trajectories guided by kinematic priors for autonomous driving according to claim 1, characterized in that: The step S1 comprises: Step S11, dividing the physical state variables of the vehicle into attributes, state variables, and control variables, constructing differential equations based on a dynamic bicycle model, and calculating the instantaneous rates of change of the vehicle's position, speed, heading angle, and yaw angle respectively; Step S12: constructing a continuous-time differential equation based on the dynamic bicycle model, and using the explicit Euler method to perform full-state discretization update on the instantaneous rate of change of the vehicle position, velocity, heading angle, and yaw angle with a fixed step size Δt, to obtain the kinematic equation for calculating the vehicle's physical state at time t+1; Step S13: The forward physical state generation flow of the dual-flow mechanism performs nonlinear compensation on the vehicle physical state at time t+1 to obtain the vehicle explicit physical state at time t+1. Step S14, reverse control of the dual flow mechanism modifies the flow to control variable C t The error generated at time step △t is dynamically corrected, and the corrected control variable is used as the initial value of the control variable in the next time step.

3. The method for generating vehicle trajectories guided by kinematic priors for autonomous driving according to claim 2, characterized in that: The step S2 comprises: Step S21: The vehicle's explicit physical state is transformed into Convert to The weight matrix is Bias Then the first feature representation unit is stabilized Feature representation of Step S22, the implicit state encoder uses a multi-head attention mechanism to Split into Q, K, V vectors, where the V vector carries the physical feature representation The Q, K, V vectors are passed through the implicit state encoder to generate a temporal implicit physical feature sequence {Z t }; Step S23, aggregate the temporal implicit physical feature sequence {Z t }, forming the implicit physical state of the vehicle Step S24, the second fully connected layer linearly expands Convert to The weight matrix is Bias Then the second feature representation unit is stabilized feature representation.

4. The method for generating vehicle trajectories guided by kinematic priors for autonomous driving according to claim 3, characterized in that: The step S3 comprises: Step S31: Using radar and cameras to collect raw data around the vehicle, the raw data around the vehicle includes point cloud data and image data around the vehicle. The perception module performs feature extraction and target detection on the raw data around the vehicle to generate information about the intelligent agent around the vehicle. In step S32, the perception module performs feature extraction and multimodal fusion on the original data around the vehicle to generate environmental information of the scene where the vehicle is located, including a high-precision map and a bird's-eye view.

5. The method for generating vehicle trajectories guided by kinematic priors for autonomous driving according to claim 4, characterized in that: The step S4 comprises: Step S41: Determine the starting point of the distribution simulation based on the vehicle's position, determine the road topology and the lane the vehicle is in based on the environmental information of the scene in which the vehicle is located, analyze the vehicle's historical trajectory data, simulate multiple possible future trajectory distributions of the vehicle in the current scene by anchoring the Gaussian distribution, sample and add noise to the multiple possible future trajectory distributions of the vehicle in the current scene based on the conditional diffusion model, and generate noise trajectory candidate samples. Step S42: perform reverse denoising reasoning on the candidate noise trajectory samples based on the conditional diffusion model, extract the historical driving data of the vehicle according to the K-means clustering algorithm, generate trajectory anchor points, and generate inference noise trajectory N around the trajectory anchor points. infer ; Step S43: Minimizing the noise prediction error using the loss function to optimize the denoising reasoning capability of the conditional diffusion model.

6. The method for generating vehicle trajectories guided by kinematic priors for autonomous driving according to claim 5, characterized in that: According to the data output direction, the structure of the diffusion decoder includes a spatial cross attention layer, an agent cross attention layer, a physical state cross attention layer, a feedforward network layer, a time step modulation layer, and a multi-layer perceptron. The diffusion decoder adopts the implicit physical state H phy As a physical prior to guide model trajectory generation; step S5 includes: Step S51: the implicit physical state H phy , the inference noise trajectory N infer , the agent information, the high-precision map and the bird's-eye view are used as inputs of the diffusion decoder, and the inference noise trajectory N infer Define the sampling space of the diffusion decoder, the spatial cross attention layer makes the inference noise trajectory N infer performing feature interactions with the bird's-eye view; Step S52: the agent cross attention layer enables the output of the spatial cross attention layer to perform feature interaction with the agent information and the high-precision map; Step S53: the physical state cross attention layer makes the output of the agent cross attention layer and the implicit physical state H phy Perform feature interaction; Step S54, the feedforward network layer performs nonlinear abstract processing on the output of the physical state cross attention layer to extract deep features; Step S55, the time step modulation layer synchronizes time step information; Step S56: the multi-layer perceptron performs nonlinear mapping on the modulated features, and the fully connected layer further extracts trajectory features; Step S57, repeating steps S51-S57 to iterate, and using the trajectory features obtained in step S56 to define the sampling space of the diffusion decoder during the iterative training process, wherein the implicit physical state H phy As a kinematic prior, it guides the generation of vehicle trajectories during the denoising process of the diffusion decoder iterative training. After the model converges, iterative training stops and the future trajectory of the vehicle is output.

Citation Information

Cited By

  • Dynamic adaptive BEV perception multi-scale feature fusion method

    CN120997790A

  • A dynamic adaptive BEV perception multi-scale feature fusion method

    CN120997790B

  • End-to-end automatic driving system and method based on diffusion model and safety guidance

    CN121157970A

  • Data expansion method and device, computer equipment and storage medium

    CN121353563A

  • Vehicle control method and device, vehicle, storage medium, program product and chip

    CN121375835A