A reinforcement learning guidance based adaptive diffusion trajectory planning method

By adopting an adaptive diffusion trajectory planning method guided by reinforcement learning, the safety and computational efficiency issues of the denoising process in autonomous driving are solved, achieving safe, smooth, and efficient trajectory generation in complex scenarios and enhancing cross-scenario adaptability.

CN121432938BActive Publication Date: 2026-03-24NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202512017065.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-24
Estimated Expiration
2045-12-30

AI Technical Summary

Technical Problem

Existing diffusion models in autonomous driving suffer from problems such as lack of safe guidance in the denoising process, inefficient computation of fixed sampling strategies, and weak generalization ability across scenarios, making it difficult to generate safe, smooth, and efficient trajectories.

Method used

An adaptive diffusion trajectory planning method based on reinforcement learning is adopted. By reconstructing the diffusion denoising process and combining multi-stage training to optimize cross-scene adaptability, a safe and smooth trajectory is generated, the number of sampling steps is dynamically adjusted, and safety and smoothness constraints are integrated.

Benefits of technology

It improves the safety and computational efficiency of autonomous driving trajectories, enhances cross-scenario generalization capabilities, and can generate multimodal trajectories that conform to real driving patterns in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121432938B_ABST
    Figure CN121432938B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning guide's adaptive diffusion trajectory planning method, including according to scene cognitive coding, construct multimodal environment characteristic representation, utilize reinforcement learning guide's diffusion denoising, generate safe smooth trajectory;The denoising step in the iteration process of safe smooth trajectory is realized self-adapting adjustment by step number decision model, according to scene complexity dynamically allocates computing resource, reduces redundant calculation in simple scene, guarantees trajectory accuracy in complex scene, satisfies real-time demand, entire planning method is trained by multi-stage and is combined with offline imitation learning and closed loop reinforcement learning fine tuning, so that model still maintains stable performance in the scene outside training distribution, and the iterative denoising process of diffusion model naturally supports multimodal trajectory generation, combined with the reward guide of reinforcement learning, can generate diversified decision in complex scene in accordance with real driving law.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving, specifically relating to an adaptive diffusion trajectory planning method based on reinforcement learning guidance. Background Technology

[0002] The core performance of an autonomous driving system depends on the decision-making ability of the trajectory planning module in complex dynamic environments. It must simultaneously meet three core requirements: multimodal behavior generation, safety assurance, and real-time response. Traditional trajectory planning methods are mainly divided into two categories: one is rule-based and optimization-based methods, which generate trajectories by pre-setting traffic rules and vehicle dynamics models. Although they perform stably on structured roads (such as highways), their robustness and generalization ability are significantly insufficient when facing unstructured scenarios (such as complex intersections and sudden obstacles). The other is data-driven end-to-end learning methods, which simplify the coupling between modules by directly mapping environmental perception information and vehicle control commands. However, they suffer from problems such as poor interpretability due to their "black box" nature, strong dependence on training data distribution, and weak adaptability to long-tail scenarios.

[0003] In recent years, diffusion models have gained widespread attention in the field of autonomous driving trajectory planning due to their ability to model complex data distributions. Diffusion models generate multimodal trajectories that conform to real-world driving patterns from Gaussian noise through an iterative denoising process, naturally adapting to the "multi-behavior selection" requirements in autonomous driving scenarios. However, existing diffusion models still have three core limitations when applied to safety-critical autonomous driving tasks:

[0004] 1. Lack of safety guidance in the denoising process: The denoising process of traditional diffusion models relies solely on data distribution fitting and does not systematically integrate safety constraints such as collision risk and trajectory smoothness into the iterative process, which may result in generated trajectories that violate physical safety rules;

[0005] 2. Fixed sampling strategy is computationally inefficient: Existing methods use a fixed number of steps for denoising sampling, which results in redundant calculations in simple scenarios (such as straight-line cruising) and insufficient trajectory accuracy in complex scenarios (such as unprotected left turns), making it difficult to meet the real-time requirements of embedded systems.

[0006] 3. Weak cross-scenario generalization ability: Traditional diffusion models rely on offline open-loop training and cannot optimize the model through interaction with dynamic environments. Their performance degrades significantly in scenarios outside the training distribution (such as sudden changes in traffic flow). Summary of the Invention

[0007] To address the aforementioned problems, the present invention aims to provide an adaptive diffusion trajectory planning method based on reinforcement learning. This method reconstructs the diffusion denoising process using a reinforcement learning framework, dynamically adjusts the number of sampling steps, and combines multi-stage training to optimize cross-scenario adaptability, ultimately generating a safe, smooth, and efficient autonomous driving trajectory.

[0008] The specific technical solution for achieving the objective of this invention is as follows:

[0009] An adaptive diffusion trajectory planning method based on reinforcement learning includes the following steps:

[0010] Step 1: Construct a multimodal environmental feature representation based on scene cognition encoding;

[0011] Step 2: Use reinforcement learning-guided diffusion denoising to generate a safe and smooth trajectory.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0013] (1) Significantly improved safety: The solution of this invention integrates safety and smoothness constraints directly into the diffusion denoising process through the GRPO reinforcement learning framework, which effectively reduces the risk of trajectory collision and improves driving safety;

[0014] (2) Optimization of computational efficiency: The adaptive step-counting reasoning mechanism of the present invention dynamically allocates computational resources according to the complexity of the scene, reduces redundant computation in simple scenes, and ensures trajectory accuracy in complex scenes, thus meeting the real-time requirements.

[0015] (3) Enhanced cross-scene generalization ability: Multi-stage training combined with offline imitation learning and closed-loop reinforcement learning fine-tuning enables the model to maintain stable performance in scenes outside the training distribution;

[0016] (4) Multimodal behavior adaptation: The iterative denoising process of the diffusion model naturally supports multimodal trajectory generation. Combined with the reward guidance of reinforcement learning, it can generate diverse decisions that conform to the real driving rules in complex scenarios.

[0017] The present invention will be further described below with reference to specific embodiments. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the adaptive diffusion trajectory planning method based on reinforcement learning guided by the present invention.

[0019] Figure 2 This is a schematic diagram of the iterative process of reinforcement learning-guided denoising in this invention.

[0020] Figure 3 This is a schematic diagram illustrating the step distribution of the adaptive step-counting reasoning of the present invention in different scenarios.

[0021] Figure 4 This is a comparison diagram of trajectory generation effects in complex intersection scenarios according to embodiments of the present invention. Detailed Implementation

[0022] Example

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0025] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps described in these embodiments do not limit the scope of this application. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0026] Combination Figure 1 An adaptive diffusion trajectory planning method based on reinforcement learning includes the following steps:

[0027] Step 1: Construct a multimodal environment feature representation based on scene cognition encoding:

[0028] Step 1-1: Collect multi-source heterogeneous data of autonomous driving scenarios, including dynamic intelligent agent information, static obstacle information and map navigation information;

[0029] The dynamic intelligent agent information includes the position, speed, acceleration, and heading angle of the main vehicle and surrounding vehicles; the static obstacle information includes the geometric position and type of road curbs, guardrails, and stop lines; and the map navigation information includes lane topology, traffic light status, and global waypoints.

[0030] The data collected in this embodiment includes:

[0031] Dynamic agent data: Collect the status of the main vehicle and 32 surrounding agents (one frame every 0.1s, for a total of 21 frames), including x / y coordinates (accuracy ±0.1m), heading angle (sin / cos representation), x / y direction velocity, x / y direction acceleration, vehicle size (width / length), and vehicle category (4-dimensional one-hot encoding).

[0032] Static obstacle data: Collect the geometric position (x / y coordinates) and category (4-dimensional one-hot encoding) of 5 static obstacles;

[0033] Map navigation data: Collected 70 lane polylines (20 points per segment), including lane centerline coordinates, left and right boundary offsets, and traffic light status (green / yellow / red / unknown, 4D one-hot encoding);

[0034] Steps 1-2: Transform the acquired multi-source heterogeneous data into the local coordinate system of the main vehicle, with the current position of the main vehicle as the origin and the vehicle's heading as the positive x-axis. Standardization is used to eliminate the influence of dimensions, and feature encoding is performed separately, including:

[0035] A bidirectional LSTM-based intelligent agent encoder is constructed to process historical trajectory data of dynamic intelligent agents, such as the state sequence within the past n seconds, and outputs dynamic features with dimensions [M×L×D], where M is the maximum number of intelligent agents, L is the number of time frames, and D is the feature dimension. In this embodiment, the hidden layer dimension of the bidirectional LSTM network is 256, the number of layers is 2, and the dropout rate is 0.1.

[0036] A static obstacle encoder is constructed based on MLP-Mixer to process static obstacle information and output static features with dimensions [H×W×D], where H and W are the height and width of the static feature, respectively. In this embodiment, MLP-Mixer contains 4 Mixer blocks (including channel mixing MLP and spatial mixing MLP) and outputs a feature dimension of 256.

[0037] A map navigation encoder based on MLP-Mixer is constructed to process map navigation information and output navigation features with a dimension of [K×D], where K is the number of global waypoints. In this embodiment, MLP-Mixer has the same structure as the static obstacle encoder and outputs features with a dimension of 256.

[0038] Steps 1-3: Employ a multi-head cross-attention mechanism to fuse dynamic features, static features, and navigation features to generate a unified high-dimensional scene feature vector, which serves as the input for subsequent diffusion trajectory denoising.

[0039]

[0040] in, This represents the noisy trajectory feature data, where t is the diffusion time step, ranging from [0,1]. t=1 corresponds to the initial noisy trajectory, and t=0 corresponds to the final denoised trajectory. It is a static feature. As a dynamic feature, For navigation features, This represents a multi-head cross-attention function;

[0041] Noisy trajectory feature data processed by multi-head cross-attention function Updated to noisy trajectory feature data .

[0042] In addition, the initial noisy trajectory feature data is generated by combining Gaussian noise and the current information of the main vehicle in the dynamic intelligent agent information;

[0043] That is, in the subsequent iterative generation of the denoised safe and smooth trajectory in step 2, the trajectory at each step needs to be fused with the scene feature data generated in this step by multi-head cross attention to form a safe trajectory generation that takes into account scene features.

[0044] That is, in the initial iteration of step 2, the generated initial Gaussian noise trajectory is... A multi-head cross-attention mechanism is employed to fuse dynamic, static, and navigation features to obtain initial noisy trajectory feature data. , which serves as the input for the trajectory denoising step.

[0045] Step 2, Combining Figure 2 By utilizing reinforcement learning-guided diffusion denoising, a safe and smooth trajectory is generated.

[0046] Step 2-1: Initialize the diffusion process and obtain the high-dimensional scene feature vector generated in Step 1 as the initial noisy trajectory feature data. ,

[0047] Where t is the diffusion time step, with a value in the range [0,1]. t=1 corresponds to the pure noise trajectory, and t=0 corresponds to the final denoised trajectory. It is discretized into 1000 time steps and a cosine noise scheduling strategy is adopted. Dimensions ,in The time step of the diffusion time step corresponds to the four dimensions of the main vehicle: x-coordinate, y-coordinate, sine of the heading angle, and cosine of the heading angle.

[0048] Step 2-2: Construct a reinforcement learning Markov decision process (MDP), defined as a tuple. ;

[0049] Where S represents the state space, including the scene features extracted in step 1, the current diffusion time step t, and the current noisy trajectory. ;

[0050] A represents the action space, defined as the noise prediction parameters for diffusion denoising (based on Group Relative Policy Optimization, GRPO), used to adjust the direction and magnitude of denoising at each step;

[0051] The reward function r is a multi-objective weighted reward, including trajectory safety costs. With smoothing cost ,in Collision risk is assessed by calculating the minimum distance between the trajectory and surrounding agents and static obstacles. Smoothness is assessed by analyzing changes in trajectory curvature or acceleration curves; both together constitute a reward signal to guide the model in generating a safe and smooth trajectory.

[0052] P represents the state transition probability. This represents the initial state distribution of the task. Indicates the discount factor. Indicates the time span of the decision-making process;

[0053] In this example, the policy network uses a 3-layer MLP, with input dimensions... The hidden layer has a dimension of 512, and the output dimension is the action space dimension.

[0054] Value Network: The structure is the same as the policy network, with output dimension 1 (state value).

[0055] Learning rate: 3×10⁻ 4 The optimizer is AdamW, and the batch size is 512.

[0056] Discount factor The advantage function was calculated using GAE (Generalized Advantage Estimation), with λ=0.95.

[0057] Steps 2-3: Iterative noise reduction optimization:

[0058] Given a denoising step size k, for each step i in the denoising process, i ∈ [0, k], the diffusion time step is: The scene features are fused with the cross-attention features of the current noisy trajectory. Input a noise prediction network built based on Diffusion Transformer (DiT);

[0059] In this embodiment, the noise prediction network adopts a DiT architecture, containing 6 Transformer blocks (each block contains multi-head self-attention, cross-attention, and feedforward networks), with an input dimension of [missing information]. (32 surrounding agents + 1 master vehicle, 21 time frames, 4 state dimensions).

[0060] Based on the GRPO reinforcement learning framework, action A is selected according to the current state S, i.e., the noise prediction parameters, and the noisy trajectory is updated to... ; Calculate the reward value for this step, and update the policy network and value network of GRPO through temporal difference error;

[0061] Enter the denoising process at step i=i+1, repeat the above process until t=0, and output the final denoising trajectory. .

[0062] Furthermore, the denoising step size k is adaptively adjusted in each iteration of denoising optimization, specifically as follows:

[0063] A step-counting decision model based on reinforcement learning is constructed. The state input of the model includes the feature encoding generated in step 1 (including dynamic agent motion patterns, static obstacle attributes, and lane topology information), the current diffusion time step t, and the current trajectory estimate. ;

[0064] The model's action output includes the updated discretized diffusion step count. , The range of values ​​is T is the step size and the number of iterations. These are the minimum and maximum step sizes, respectively; for example... It is discrete into five levels: 5, 8, 12, 16, and 20.

[0065] The reward function of the step-counting decision model is designed as a trade-off function that considers trajectory quality and computational cost, penalizing the long-step diffusion process and guiding the reinforcement learning framework to reduce the number of diffusion steps while ensuring trajectory generation capability.

[0066] The step-count decision model, during the initialization of the diffusion process, outputs the initial step count based on scene feature encoding. ;

[0067] Each completed , After denoising, the scene complexity and current trajectory quality are reassessed, and the number of steps for the next step is adjusted. ;

[0068] Repeat the above process until noise reduction is complete.

[0069] The network structure of the step decision model in this embodiment is a 3-layer MLP based on PPO (input dimension 512+1=513, hidden layer dimension 256, output dimension 5 (corresponding to 5 discrete steps)).

[0070] Learning rate: 1×10⁻ 4 The batch size is 256, and the training epochs are 120.

[0071] Scenario complexity assessment is based on a comprehensive calculation of the number of dynamic agents, lane change frequency, and traffic light state changes. Low-complexity scenarios (such as lane keeping) correspond to 5-8 steps, medium-complexity scenarios (such as lane changing) correspond to 12-16 steps, and high-complexity scenarios (such as unprotected left turns) correspond to 16-20 steps. Figure 3 As shown;

[0072] In the data of this embodiment, by adjusting the number of steps, the calculation time for the lane keeping scenario is reduced to 51ms, the lane changing scenario to 107ms, and the unprotected left turn scenario to 169ms, with the average calculation time reduced by 40% compared to the fixed number of steps strategy.

[0073] In addition, the intelligent agent encoder, static obstacle encoder, map navigation encoder, and noise prediction network are based on the large-scale offline autonomous driving dataset NuPlan, with a resolution of 10. 6 The model is trained using expert trajectories as supervisory signals to obtain initial parameters, enabling it to initially learn the trajectory distribution of real driving.

[0074] The noise prediction network and step-count decision model are trained using a "layered alternating fine-tuning" strategy, including:

[0075] In the CARLA simulation environment, three scenarios—highway, urban road, and complex intersection—are constructed, each containing 1000 dynamic cases, forming a closed-loop training architecture. The main vehicle executes the diffusion-generated trajectory and interacts with the dynamic environment. In the first stage, the noise prediction network is frozen, and the step-counting decision model is trained. In the second stage, the step-counting decision model is frozen, and the GRPO strategy of the noise prediction network is fine-tuned. After each round of fine-tuning, the scene-action-reward sequence is stored in the replay buffer to avoid the model overfitting to specific scenarios.

[0076] Repeat the above training process until the noise prediction network and the step decision model both reach the preset performance indicators in the open loop benchmark (NuPlanVal14 / Test14) and closed loop simulation, ensuring that the model has robust performance in both visible and invisible scenarios.

[0077] like Figure 4 As shown, the noise reduction step size adaptive adjustment strategy in this embodiment is dynamically adjusted at different stages of the vehicle's right turn:

[0078] In the initial stage, a step size of 16 is used to capture trajectory trends. During complex turns, this step size is increased to 20 to improve noise reduction precision. After navigating complex road sections, the step size is reduced back to 16 to balance performance and efficiency. In the right-turn final stage, the step size is reduced to 8 to save computational resources. This strategy effectively improves the fit between the diffusion planner trajectory and the expert trajectory, making the changes in motion parameters smoother. In the subsequent straight-line holding stage, although there are deviations between the trajectory and the expert trajectory, the requirements for smoothness and safety are still met.

[0079] The solution of this invention integrates scene features, adaptive step inference mechanism and GRPO reinforcement learning framework with smooth constraint fusion diffusion denoising process. It combines the multimodal generation capability of diffusion model with the constraint optimization capability of reinforcement learning to achieve high-performance end-to-end autonomous driving trajectory planning, and can generate diversified decisions that conform to real driving rules in complex scenarios.

[0080] This solution also provides an adaptive diffusion trajectory planning system based on reinforcement learning, which includes the following modules:

[0081] Scene encoding module: used to construct multimodal environmental feature representations based on scene cognition encoding;

[0082] Denoising optimization module: Used to generate safe and smooth trajectories by using reinforcement learning-guided diffusion denoising;

[0083] Adaptive step size adjustment module: used to adaptively adjust the denoising step size during the generation of a safe and smooth trajectory;

[0084] Closed-loop training module: used to train and optimize the parameters of the network for scene cognition encoding and trajectory diffusion denoising.

[0085] This solution also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0086] Step 1: Construct a multimodal environmental feature representation based on scene cognition encoding;

[0087] Step 2: Use reinforcement learning-guided diffusion denoising to generate a safe and smooth trajectory.

[0088] This solution also provides a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, performs the following steps:

[0089] Step 1: Construct a multimodal environmental feature representation based on scene cognition encoding;

[0090] Step 2: Use reinforcement learning-guided diffusion denoising to generate a safe and smooth trajectory.

[0091] The embodiments described above are merely one implementation method of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An adaptive diffusion trajectory planning method based on reinforcement learning guidance, characterized in that, Includes the following steps: Step 1: Construct a multimodal environment feature representation based on scene cognition encoding: Step 1-1: Collect multi-source heterogeneous data of autonomous driving scenarios, including dynamic intelligent agent information, static obstacle information and map navigation information; The dynamic intelligent agent information includes the position, speed, acceleration, and heading angle of the main vehicle and surrounding vehicles; the static obstacle information includes the geometric position and type of road curbs, guardrails, and stop lines; and the map navigation information includes lane topology, traffic light status, and global waypoints. Steps 1-2: Transform the acquired multi-source heterogeneous data into the main vehicle's local coordinate system and perform feature encoding on each, including: A bidirectional LSTM-based agent encoder is constructed to process dynamic agent information and output dynamic features with dimensions [M×L×D], where M is the maximum number of agents, L is the number of time frames, and D is the feature dimension. A static obstacle encoder is built based on MLP-Mixer to process static obstacle information and output static features with dimensions [H×W×D], where H and W are the height and width of the static feature, respectively. A map navigation encoder is built based on MLP-Mixer to process map navigation information and output navigation features with dimension [K×D], where K is the number of global waypoints; Steps 1-3: Employ a multi-head cross-attention mechanism to fuse dynamic features, static features, and navigation features to generate a unified high-dimensional scene feature vector, which serves as the input for subsequent diffusion trajectory denoising. ; in, This represents the noisy trajectory feature data, where t is the diffusion time step, ranging from [0,1]. t=1 corresponds to the initial noisy trajectory, and t=0 corresponds to the final denoised trajectory. It is a static feature. As a dynamic feature, For navigation features, This represents a multi-head cross-attention function; Noisy trajectory feature data processed by multi-head cross-attention function Updated to noisy trajectory feature data ; Step 2: Utilize reinforcement learning-guided diffusion denoising to generate a safe and smooth trajectory: Step 2-1: Initialize the diffusion process and obtain the high-dimensional scene feature vector generated in Step 1 as the initial noisy trajectory feature data. , Where t is the diffusion time step, with a value in the range [0,1]. t=1 corresponds to pure noise, and t=0 corresponds to the final denoised trajectory. Dimensions ,in The time step of the diffusion time step corresponds to the four dimensions of the main vehicle: x-coordinate, y-coordinate, sine of the heading angle, and cosine of the heading angle. Step 2-2: Construct a reinforcement learning Markov decision process (MDP), defined as a tuple. ; Where S represents the state space, including the scene features extracted in step 1, the current diffusion time step t, and the current noisy trajectory. ; A represents the action space, defined as the noise prediction parameter for diffusion denoising, used to adjust the direction and amplitude of denoising at each step; The reward function r is a multi-objective weighted reward, including trajectory safety costs. With smoothing cost ,in Collision risk is assessed by calculating the minimum distance between the trajectory and surrounding agents and static obstacles. Smoothness is assessed by analyzing changes in trajectory curvature or acceleration curves; both together constitute a reward signal to guide the model in generating a safe and smooth trajectory. P represents the state transition probability. This represents the initial state distribution of the task. Indicates the discount factor. Indicates the time span of the decision-making process; Steps 2-3: Iterative noise reduction optimization: Given a denoising step size k, for each step i in the denoising process, i ∈ [0, k], the diffusion time step is: The scene features are fused with the cross-attention features of the current noisy trajectory. Input a noise prediction network built based on DiffusionTransformer; Based on the GRPO reinforcement learning framework, action A is selected according to the current state S, and the noisy trajectory is updated to... ; Calculate the reward value for this step, and update the policy network and value network of GRPO through temporal difference error; Enter the denoising process at step i=i+1, repeat the above process until t=0, and output the final denoising trajectory. .

2. The adaptive diffusion trajectory planning method based on reinforcement learning as described in claim 1, characterized in that, The initial noisy trajectory feature data is generated by combining Gaussian noise and the current information of the main vehicle in the dynamic intelligent agent information.

3. The adaptive diffusion trajectory planning method based on reinforcement learning as described in claim 1, characterized in that, The denoising step size k is adaptively adjusted in each iteration of denoising optimization, specifically as follows: Construct a step-counting decision model based on reinforcement learning. The state input of the model includes the feature encoding generated in step 1, the current diffusion time step t, and the current trajectory estimate. ; The model's action output includes the updated discretized diffusion step count. The range of values ​​is M is the step size and the number of iterations. These are the minimum and maximum step sizes set, respectively. The reward function of the step-counting decision model is designed as a trade-off function that considers trajectory quality and computational cost, penalizing the long-step diffusion process and guiding the reinforcement learning framework to reduce the number of diffusion steps while ensuring trajectory generation capability. The step-count decision model, during the initialization of the diffusion process, outputs the initial step count based on scene feature encoding. ; Each completed , After denoising, the scene complexity and current trajectory quality are reassessed, and the number of steps for the next step is adjusted. ; Repeat the above process until noise reduction is complete.

4. The adaptive diffusion trajectory planning method based on reinforcement learning as described in claim 1, characterized in that, The intelligent agent encoder, static obstacle encoder, map navigation encoder, and noise prediction network are trained based on the large-scale offline autonomous driving dataset NuPlan, using expert trajectories as supervision signals to obtain initial parameters, enabling the model to initially learn the trajectory distribution of real driving. The noise prediction network and step-counting decision model are trained using a "layered alternating fine-tuning" strategy, including: A closed-loop training architecture is constructed in a simulation environment. The main vehicle executes the diffusion-generated trajectory and interacts with the dynamic environment. In the first stage, the noise prediction network is frozen and the step decision model is trained. In the second stage, the step decision model is frozen and the GRPO strategy of the noise prediction network is fine-tuned. After each round of fine-tuning, the scene-action-reward sequence is stored in the replay buffer to avoid the model from overfitting to a specific scene. Repeat the above training process until the noise prediction network and the step decision model both reach the preset performance indicators in open-loop benchmarks and closed-loop simulations, ensuring that the model has robust performance in both visible and invisible scenarios.

5. An adaptive diffusion trajectory planning system based on reinforcement learning, used to execute the method of claim 1, characterized in that, Includes the following modules: Scene encoding module: used to construct multimodal environmental feature representations based on scene cognition encoding; Denoising optimization module: Used to generate safe and smooth trajectories by using reinforcement learning-guided diffusion denoising; Adaptive step size adjustment module: used to adaptively adjust the denoising step size during the generation of a safe and smooth trajectory; Closed-loop training module: used to train and optimize the parameters of the network for scene cognition encoding and trajectory diffusion denoising.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-4.

7. A computer-storable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Automatic driving control method, device, system and equipment and storage medium

    CN118393973A

  • Key safety scene generation system for automatic driving automobile based on diffusion model

    CN120107972A