Safe automatic driving method based on double-alternative diffusion strategy

By adopting a safe autonomous driving method with a dual alternative diffusion strategy in autonomous driving, using diffusion models and integrated networks to train and evaluate actions in offline environments, uncertainty is solved, and the problems of inefficiency and safety risks in traditional reinforcement learning in autonomous driving are achieved, achieving higher safety and task success rates.

CN120057034AActive Publication Date: 2025-05-30CHINA UNIV OF MINING & TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510126707.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

Traditional reinforcement learning has problems of inefficiency and safety risks in autonomous driving, especially insecure decisions caused by extrapolation errors.

Method used

Using a safe autonomous driving method based on a dual alternative diffusion strategy, by building two diffusion models and an integrated network, these networks are trained in offline environments, and using the integrated network to evaluate action uncertainty during the deployment phase, selecting action execution with less uncertainty.

Benefits of technology

It significantly improves the safety of autonomous driving tasks, avoids unsafe decisions caused by extrapolation errors, and achieves a high task success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120057034A_ABST
    Figure CN120057034A_ABST
Patent Text Reader

Abstract

The invention discloses a safe automatic driving method based on a double-alternative diffusion strategy, and provides a strategy using two diffusion models as mutual substitution for the problem that unsafe behaviors may occur in an automatic driving task due to extrapolation errors in traditional offline reinforcement learning. Taking an integrated network comprising a plurality of action value networks as a strategy evaluation network; two diffusion models and an integrated network are trained in an offline environment, the integrated network is utilized to perform uncertainty evaluation on actions generated by two strategies in a deployment stage, and a strategy with low uncertainty is selected as a final driving strategy, so that the safety of an automatic driving task in the deployment stage is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an offline reinforcement learning method, and in particular to a safe automatic driving method. Background Art

[0002] Reinforcement learning can effectively solve sequential decision-making problems through trial and error methods and is widely used in autonomous driving. However, traditional reinforcement learning requires online interaction with the environment, which is inefficient and has safety risks, limiting its development. Recent studies have shown that reinforcement learning can first learn the initial strategy from a static data set and then optimize the strategy performance through limited online interaction. This method is called offline reinforcement learning (ORL). The development of ORL provides strong support for the application of reinforcement learning in autonomous driving.

[0003] Autonomous driving tasks are about safety, and any unsafe behavior may lead to serious consequences. In order to better apply ORL to autonomous driving tasks, the first challenge is to solve the problem of extrapolation error. Extrapolation error occurs because ORL relies on static datasets for training, which are usually collected by specific behavioral policies. Therefore, the state-action distribution in the dataset is restricted by the behavioral policy, resulting in inaccurate or even wrong value estimates of the model in areas not covered by the data. This problem is particularly serious: if ORL is directly deployed to autonomous driving tasks without sufficient safety measures, the strategy may make wrong decisions due to extrapolation error, which may lead to serious safety accidents. Methods to mitigate extrapolation error usually focus on dealing with out-of-distribution (OOD) state-action pairs, which can be mainly divided into the following two categories: 1) Policy constraints: By constraining the learned policy to keep it close to the behavioral policy, the possibility of sampling OOD actions is reduced during the bootstrapping process. 2) Value function penalty: assigning low values ​​to OOD actions to penalize the learned value function. By imposing constraints or penalties, the trained policy is explicitly guided to stay close to the dataset collected by the behavioral policy. However, these methods require accurate evaluation of behavioral policies or precise sampling of OOD samples, both of which are difficult problems. In addition, these methods fail to fully exploit the generalization ability of the action-value (Q-value) function and completely prohibit the intelligent car from exploring any OOD state-action pair without considering its potential value, resulting in an overly conservative strategy. If it is possible to identify those OOD data points whose Q-values ​​can be predicted with high confidence, it may be more effective to allow the intelligent car to select these data points.

[0004] Some studies have shown that when using a Q-value function ensemble network to estimate the posterior distribution of Q-values, in regions with rich data, the Q-values tend to converge, while in regions with sparse data, the Q-values show divergence. This discovery enables us to use the Q-function ensemble network to measure the uncertainty of actions and ensure that the learned policy does not deviate from the behavior policy by selecting actions with lower uncertainty. However, relying solely on uncertainty estimation may face three problems: 1) Insufficient data diversity: To enable the Q-value function ensemble network to accurately capture the true uncertainty, the training data must have sufficient diversity and comprehensiveness; if the training data is limited or has significant biases, the ensemble method may not provide reliable uncertainty estimates. 2) Lack of independence: If the initialization, training data, or optimization process of the Q-value network is highly similar, its outputs may be highly correlated, weakening the effect of the ensemble; in this case, the estimated uncertainty may be inaccurate and unable to reflect the model's true confidence in actions. 3) Ineffectiveness for highly complex distributions: When the true source of uncertainty is complex, such as a multimodal distribution or an environment with high heterogeneity, the numerical outputs of the Q-value function ensemble network may not be able to fully capture these complexities, thus limiting its effectiveness in representing uncertainty.

[0005] On the other hand, given the remarkable success of diffusion models in image generation tasks, recent studies have shown that diffusion models also exhibit high expressive power in learning policies for reinforcement learning tasks, especially suitable for generating highly diverse and robust policy or action distributions. Compared with the traditional Gaussian distribution assumption, diffusion models can more effectively handle multimodal distributions and thus perform particularly well in complex scenarios involving multiple optimal policies. However, reinforcement learning based on diffusion models also faces some challenges, including unstable training, sensitivity to hyperparameters, difficulties in reward guidance, and significant randomness in the sampling process; due to the existence of these challenges, it is difficult for the policies generated by diffusion models to maintain consistency. At the same time, the potential of diffusion models in autonomous driving decision-making tasks has not been fully explored. Its inherent reverse generation process has the characteristics of a black box, resulting in insufficient interpretability. Therefore, the key to successfully applying diffusion models to autonomous driving tasks lies in making full use of their high expressive power while ensuring the safety of the decision-making process. Summary of the Invention

[0006] Object of the Invention: Aiming at the above-mentioned prior art, a safe autonomous driving method based on a dual-alternative diffusion strategy is proposed, which belongs to a safe offline reinforcement learning method that does not require interaction with the environment during training and can ensure the safety of the policy when deployed in an autonomous driving environment.

[0007] Technical Solution: A safe autonomous driving method based on a dual-alternative diffusion strategy. First, two diffusion models are constructed as policy networks, and an ensemble network containing multiple action-value networks is used as a policy evaluation network;

[0008] Secondly, train the two diffusion models and the integrated network in an offline environment to obtain two trained diffusion models and a trained integrated network;

[0009] Then, in the deployment stage, use the trained integrated network to evaluate the uncertainty of the actions generated by the two trained diffusion models;

[0010] Finally, select the action with lower uncertainty to execute.

[0011] Furthermore, a safety autonomous driving method based on a dual-alternative diffusion strategy of the present invention includes the following steps:

[0012] Step 1, construct a policy network and an integrated network, and initialize the network parameters of the policy network and the integrated network;

[0013] The two diffusion models are diffusion model and diffusion model The two diffusion models form a policy network;

[0014] Establish an integrated network including K action value networks as the policy evaluation network;

[0015] Diffusion model Diffusion model and the parameters of the K action value networks are respectively represented by θ 1 , θ 2 and φ k ;

[0016] Step 2, respectively construct the target policy networks of diffusion model diffusion model and the K action value networks and the target integrated network where: θ′ 1 , θ′ 2 and φ′ k respectively represent the network parameters of the target policy network and the target integrated network ;

[0017] The initialization method of the parameters of the target network is: directly assign the parameters (θ 1 , θ 2 , φ k ) of the corresponding original network to the parameters (θ′ 1 , θ′ 2 , φ′ k ) of the target network;

[0018] Step 3, randomly extract samples from the offline dataset and input them into the integrated network, the policy network, and the target network, and train the diffusion model, diffusion model and K action-value networks in the offline environment, and update the parameters of the diffusion model diffusion model integrated network and its target network; obtain two trained diffusion models and a trained integrated network;

[0019] Step 4, in the deployment stage, use the trained integrated network and the two trained diffusion models for the deployment of the intelligent vehicle, use the trained integrated network to evaluate the uncertainty of the actions generated by the two trained diffusion models, and select the actions with lower uncertainty to execute.

[0020] Further, in Step 3, randomly extract samples from the offline dataset and input them into the integrated network, the policy network, and the target network, and train the diffusion model diffusion model and K action-value networks in the offline environment, and update the parameters of the diffusion model diffusion model integrated network and its target network; obtain two trained diffusion models and a trained integrated network. The specific steps are as follows:

[0021] Step 3.1, randomly extract samples (s, a, r, s′) from the offline dataset; where s in the sample (s, a, r, s′) represents the current state of the intelligent vehicle, a represents the action executed by the intelligent vehicle through the policy network, r represents the immediate reward obtained by the intelligent vehicle, and s′ represents the next state of the intelligent vehicle; Input (s, a) in the sample into the integrated network

[0022] diffusion model diffusion model and the diffusion model respectively, input (r, s′) into each action-value network in the target integrated network and input s′ into the target policy network and respectively;

[0023] Step 3.2, update the network parameters of the integrated network First, input s′ into the target policy network

[0024] and and respectively to generate actions

[0025] Then, use the target network of the integrated network to calculate the uncertainty values u′ of the actions 1 , u′ 2 ;

[0026] Compare the uncertainty value u′ of the action 1 with the uncertainty value u′ of the action 2 , and select the target value according to the smaller u′

[0027]

[0028] where [[ID=]] indicates the indicator function;

[0029] Finally, update the parameters of each action value network by minimizing the temporal difference error between the output value of each action value network in the integrated network and the target value joint reward value, that is, minimize the following loss function:

[0030]

[0031] where, represents the number of experience samples in the mini-batch, here When u′ 1 ≤u′ 2 , Otherwise

[0032] Update the parameters φ k of the k-th action value network using the gradient descent method, and the adjustment amount of the parameter φ k is:

[0033]

[0034] where, l φ represents the learning rate in this gradient descent process;

[0035] Step 3.3, update the policy network and network parameters, specifically:

[0036] Minimize the following loss function by gradient descent:

[0037]

[0038] where j = {1, 2}, representing the numbers of 2 independent diffusion models, η and σ are balance coefficients, Represents the policy regularization loss, Represents the policy improvement term, Represents the uncertainty regularization term;

[0039] Using the gradient descent method to update the parameter θ j The adjustment amount of the parameter θ j is:

[0040]

[0041] where l θ represents the learning rate in this gradient descent process, represents taking the derivative of the network parameters in each of the two independent diffusion models;

[0042] Step 3.4, update the target network parameters;

[0043] First, calculate respectively: ρθ 1 +(1 - ρ)θ′ 1 ρθ 2 +(1 - ρ)θ′ 2 and ρφ k +(1 - ρ)φ′ k ;

[0044] Then, assign the results of the above calculations to: θ′ 1 、θ′ 2 and φ′ k ;

[0045] That is, θ′ 1 = ρθ 1 +(1 - ρ)θ′ 1

[0046] θ′ 2 = ρθ 2 +(1 - ρ)θ′ 2

[0047] φ′ k = ρφ k +(1 - ρ)φ′ k

[0048] where ρ represents the target network update rate;

[0049] Step 3.5, repeat Steps 3.1 to 3.4, continuously update each network parameter, and obtain the trained integrated network and the two trained diffusion models.

[0050] Furthermore, the target network of the integrated network respectively calculates the uncertainty values u′ of the actions 1 、u′2 , which is as follows:

[0051]

[0052] Among them, represents the k-th action value function in the target integrated network , and represents the average action value.

[0053] Furthermore, the policy regularization loss in step 3.3 is equivalent to the behavior cloning loss, that is:

[0054]

[0055] where ∈ represents standard Gaussian noise, represents the noise prediction model, represents α i 's cumulative value;

[0056] Then, the policy improvement term is:

[0057]

[0058] Among them, represents the action generated by the diffusion model according to the state s, that is

[0059] Finally, regarding ensuring the safety of the actions generated by the policy, that is, the uncertainty regularization term is:

[0060]

[0061] Furthermore, in step 3.3, the balance coefficients η = 0.01 and σ = 1 in the loss function;

[0062] the learning rate e θ in the gradient descent process

[0063] = 0.003;

[0064] Furthermore, in step 4, select the action with lower uncertainty to execute, which is expressed as:

[0065]

[0066] a selected is the action finally selected and executed by the intelligent vehicle, and s E is the real-time state obtained when the intelligent vehicle interacts with the environment,

[0067] Further, the current state s of the intelligent vehicle includes the steering wheel value, the heading angle, the speed, the distances from both sides of the intelligent vehicle body to both sides of the road, and the distances from the intelligent vehicle body to surrounding obstacles;

[0068] The action a includes the throttle and brake values and the steering wheel value;

[0069] The value range of the throttle and brake values is [-1, 1];

[0070] When the throttle and brake values are in the range [-1, 0), it indicates that the intelligent vehicle is in the braking state, and when the throttle and brake value is equal to -1, it indicates the maximum braking force;

[0071] When the throttle and brake values are in the range [0, 1], it indicates that the intelligent vehicle is in the accelerating state, and when the throttle and brake value is equal to 1, it indicates the maximum acceleration;

[0072] The value range of the steering wheel value is [-1, 1], where [-1, 0) indicates turning the steering wheel to the left, and -1 indicates turning the steering wheel to the left to the maximum; [0, 1] indicates turning the steering wheel to the right, and 1 indicates turning the steering wheel to the right to the maximum.

[0073] Further, both the diffusion model and the action value network are multi-layer perceptron structures with 2 hidden layers and 256 neurons in each hidden layer.

[0074] Beneficial effects: The present invention proposes a safe autonomous driving method based on a dual-alternative diffusion strategy for the problem of safety deployment in autonomous driving. The main advantages of the present invention are: (1) This is a safe offline reinforcement learning method that does not require interaction with the environment during training. (2) This method can ensure the safety of the policy when deployed in the autonomous driving environment. (3) This method uses two diffusion models as alternative strategies, introduces an action value integration network to evaluate the alternative strategies, and selects actions with lower uncertainty, thereby significantly improving the safety of the policy. (4) This method can achieve a high task success rate while ensuring the safety of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 It is a structural diagram of a safe autonomous driving method based on a dual-alternative diffusion strategy. DETAILED DESCRIPTION OF THE INVENTION

[0076] The following further explains the present invention with reference to the accompanying drawings.

[0077] A secure autonomous driving method based on a dual-alternative diffusion strategy. First, two diffusion models are used as alternative strategies, and an integrated network containing multiple action-value networks is used as the policy evaluation network. Second, the two diffusion models and the integrated network are trained in an offline environment. Then, during the deployment phase, the integrated network is used to evaluate the uncertainty of the actions generated by the two diffusion models. Finally, the diffusion strategy with lower uncertainty is selected as the final driving strategy.

[0078] It includes the following specific steps:

[0079] Step 1: Prepare an offline dataset, establish a policy network and an integrated network, and initialize the network parameters.

[0080] Prepare an offline dataset Establish two diffusion models and and represent diffusion model 1 and diffusion model 2 respectively. They are independent of each other and jointly form the policy network; establish an integrated network containing K action-value networks where each action-value network is independent of each other; θ 1 , θ 2 and φ k represent the parameters of diffusion model 1, diffusion model 2, and K action-value networks respectively;

[0081] Policy network Policy network and the action-value network are both multi-layer perceptron structures with two hidden layers and 256 neurons in each hidden layer. Among them: in the policy network, the input dimension of each diffusion model is the state dimension, and the output dimension is the action dimension; in the integrated network, the input dimension of each action-value network is the sum of the state dimension and the action dimension, and the output dimension is 1; the network parameters are initialized randomly.

[0082] Step 2, establish a target network and initialize the network parameters.

[0083] Respectively establish the target networks of diffusion model 1, diffusion model 2, and K action-value networks and where: θ′ 1 , θ′ 2 and φ′ k represent the parameters of the target networks corresponding to diffusion model 1, diffusion model 2, and K action-value networks respectively; the structure of the target network is the same as that of the corresponding original network, and the initialization method of the parameters of the target network is: directly assign the parameters (θ 1 , θ 2 , φ k ) of the corresponding original network to the parameters (θ′1 , θ′ 2 , φ′ k ).

[0084] Step 3: Randomly extract samples from the offline dataset and input them into the integrated network, the policy network, and the target network.

[0085] First, randomly extract samples (s, a, r, s′) from the offline dataset; among them, the sample (s, a, r, s′) represents the current state s of the intelligent vehicle, executes the action a through the policy network, obtains the immediate reward r, and then transitions to the next state s′. In this specific embodiment, the current state s of the intelligent vehicle includes state information such as the steering wheel value, the heading angle, the speed, the distances from both sides of the intelligent vehicle body to both sides of the road, and the distances from the intelligent vehicle body to surrounding obstacles. The action a includes the throttle and brake values and the steering wheel angle.

[0086] Then, input (s, a) in the sample into the integrated network

[0087] and the policy network respectively, input (r, s′) into each action value network in the target integrated network and , and input s′ into the target policy network and respectively.

[0088] Step 4: Update the network parameters of the integrated network .

[0089] By sampling mini-batch samples from the offline dataset and using the gradient descent method to minimize the temporal difference error of each action value network in the integrated network, that is, the square difference between the current Q value and the combined reward value of the target Q value, the parameter update of the action value network is performed. The specific steps are as follows: First, according to the generation mechanism of the diffusion model:

[0090] where a

[0091]

[0092] where a i-1 |a i represents the action generation at a time step in the reverse diffusion process, that is, generating the action a of the previous time step based on the action a of the current time step i , a i-1 represents the action at the current time step in the diffusion process, β i i represents the variance adjustment scheme, controlling the noise level in the diffusion process, and α i i-1 = 1 - β​i , ∈ θ denotes the noise prediction model, which is composed of a neural network with 2 hidden layers. s represents the state input into the diffusion model, denotes the cumulative value of α i where i represents the current time step of the diffusion process, and t represents a time step before the current time step i, denotes the standard Gaussian noise;

[0093] s′ are respectively input into the target policy network and to generate actions

[0094] Then, according to the definition of uncertainty u(s,a):

[0095]

[0096] where Q k (s,a) is the k-th action value function, is the average action value;

[0097] Use the target network of the ensemble network to calculate the uncertainty values of actions respectively:

[0098]

[0099] where, denotes the k-th action value function in the target ensemble network , denotes the average action value to compare the uncertainty value u′ of action with the uncertainty value u′ of action 1 and action and select the target value according to the smaller u′ for calculating the temporal difference error: 2 , and select the target value according to the smaller u′ for calculating the temporal difference error:

[0100]

[0101] denotes the target value finally determined according to the uncertainty value, denotes the indicator function;

[0102] Finally, by minimizing the temporal difference error between the output value of each action value network in the ensemble network and the joint reward value of the target value, update the parameters of each action value network, that is, minimize the following loss function:

[0103]

[0104] where, denotes the number of empirical samples in a small batch, where When u′ 1 ≤u′ 2 at that time, otherwise

[0105] Use gradient descent to update the parameters φ of the k-th action value network k The adjustment amount of the parameter φ k is:

[0106]

[0107] where e φ denotes the learning rate in this gradient descent process, where e φ = 0.0003, denotes taking the derivative of the network parameters in the k-th action value network.

[0108] Step 5, update the policy network parameters.

[0109] The update of the policy network parameters is mainly achieved by optimizing three aspects of losses: policy regularization, policy improvement, and uncertainty regularization, as follows:

[0110] First, the policy regularization loss is equivalent to the behavior cloning loss, that is:

[0111]

[0112] where j = {1, 2}, represents the numbers of 2 independent diffusion models, ∈ represents standard Gaussian noise, represents the noise prediction model, represents the proportion control of the noise;

[0113] Then, regarding improving the quality of the actions generated by the policy, that is, the policy improvement term is:

[0114]

[0115] where represents the action generated by the diffusion model according to the state s, that is

[0116] Finally, regarding ensuring the safety of the actions generated by the policy, that is, the uncertainty regularization term is:

[0117]

[0118] Finally, minimize the following loss function through gradient descent:

[0119]

[0120] Among them, η and σ are balance coefficients, where η = 0.01 and σ = 1;

[0121] Using the gradient descent method to update the parameter θ j The adjustment amount of the parameter θ j is:

[0122]

[0123] where l θ represents the learning rate in this gradient descent process, where e θ = 0.003, represents the derivative of the network parameters of each of the two independent diffusion models.

[0124] Step 6, update the target network parameters.

[0125] First, calculate respectively: ρθ 1 +(1 - ρ)θ′ 1 ρθ 2 +(1 - ρ)θ′ 2 and ρφ k +(1 - ρ)φ′ k ;

[0126] Then, assign the results of the above calculations to: θ′ 1 θ′ 2 and φ′ k ;

[0127] That is, θ′ 1 = ρθ 1 +(1 - ρ)θ′ 1

[0128] θ′ 2 = ρθ 2 +(1 - ρ)θ′ 2

[0129] φ′ k = ρφ k +(1 - ρ)φ′ k

[0130] Among them, ρ represents the target network update rate, where ρ = 0.005.

[0131] Step 7, repeat Steps 3 to 6, continuously update each network parameter, the number of updates is not less than 1×10 6 times, and use the finally updated policy network as the optimal policy for the deployment of the intelligent vehicle.

[0132] Step 8, use uncertainty metrics to select relatively safe actions for execution during algorithm deployment, that is, evaluate the actions given by the two diffusion models through an integrated network, and then select the action with lower uncertainty for execution:

[0133]

[0134] a selected is the action finally selected for execution, s E is the real-time state obtained when the intelligent vehicle interacts with the environment.

[0135] The method of the present invention belongs to a safe offline reinforcement learning method that does not require interaction with the environment during training, and this method can ensure the safety of the policy when deployed in an autonomous driving environment. Specifically, aiming at the problem that traditional offline reinforcement learning may lead to unsafe behaviors in autonomous driving tasks due to extrapolation errors, it is proposed to use two diffusion models as alternative policies, and use an integrated network containing multiple action value networks as the policy evaluation network; by training the two diffusion models and the integrated network in an offline environment, during the deployment phase, the integrated network is used to evaluate the uncertainty of the actions generated by the two policies, and the policy with lower uncertainty is selected as the final driving policy, so as to ensure the safety of the autonomous driving task during the deployment phase.

[0136] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A safe autonomous driving method based on a dual alternative diffusion strategy, characterized in that: Firstly, two diffusion models are constructed as policy networks, and an integrated network containing multiple action-value networks is used as the policy evaluation network. Secondly, the two diffusion models and the integrated network are trained in an offline environment to obtain two trained diffusion models and a trained integrated network; Then, in the deployment phase, the trained integrated network is used to perform uncertainty assessment on the actions generated by the two trained diffusion models; Finally, choose the action with lower uncertainty.

2. According to claim 1, a safe autonomous driving method based on a dual alternative diffusion strategy is characterized in that: The steps include: Step 1: construct the policy network and the integrated network, and initialize the network parameters of the policy network and the integrated network; The two diffusion models are diffusion models and diffusion model The two diffusion models constitute a strategy network; Build an integrated network consisting of K action-value networks As the strategy evaluation network; Diffusion Model Diffusion Model The parameters of the K action-value networks are θ1, θ2 and φ respectively. k express; Step 2: Build diffusion models separately Diffusion Model and a target policy network with K action-value networks and target integrated network Among them: θ′1, θ′2 and φ′ k Represent the target strategy network and target integrated network Network parameters; The method for initializing the parameters of the target network is as follows: the parameters of the corresponding original network (θ1, θ2, φ k ) is directly assigned to the parameters of the target network (θ′1,θ′2,φ′ k ); Step 3: Randomly extract samples from the offline dataset and input them into the integrated network, policy network, and target network to test the diffusion model in an offline environment. Diffusion Model Train with K action value networks and update the diffusion model Diffusion Model Integrated Network and the parameters of its target network; obtain two trained diffusion models and a trained integrated network; Step 4: In the deployment phase, the trained integrated network and the two trained diffusion models are used for deployment of smart cars, and the trained integrated network is used to perform uncertainty assessment on the actions generated by the two trained diffusion models, and the action with lower uncertainty is selected for execution.

3. The method for safe autonomous driving based on a dual alternative diffusion strategy according to claim 1, characterized in that: Step 3: Randomly extract samples from the offline dataset and input them into the integrated network, policy network, and target network to test the diffusion model in an offline environment. Diffusion Model Train with K action value networks and update the diffusion model Diffusion Model Integrated Network and the parameters of its target network; obtain two trained diffusion models and a trained integrated network. The specific steps are as follows: Step 3.1, from offline dataset Randomly extract samples (s, a, r, s′); where s in the sample (s, a, r, s′) represents the current state of the smart car, a represents the action performed by the smart car through the policy network, r represents the immediate reward obtained by the smart car, and s′ represents the state of the smart car at the next moment; Input the samples (s, a) into the integrated network respectively Diffusion Model Diffusion Model In the example, (r, s′) is input into the target integrated network. In each action value network in , s′ is input into the target strategy network and middle; Step 3.2, Update the integrated network Network parameters First, s′ is input into the target policy network and Generate actions in Then, the target network of the integrated network is used Calculate actions separately The uncertainty values ​​u′1, u′2; Comparison Actions The uncertainty value u′1 and the action The uncertainty value u′2 is selected, and the target value is selected according to the smaller u′ represents the indicator function; Finally, by minimizing the time difference error between the output value of each action value network in the integrated network and the target value combined with the reward value, the parameters of each action value network are updated, that is, the following loss function is minimized: in, represents the number of experience samples in a small batch, here When u′1≤u′2, otherwise Use the gradient descent method to adjust the parameter φ of the k-th action value network k Update the parameter φ k The adjustment amount is: Among them, e φ Represents the learning rate during the gradient descent process; Step 3.3, Update the policy network and The network parameters are: The following loss function is minimized by gradient descent: Where j = {1,2}, which indicates the numbers of two independent diffusion models, η and σ are equilibrium coefficients, represents the policy regularization loss, represents the strategy improvement item, represents the uncertainty regularization term; Using the gradient descent method, we can adjust the parameter θ j Update the parameter θ j The adjustment amount is: Among them, l θ represents the learning rate in the gradient descent process, It means to derive the network parameters of two independent diffusion models; Step 3.4, update the target network parameters; First, calculate: ρθ1+(1-ρ)θ′1, ρθ2+(1-ρ)θ′2 and ρφ respectively. k +(1-ρ)φ′ k ; Then, the results of the above calculations are assigned to: θ′1, θ′2 and φ′ k ; That is, θ′1=ρθ1+(1-ρ)θ′1 θ′2=ρθ2+(1-ρ)θ′2 f′ k =rφ k +(1-ρ)φ′ k Where, ρ represents the target network update rate; Step 3.5, repeat steps 3.1 to 3.4, continuously update the network parameters, and obtain the trained integrated network and the two trained diffusion models.

4. According to claim 3, a safe autonomous driving method based on a dual alternative diffusion strategy is characterized in that: The target network using the integrated network Calculate actions separately The uncertainty values ​​u′1 and u′2 are as follows: in, Represents the target integrated network The k-th action value function in represents the average action value.

5. According to claim 3, a safe autonomous driving method based on a dual alternative diffusion strategy is characterized in that: Strategy regularization loss in step 3.3 Equivalent to behavioral cloning loss, that is: Among them, ∈ represents standard Gaussian noise, represents the noise prediction model, Represents α i The cumulative value of Then, the policy improvement term is: in, represents the action generated by the diffusion model according to the state s, that is Finally, to ensure the security of the strategy-generated actions, the uncertainty regularization term is:

6. According to claim 3, a safe automatic driving method based on a dual alternative diffusion strategy is characterized in that: In step 3.3, the balance coefficient in the loss function is η = 0.01, σ = 1; The learning rate l during gradient descent θ =0.003; In step 3.4, the target network update rate ρ = 0.

005.

7. The method for safe autonomous driving based on a dual alternative diffusion strategy according to claim 2, characterized in that: In step 4, the action with lower uncertainty is selected for execution, which is expressed as: a selected is the action that the smart car finally chooses to execute, s E It is the real-time status obtained when the smart car interacts with the environment.

8. The method for safe autonomous driving based on a dual alternative diffusion strategy according to claim 2, characterized in that: The current state s of the smart car includes the steering wheel value, heading angle, speed, the distances from both sides of the smart car to both sides of the road, and the distances from the smart car to surrounding obstacles; Action a includes throttle and brake values ​​and steering wheel values; The throttle and brake values ​​are in the range of [-1,1]; When the throttle and brake values ​​are [-1,0), it means the smart car is in braking state. When the throttle and brake values ​​are equal to -1, it means the braking force is the maximum. When the throttle and brake values ​​are [0,1], it means that the smart car is in the throttle state. When the throttle and brake values ​​are equal to 1, it means that the throttle is at its maximum. The value range of the steering wheel value is [-1,1], wherein [-1,0) indicates turning the steering wheel to the left, and -1 indicates turning the steering wheel all the way to the left; [0,1] indicates turning the steering wheel to the right, and 1 indicates turning the steering wheel all the way to the right.

9. The method for safe autonomous driving based on a dual alternative diffusion strategy according to claim 2, characterized in that: Both the diffusion model and the action value network contain two hidden layers, and the number of neurons in the hidden layer is 256.

Citation Information

Patent Citations

  • Automatic driving reinforcement learning method based on approximate safety action

    CN115542915A

  • Automatic driving hybrid decision control method and system

    CN116661299A

  • Automatic driving model, method and device based on generative diffusion model and vehicle

    CN117519206A

  • Offline reinforcement learning method based on inverse diffusion guidance strategy

    CN117952186A

  • Driver behavior strategy generation method, device and equipment and readable storage medium

    CN118627276A