Automatic driving robustness confrontation training method based on adaptive background vehicle model

By constructing an adaptive background vehicle model, using a hybrid risk field and an adaptive alternating training mechanism, the problems of insufficient versatility and poor stability in autonomous driving confrontation training are solved, and robustness and safety improvements in complex traffic environments are achieved.

CN120354744AActive Publication Date: 2025-07-22HARBIN INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510522601.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing autonomous driving confrontation training methods have problems such as insufficient generality, inaccurate risk quantification and poor training stability. It is difficult to effectively deal with emergencies in complex traffic scenarios, affecting vehicle safety.

Method used

Build an adaptive background vehicle model, and dynamically adjust the adversarial intensity through a mixed risk field, a probabilistic risk assessment framework and an adaptive alternating training mechanism to achieve robustness and safety of adversarial training.

Benefits of technology

It improves the robustness of autonomous driving systems in complex traffic environments, reduces accident rates, provides a general adversarial training framework, and enhances learning efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354744A_ABST
    Figure CN120354744A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving robustness confrontation training method based on an adaptive background vehicle model, and belongs to the technical field of intelligent traffic control. In order to solve the problem of insufficient universality in the existing automatic driving confrontation training, the method comprises the following steps: constructing an automatic driving model robustness training environment; a combined attenuation factor is constructed to analyze the risk between the background vehicle and the target vehicle, and TTC is adopted to construct a mixed risk field of the background vehicle; constructing a probability risk assessment framework; a scoring function is introduced to calculate the expected risk of the background vehicle and the expected risk of the surrounding environment, and an activation function is constructed to screen the background vehicles and behaviors confronting the target vehicle; a self-adaptive alternate training mechanism is constructed, the background vehicle model determines background vehicle behaviors according to an activation function, the target vehicle model makes a decision based on a deep reinforcement learning strategy, and the self-adaptive alternate training mechanism dynamically adjusts the intensity of adversarial behaviors in the simulation training process, so that automatic driving robustness adversarial training is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent traffic control and management, and specifically provides an autonomous driving robust adversarial training method based on an adaptive background vehicle model. Background Art

[0002] In recent years, autonomous driving technology has developed rapidly and has become an important research direction in the field of intelligent transportation. With the gradual commercialization of autonomous vehicles, their safety and reliability in actual traffic environments have received widespread attention. However, current autonomous vehicles still face many safety challenges in mixed traffic flows with human-driven vehicles.

[0003] The behavior patterns of manually driven vehicles are highly unpredictable. Drivers may make sudden lane changes, emergency brakes, and other operations based on various subjective factors during driving. These behaviors are often difficult to be accurately predicted by the autonomous driving system in advance, which greatly increases the risk of traffic accidents. The autonomous driving system itself also has some problems that need to be solved. Existing autonomous driving models are mainly divided into two categories: traditional modular and end-to-end models. Although the traditional modular model has a clearer functional division, in the face of some extreme scenarios, there will be problems with the coordination between modules, resulting in a decrease in the overall performance of the system; the end-to-end model is insufficient in the generalization ability of complex scenarios, and it is difficult to effectively deal with adversarial interference. Once encountering malicious attacks or complex interference signals, the system may misjudge or even fail.

[0004] In the training phase of autonomous driving systems, the currently commonly used training methods have obvious limitations. Most training environments are built based on static or regularized scenarios, lacking the dynamic adversarial elements of real traffic environments. This single training environment means that after training, the autonomous driving model can perform well in ideal scenarios, but when faced with real, complex and changeable traffic scenarios, it often has problems of insufficient adaptability and cannot effectively respond to various emergencies, which in turn affects the safe driving of the vehicle.

[0005] To address the above problems, some studies have begun to attempt to improve the robustness of autonomous driving systems through adversarial training. However, these adversarial training methods have also revealed some deficiencies in practical applications. On the one hand, most of these methods rely on specific vehicle models and lack sufficient generality, making it difficult to directly apply them to different autonomous driving systems and limiting the possibility of their wide application. On the other hand, during the training process, the adversarial intensity is usually fixed, which makes the system prone to collapse or difficult convergence during training, seriously affecting the stability and efficiency of training. In addition, existing adversarial training methods lack a set of dynamic risk assessment mechanisms and cannot accurately quantify the risks brought by adversarial behaviors. It is also difficult to effectively ensure the safety of the system while enhancing the adversarial challenge, making the entire adversarial training process have a certain degree of blindness and risk. Summary of the Invention

[0006] The problem to be solved by the present invention is the problems of insufficient generality, inaccurate risk quantification, and poor training stability in existing autonomous driving adversarial training. A method for robust adversarial training of autonomous driving based on an adaptive background vehicle model is proposed.

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] A method for robust adversarial training of autonomous driving based on an adaptive background vehicle model includes the following steps:

[0009] S1. Define the autonomous driving vehicle under training as the target vehicle model, define the vehicle models around the target vehicle model as the background vehicle models, construct a robust training environment for the autonomous driving model, and collect the driving parameters of the target vehicle model and the background vehicle models.

[0010] S2. Construct a combined attenuation factor to analyze the risks between the background vehicles and the target vehicle, and use the time to collision (TTC) before collision to construct a mixed risk field of the background vehicles for reverse risk assessment of the safety of the target vehicle.

[0011] S3. Construct a probability risk assessment framework, hierarchically divide the risk levels into three levels: high, medium, and low, and calculate the conditional probability, posterior probability, and surrounding environment probability between the background vehicles and the target vehicle at different risk levels based on the mixed risk field obtained in step S2.

[0012] S4. Introduce a scoring function to map the risk levels to numerical scores, calculate the expected risks of the background vehicles and the surrounding environment, and construct an activation function to screen the background vehicles and behaviors that generate adversarial effects on the target vehicle.

[0013] S5. Build an adaptive alternating training mechanism. Set the simulation training process to be carried out for M rounds, and each round of training executes T time steps. In each simulation time step, the background vehicle model determines the behavior of the background vehicle according to the activation function in step S4, and the target vehicle model makes decisions based on the deep reinforcement learning strategy. During the simulation training process, the adaptive alternating training mechanism dynamically adjusts the intensity of the adversarial behavior according to the safety performance of the target vehicle model, and completes the robust adversarial training of autonomous driving based on the adaptive background vehicle model.

[0014] Further, the specific implementation method of step S1 includes the following steps:

[0015] S1.1. Define the trained autonomous vehicle as the target vehicle, and the target vehicle model as π α ; Define the vehicles around the target vehicle as background vehicles, and the background vehicle model as π β ;

[0016] Set the target vehicle and the background vehicle as independent and competing entities. The target vehicle and the background vehicle lack white-box access to each other's input states without communication, and do not share parameter information in the weight configuration;

[0017] Define the interaction process between the target vehicle model and the background vehicle model as a Markov process, and the expression is:

[0018] M = <S, (O α , O β ), (A α , A β ), P, (R α , R β ), γ> (1)

[0019] Where S represents the state space, O α and O β are the state spaces of the target vehicle model and the background vehicle model respectively, A α and A β are the action spaces of the target vehicle model and the background vehicle model respectively, P is the transition probability function, R α and R β are the reward functions of the target vehicle model and the background vehicle model respectively, and γ is the discount factor;

[0020] P: S × A α × A β → S′

[0021] Where S′ represents the probability distribution of the next state;

[0022] With the goal of maximizing the cumulative reward R: S × A α × A β×S′→R;

[0023] S1.2. Define that the system uses V2X communication technology to collect the driving parameters of the target vehicle and six specific-position background vehicles around it in real time. The background vehicle set Ω b is defined as:

[0024] b∈Ω b ={front, left front, right front, rear, left rear, right rear} (2)

[0025] where b is the background vehicle;

[0026] Define that the state space of the background vehicle includes the target vehicle state, the background vehicle state, and lane features. The target vehicle state includes the lateral position x t , speed v t , acceleration a t and heading angle θ t , the background vehicle state includes the relative distance Δd b between the nearby background vehicle and the target vehicle, the speed v b and acceleration a b of the background vehicle, and the lane features include the average speed and density ρ lane of each lane. The data collection range covers 300 meters around the target vehicle;

[0027] Define that the action space of the background vehicle includes candidate background vehicles candidate adversarial behaviors aa, and candidate adversarial behaviors include lane change, hard acceleration, hard deceleration, and normal behavior na.

[0028] Furthermore, the specific implementation method of step S2 includes the following steps:

[0029] S2.1. Construct the longitudinal attenuation factor φ x,b (Δx,Δy,t), which quantifies the risk attenuation along the longitudinal axis of the vehicle. The expression is:

[0030]

[0031] where Δx represents the longitudinal relative position, Δy is the lateral relative position, x is the longitudinal position coordinate, L b is the length of the background vehicle, v x,b (t) represents the longitudinal speed of the vehicle at time t, and the parameters μ x,b and ρ x,b represent the longitudinal risk attenuation intensity based on vehicle type and the longitudinal risk attenuation intensity based on speed influence respectively, and max is the maximum value function;

[0032] S2.2. Construct the lateral attenuation factor φy,b (Δx, Δy, t), which quantifies the risk decay along the lateral axis of the vehicle, and the expression is:

[0033]

[0034] where y is the lateral position coordinate, W b is the width of the background vehicle, v y,b (t) is the lateral speed of the vehicle at time t, and the parameters μ y,b and ρ y,b respectively represent the lateral risk decay intensity based on the vehicle type and the lateral risk decay intensity based on the speed influence;

[0035] S2.3. Combine the longitudinal decay factor and the lateral decay factor using the Euclidean norm to obtain the combined decay factor φ k (Δx, Δy, t), and the expression is:

[0036]

[0037] S2.4. Introduce the TTC collision time into the artificial risk domain. The calculation formulas of TTC on different axes are as follows:

[0038]

[0039] where TTC x is the longitudinal collision time, TTC y is the lateral collision time, Δv x , Δv y are respectively the longitudinal relative speed and the lateral relative speed between the background vehicle and the target vehicle;

[0040] Construct the hybrid risk field of the background vehicle, and the expression is:

[0041]

[0042] where η b represents the risk value of the hybrid risk field from the background vehicle b to the target vehicle, indicating the risk value at the position (Δx, Δy) around the background vehicle b at time t with the relative speed (Δv x , Δv y ).

[0043] Furthermore, the specific implementation method of step S3 includes the following steps:

[0044] S3.1. Divide the risk level into three risk levels: high, medium, and low hierarchically, and the risk level is defined as:

[0045] τ ∈ Ω = {dangerous, attentive, safe} = {D, A, S} (9)

[0046] Among them, τ represents the risk level, Ω represents the set of risk levels, D represents high risk, A represents medium risk, and S represents low risk;

[0047] S3.2. Define the conditional probability distribution function F(η b | τ b ) of the background vehicle b under different risk levels, and its definition is as follows:

[0048]

[0049]

[0050] Among them, F(η b | τ b = D) is the conditional probability distribution function under the high-risk level, f(η b | τ b = A) is the conditional probability distribution function under the medium-risk level, F(η b | τ b = S) is the conditional probability distribution function under the low-risk level, η D , η S , σ S , σ D are respectively the predefined high-risk threshold hyperparameter, low-risk threshold hyperparameter, low-risk standard deviation hyperparameter, and high-risk standard deviation hyperparameter;

[0051] S3.3. Use Bayesian inference to calculate the posterior probability P(τ b | η b ) of the risk level τ b of the background vehicle b given η b , and the expression is:

[0052]

[0053] Among them, P(τ b ) represents the prior probability of the risk level τ b ;

[0054] For simplicity of calculation, assume that the prior probability distributions of different risk levels are uniform and satisfy the normalization condition, and the expression is:

[0055]

[0056] S3.4. Comprehensively consider the risk levels of all background vehicles in the surrounding environment to determine the overall risk distribution of the entire environment. The probability distribution of the risk level of the surrounding environment is defined as:

[0057] P(τ env =D)=1 - ∏ b∈B (1 - P(τ b =D|η b )) (12)

[0058] P(τ env =S)=∏ b∈B P(τ b =S|η b ) (13)

[0059] P(τ env =A)=1 - P(τ env =S) - P(τ env =D) (14)

[0060] Among them, P(τ env =D), P(τ env =A), and P(τ env =S) respectively represent the probabilities of the surrounding environment being in three risk levels of high, medium, and low, and are represented by P(τ b =D|η b ), P(τ b =A|η b ), P(τ b =S|η b ) to represent the risk level probabilities of the background vehicle b.

[0061] Furthermore, the specific implementation method of step S4 includes the following steps:

[0062] S4.1. Introduce a scoring function Sc(τ) to map the risk level to a numerical score, assign 2, 1, and 0 to D, A, and S respectively, and calculate the expected risk of the background vehicle and the expected risk of the surrounding environment. The expressions are:

[0063]

[0064]

[0065] Among them, ∈ b is the expected risk of the background vehicle, ∈ env is the expected risk of the surrounding environment, P(τ b |n b ) is the conditional probability of the risk level τ b under the given state n b , Sc(τ b ) is the scoring value of the risk level τ b , Sc(τ env ) is the scoring value of the risk level τ env ;

[0066] S4.2. According to the expected risk of the candidate background vehicle The expected risk relative to the surrounding environment ∈ env Evaluate whether the candidate background vehicle is suitable for performing adversarial actions, and define the activation function as:

[0067]

[0068] where k is a positive constant that determines the steepness of the function, M l and M u are the upper and lower bounds of the activation function respectively. If the result of the activation function is 1, it is determined that the candidate background vehicle performs the candidate adversarial behavior, and if it is 0, it performs the normal behavior.

[0069] Furthermore, the specific implementation method of step S5 includes the following steps:

[0070] S5.1. Construct an adaptive alternating training mechanism, and adaptively switch the training process between the target vehicle and the background vehicle according to the switching indicator C R using as an indicator, where C c represents the number of training rounds between two accidents;

[0071] The switching is controlled by the adaptive alternating training threshold H. When C R > H, the training of the background vehicle model is turned off, and the training of the target vehicle model is started to enhance its decision-making strategy under adversarial conditions;

[0072] S5.2. Set the goal of the background vehicle model in the training phase to utilize the fixed strategy of the target vehicle model during the non-training period of the target vehicle, and optimize the strategy of the background vehicle model π β to maximize the cumulative discounted reward, and the expression is:

[0073]

[0074] where, is the optimal strategy of the background vehicle model β, that is, the strategy that maximizes the long-term discounted reward; is the expectation under the strategy of the background vehicle model π β ; R β is the reward function of the background vehicle model β, indicating the immediate reward obtained when transferring to the state s (t) after executing the action and transferring to the state s (t+1) ; γ ∈ [0, 1) is the discount factor, which is used to weigh the importance of future rewards; s (t) and s (t+1)The states at times t and t+1 respectively, and is the action executed by the background vehicle β at time t; is the background vehicle model π α under the influence of the target vehicle model π β state transition probability function, indicating the probability of transitioning to s (t) after executing the action ; s0 = s represents the initial state as s; (t+1) The probability of;

[0075] S5.3. When the background vehicle model is trained, by introducing adversarial behavior into the deep reinforcement learning policy π α of the target vehicle model for decision-making, when the adversarial behavior effectively creates failure cases for the target vehicle, update the weight parameters of the target vehicle model by retraining, while keeping the parameters of the background vehicle model fixed in the training environment. The policy π α of the target vehicle model is to maximize the cumulative discounted reward, and the expression is

[0076]

[0077] where, is the optimal policy of the target vehicle model α, that is, the policy that maximizes the long-term discounted reward; is the expectation under the policy of the target vehicle model π α ; R α is the reward function of the target vehicle model α, indicating the immediate reward obtained when transitioning to the state s (t) after executing the action ; γ ∈ [0, 1) is the discount factor, used to weigh the importance of future rewards; s (t+1) and s (t) and s (t+1) are the states at times t and t+1 respectively, and is the action executed by the target vehicle α at time t; is the target vehicle model π α under the influence of the background vehicle model π α state transition probability function, indicating the probability of transitioning to s (t) after executing the action ; s0 = s represents the initial state as s. (t+1) The probability of;

[0078] Advantages of the present invention:

[0079] A robust adversarial training method for autonomous driving based on an adaptive background vehicle model constructs a hybrid risk field to finely quantify the interaction risks between the target vehicle and background vehicles, and uses a reinforcement learning model to generate controllable adversarial actions. Meanwhile, an adaptive activation function and an alternating training mechanism are designed to dynamically balance the adversarial intensity and training efficiency, avoiding model collapse caused by excessive interference. Ultimately, the robustness of the autonomous driving system is improved in complex traffic environments, the accident rate is significantly reduced, and it is compatible with multiple reinforcement learning algorithms, providing a general adversarial training framework for different autonomous driving models.

[0080] A robust adversarial training method for autonomous driving based on an adaptive background vehicle model constructs an adaptive alternating training mechanism, alleviating the inherent instability problem of the reinforcement learning framework. In addition, the present invention allows for the gradual adjustment of the adversarial intensity faced by the target vehicle, simulating a series of new environmental conditions after switching, thereby smoothing the change curve of the adversarial intensity. This process not only enhances the adaptability and learning efficiency at different stages but also maintains the generality of training, ensuring robustness and safety in various driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 is a flowchart of a robust adversarial training method for autonomous driving based on an adaptive background vehicle model according to the present invention;

[0082] Figure 2 is a schematic diagram of collecting the driving parameters of the target vehicle and six specific-position background vehicles around it according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described specific embodiments are only a part of the embodiments of the present invention, rather than all of the specific embodiments. Usually, the components of the specific embodiments of the present invention described and shown in the accompanying drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0084] Therefore, the detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents the selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0085] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and described in detail in conjunction with the appended Figure 1 and the appended Figure 2 as follows:

[0086] Example 1:

[0087] A robust adversarial training method for autonomous driving based on an adaptive background vehicle model, comprising the following steps:

[0088] S1. Define the trained autonomous driving vehicle as the target vehicle model, define the vehicle models around the target vehicle model as the background vehicle models, construct a robust training environment for the autonomous driving model, and collect the driving parameters of the target vehicle model and the background vehicle models;

[0089] Furthermore, the specific implementation method of step S1 includes the following steps:

[0090] S1.1. Define the trained autonomous driving vehicle as the target vehicle, and the target vehicle model as π α ; define the vehicles around the target vehicle as the background vehicles, and the background vehicle model as π β ;

[0091] Set the target vehicle and the background vehicle as independent and competitive entities. The target vehicle and the background vehicle operate without communication and lack white-box access to each other's input states, and parameter information is not shared in the weight configuration;

[0092] Define the interaction process between the target vehicle model and the background vehicle model as a Markov process, and the expression is:

[0093] M = <S, (O α , O β ), (A α , A β ), P, (R α , R β ), γ> (1)

[0094] Wherein, S represents the state space, O α and O β are the state spaces of the target vehicle model and the background vehicle model respectively, A α and A β are the action spaces of the target vehicle model and the background vehicle model respectively, P is the transition probability function, R α and R β are the reward functions of the target vehicle model and the background vehicle model respectively, and γ is the discount factor;

[0095] P: S × A α × A β → S′

[0096] where S' represents the probability distribution of the next state;

[0097] With the goal of maximizing the cumulative reward, R: S × A α × A β × S′ → R;

[0098] S1.2. Define that the system uses V2X communication technology to collect the driving parameters of the target vehicle and six background vehicles at specific positions around it in real time. The set of background vehicles Ω b is defined as:

[0099] b ∈ Ω b = {front, left front, right front, rear, left rear, right rear} (2)

[0100] where b is the background vehicle;

[0101] Define the state space of the background vehicle to include the target vehicle state, background vehicle state, and lane characteristics. The target vehicle state includes the lateral position x of the target vehicle t , speed v t , acceleration a t and heading angle θ t , the background vehicle state includes the relative distance Δd between the nearby background vehicle and the target vehicle b , the speed v of the background vehicle b and acceleration a b , the lane characteristics include the average speed of each lane and density ρ lane of the information, and the data collection range covers 300 meters around the target vehicle;

[0102] Define the action space of the background vehicle to include candidate background vehicles Candidate adversarial behaviors aa, candidate adversarial behaviors include lane change, hard acceleration, hard deceleration, and normal behavior na.

[0103] S2. Construct a combined attenuation factor to analyze the risk between the background vehicle and the target vehicle, and use the time to collision TTC to construct a mixed risk field of the background vehicle for reverse risk assessment of the safety of the target vehicle;

[0104] Furthermore, the specific implementation method of step S2 includes the following steps:

[0105] S2.1. Construct a longitudinal attenuation factor φ x,b (Δx, Δy, t), which quantifies the risk attenuation along the longitudinal axis of the vehicle. The expression is:

[0106]

[0107] where Δx represents the longitudinal relative position, Δy is the lateral relative position, x is the longitudinal position coordinate, and L b is the length of the background vehicle, v x,b (t) represents the longitudinal speed of the vehicle at time t, and the parameters μ x,b and ρ x,b represent the longitudinal risk attenuation intensity based on vehicle type and the longitudinal risk attenuation intensity based on speed influence respectively, and max is the maximum value function;

[0108] S2.2. Construct the lateral attenuation factor φ y,b (Δx, Δy, t), which quantifies the risk attenuation along the lateral axis of the vehicle. The expression is:

[0109]

[0110] where y is the lateral position coordinate, W b is the width of the background vehicle, v y,b (t) is the lateral speed of the vehicle at time t, and the parameters μ y,b and ρ y,b represent the lateral risk attenuation intensity based on vehicle type and the lateral risk attenuation intensity based on speed influence respectively;

[0111] S2.3. Use the Euclidean norm to combine the longitudinal attenuation factor and the lateral attenuation factor to obtain the combined attenuation factor φ k (Δx, Δy, t). The expression is:

[0112]

[0113] S2.4. Introduce the TTC collision time into the artificial risk field. The calculation formulas of TTC on different axes are as follows:

[0114]

[0115] where TTC x is the longitudinal collision time, TTC y is the lateral collision time, and Δv x and Δv y are the longitudinal relative speed and the lateral relative speed between the background vehicle and the target vehicle respectively;

[0116] Construct the hybrid risk field of the background vehicle. The expression is:

[0117]

[0118] where η b represents the risk value of the hybrid risk field from the background vehicle b to the target vehicle, indicating the relative speed at time t (Δv x , Δvy ) The risk value at the position (Δx, Δy) around the background vehicle b.

[0119] Further, η b Ensure that the risk is maximum within the area occupied by the background vehicle and smoothly decreases to near zero as the distance increases. The constant 1 in the denominator ensures that the risk function remains valid even when the background vehicle speed is zero.

[0120] S3. Construct a probabilistic risk assessment framework, hierarchically divide the risk levels into three levels: high, medium, and low risk levels, and calculate the conditional probability, posterior probability, and surrounding environment probability between the background vehicle and the target vehicle at different risk levels based on the mixed risk field obtained in step S2;

[0121] Further, the specific implementation method of step S3 includes the following steps:

[0122] S3.1. Hierarchically divide the risk levels into three levels: high, medium, and low risk levels, and the risk levels are defined as:

[0123] τ ∈ Ω = {dangerous, attentive, safe} = {D, A, S} (9)

[0124] Where τ represents the risk level, Ω represents the set of risk levels, D represents high risk, A represents medium risk, and S represents low risk;

[0125] S3.2. Define the conditional probability distribution function F(η b |τ b ) of the background vehicle b at different risk levels, and its definition is as follows:

[0126]

[0127]

[0128] Where F(η b |τ b = D) is the conditional probability distribution function at the high risk level, F(η b |τ b = A) is the conditional probability distribution function at the medium risk level, F(η b |τ b = S) is the conditional probability distribution function at the low risk level, η D 、η S 、σ S 、σ D are respectively predefined high risk threshold hyperparameters, low risk threshold hyperparameters, low risk standard deviation hyperparameters, and high risk standard deviation hyperparameters;

[0129] S3.3. Calculate the posterior probability P(τ b of the risk level τ of the background vehicle b given η b using Bayesian inference. The expression is as follows: b |η b )

[0130]

[0131] where P(τ b ) represents the prior probability of the risk level τ b .

[0132] For simplicity in calculation, assume that the prior probability distributions for different risk levels are uniform and satisfy the normalization condition. The expression is:

[0133]

[0134] S3.4. Considering the risk levels of all background vehicles in the surrounding environment comprehensively, determine the overall risk distribution of the entire environment. The probability distribution of the risk level of the surrounding environment is defined as:

[0135] P(τ env = D) = 1 - ∏ b∈B (1 - P(τ b = D|η b )) (12)

[0136] P(τ env = S) = ∏ b∈B P(τ b = S|η b ) (13)

[0137] P(τ env = A) = 1 - P(τ env = S) - P(τ env = D) (14)

[0138] where P(τ env = D), P(τ env = A), and P(τ env = S) represent the probabilities of the surrounding environment being in high, medium, and low risk levels respectively, and are represented by P(τ b = D|η b ), P(τ b = A|η b ), P(τ b = S|η b ) which represent the probability of the risk level of the background vehicle b.

[0139] S4. Introduce a scoring function to map the risk level to a numerical score, calculate the expected risk of background vehicles and the expected risk of the surrounding environment, and construct an activation function to screen background vehicles and behaviors that are hostile to the target vehicle;

[0140] Furthermore, the specific implementation method of step S4 includes the following steps:

[0141] S4.1. The scoring function Sc(τ) is introduced to map the risk level to a numerical score, and D, A, and S are assigned 2, 1, and 0 respectively. The expected risk of background vehicles and the expected risk of the surrounding environment are calculated. The expression is:

[0142]

[0143] Among them, ∈ b is the expected risk of background vehicles, ∈ env is the expected risk of the surrounding environment, P(τ b |n b ) is the value in a given state n b Lower risk level τ b The conditional probability, Sc(τ b ) is the risk level τ b The score value, Sc(τ env ) is the risk level τ env Rating value of

[0144] S4.2. Based on the expected risk of candidate background vehicles Expected risk relative to the surrounding environment∈ env Evaluate whether the candidate background vehicle is suitable for performing adversarial actions and define the activation function for:

[0145]

[0146] Among them, k is a positive constant that determines the steepness of the function, M l and M u They are the upper and lower levels of the activation function respectively. If the result of the activation function is 1, it is determined that the candidate background vehicle performs the candidate adversarial behavior. If it is 0, it performs the normal behavior.

[0147] S5. Construct an adaptive alternating training mechanism, set the simulation training process to perform M rounds in total, and each round of training executes T time steps in total. In each simulation time step, the background vehicle model determines the background vehicle behavior according to the activation function of step S4, and the target vehicle model makes decisions based on the deep reinforcement learning strategy. During the simulation training process, the adaptive alternating training mechanism dynamically adjusts the intensity of the adversarial behavior according to the safety performance of the target vehicle model, and completes the autonomous driving robust adversarial training based on the adaptive background vehicle model.

[0148] Further, the specific implementation method of step S5 includes the following steps:

[0149] S5.1. Construct an adaptive alternating training mechanism, and adaptively switch the training process between the target vehicle and the background vehicle according to the switching indicator C R and use as an indicator, where C c represents the number of training rounds between two accidents;

[0150] The switching is controlled by the adaptive alternating training threshold H. When C R > H, the training of the background vehicle model is turned off, and the training of the target vehicle model is started to enhance its decision-making strategy under adversarial conditions;

[0151] S5.2. Set the goal of the background vehicle model during the training phase to use the fixed strategy of the target vehicle model during the non-training period of the target vehicle to optimize the strategy of the background vehicle model π β to maximize the cumulative discounted reward, and the expression is:

[0152]

[0153] where is the optimal strategy of the background vehicle model β, that is, the strategy that maximizes the long-term discounted reward; is the expectation under the background vehicle model π β strategy; π β is the reward function of the background vehicle model β, indicating the immediate reward obtained when transferring to the state s (t) after executing the action ; γ ∈ [0, 1) is the discount factor, which is used to weigh the importance of future rewards; s (t+1) and s (t) and s (t+1) are the states at times t and t + 1 respectively, and is the action executed by the background vehicle β at time t; is the state transition probability function of the background vehicle model π α under the influence of the target vehicle model π β , indicating the probability of transferring to s (t) after executing the action ; s0 = s represents the initial state as s; (t+1) ;

[0154] S5.3. When the background vehicle model is trained, introduce adversarial behavior into the deep reinforcement learning strategy π of the target vehicle model αMake a decision. When the adversarial behavior effectively creates failure cases for the target vehicle, update the weight parameters of the target vehicle model through retraining, while keeping the background vehicle model parameters fixed in the training environment. The policy π of the target vehicle model α is to maximize the cumulative discounted reward, and the expression is

[0155]

[0156] where is the optimal policy of the target vehicle model α, that is, the policy that maximizes the long-term discounted reward; is the expectation under the target vehicle model π α policy; R α is the reward function of the target vehicle model α, representing the immediate reward obtained when transferring to state s (t) after executing the action ; s (t+1) is the immediate reward obtained when transferring to state s (t) ; γ∈[0,1) is the discount factor, which is used to weigh the importance of future rewards; s (t+1) and s are the states at times t and t + 1 respectively, and is the influence of the background vehicle model π α on the target vehicle model π α state transition probability function, representing the probability of transferring to s (t) after executing the action ; s0 = s represents the initial state as s. (t+1)

[0157] Furthermore, the adoption of the adaptive alternating training mechanism alleviates the inherent instability problem of the reinforcement learning framework. In addition, the adaptive alternating training mechanism allows for the gradual adjustment of the adversarial intensity faced by the target vehicle, simulating a series of new environmental conditions after switching, thereby smoothing the change curve of the adversarial intensity. This process not only enhances the adaptability and learning efficiency at different stages but also maintains the generality of training, ensuring robustness and safety in various driving scenarios.

[0158] ​It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0159] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A robust adversarial training method for autonomous driving based on an adaptive background vehicle model, characterized in that The method includes the following steps: S1. Define the trained autonomous vehicle as the target vehicle model, define the vehicle models around the target vehicle model as the background vehicle models, construct the robustness training environment of the autonomous driving model, and collect the driving parameters of the target vehicle model and the background vehicle models; S2. Construct a combined attenuation factor to analyze the risk between the background vehicle and the target vehicle, and construct a hybrid risk field of the background vehicle using the time to collision (TTC) before collision for reverse risk assessment of the safety of the target vehicle; S3. Construct a probabilistic risk assessment framework, hierarchically divide the risk levels into three levels: high, medium, and low, and calculate the conditional probability, posterior probability, and surrounding environment probability between the background vehicle and the target vehicle at different risk levels based on the hybrid risk field obtained in step S2; S4. Introduce a scoring function to map the risk levels to numerical scores, calculate the expected risk of the background vehicle and the expected risk of the surrounding environment, and construct an activation function to screen the background vehicles and behaviors that pose an adversarial effect on the target vehicle; S5. Construct an adaptive alternating training mechanism, set the simulation training process to be carried out for M rounds, and each round of training executes T time steps. In each simulation time step, the background vehicle model determines the background vehicle behavior according to the activation function in step S4, and the target vehicle model makes decisions based on the deep reinforcement learning strategy. During the simulation training process, the adaptive alternating training mechanism dynamically adjusts the intensity of the adversarial behavior according to the safety performance of the target vehicle model to complete the robustness adversarial training of autonomous driving based on the adaptive background vehicle model.

2. The robust adversarial training method for autonomous driving based on an adaptive background vehicle model according to claim 1, characterized in that The specific implementation method of step S1 includes the following steps: S1.

1. Define the trained autonomous vehicle as the target vehicle, and the target vehicle model as π α ; Define the vehicles around the target vehicle as background vehicles, and the background vehicle model as π β ; Set the target vehicle and the background vehicle as independent and competing entities. The target vehicle and the background vehicle operate without communication and lack white-box access to each other's input states, and no parameter information is shared in the weight configuration; Define the interaction process between the target vehicle model and the background vehicle model as a Markov process, and the expression is: M = <S, (O α , O β ), (A α , A β ), P, (R α , R β ), γ> (1) Among them, S represents the state space, O α and O β are the state spaces of the target vehicle model and the background vehicle model respectively, A α and A β are the action spaces of the target vehicle model and the background vehicle model respectively, P is the transition probability function, R α and R β are the reward functions of the target vehicle model and the background vehicle model respectively, and γ is the discount factor; P: S × A α × A β → S′ where S′ represents the probability distribution of the next state; With the goal of maximizing cumulative rewards \(R: S\times A\) α \(\times A\) β \(\times S'\to R\); S1.

2. Define that the system uses V2X communication technology to collect the driving parameters of the target vehicle and six background vehicles at specific positions around it in real time. The set of background vehicles Ω b is defined as: b ∈ Ω b = {front, front left, front right, rear, rear left, rear right} (2) where b is the background vehicle; The state space of the defined background vehicle includes the target vehicle state, the background vehicle state, and lane characteristics. The target vehicle state includes the lateral position x of the target vehicle t , speed v t , acceleration a t and heading angle θ t . The background vehicle state includes the relative distance Δd between the nearby background vehicle and the target vehicle b , the speed v of the background vehicle b and acceleration a b . The lane characteristics include the average speed of each lane and density ρ lane of the information. The data acquisition range covers 300 meters around the target vehicle; Define the action space of the background vehicle to include candidate background vehicles Candidate adversarial behaviors aa, where candidate adversarial behaviors include lane changes, hard accelerations, hard decelerations, and normal behavior na.

3. A method for robust adversarial training of an autonomous driving based on an adaptive background vehicle model according to claim 1 or 2, characterized in that The specific implementation method of step S2 includes the following steps: S2.

1. Construct the longitudinal attenuation factor φ x,b (Δx, Δy, t), which quantifies the risk attenuation along the longitudinal axis of the vehicle, and the expression is: where Δx represents the longitudinal relative position, Δy is the lateral relative position, x is the longitudinal position coordinate, L b is the length of the background vehicle, v x,b (t) represents the longitudinal speed of the vehicle at time t, and the parameter μ x,b and ρ x,b respectively represent the longitudinal risk attenuation intensity based on the vehicle type and the longitudinal risk attenuation intensity based on the speed influence, and max is the maximum value function; S2.

2. Construct the lateral attenuation factor φ y,b (Δx, Δy, t), which quantifies the risk attenuation along the lateral axis of the vehicle, and the expression is: where y is the horizontal position coordinate, W b is the width of the background vehicle, v y,b (t) is the lateral speed of the vehicle at time t, and the parameters μ y,b and ρ y,b represent the lateral risk attenuation intensity based on vehicle type and the lateral risk attenuation intensity based on speed influence, respectively; S2.

3. Combine the longitudinal attenuation factor and the transverse attenuation factor using the Euclidean norm to obtain the combined attenuation factor φ k (Δx, Δy, t), and the expression is as follows: S2.

4. Introduce the TTC collision time into the artificial risk domain. The calculation formulas of TTC on different axes are as follows: Among them, TTC x is the longitudinal collision time, and TTC y is the lateral collision time. Δv x and Δv y are the longitudinal relative speed and the lateral relative speed between the background vehicle and the target vehicle, respectively; Construct the hybrid risk field of the background vehicle, and the expression is: Among them, η b represents the risk value of the hybrid risk field from the background vehicle b to the target vehicle, indicating the risk value at time t at the relative speed (Δv x , Δv y ) at the position (Δx, Δy) around the background vehicle b.

4. A method for robust adversarial training of autonomous driving based on an adaptive background vehicle model according to claim 3, wherein The specific implementation method of step S3 includes the following steps: S3.

1. Hierarchically divide the risk levels into three levels: high, medium, and low, and the risk levels are defined as: τ∈Ω={dangerous,attentive,safe}={D,A,S}(9) where τ represents the risk level, Ω represents the set of risk levels, D represents high risk, A represents medium risk, and S represents low risk; S3.

2. Define the conditional probability distribution function F(η b |τ b ) of the background vehicle b at different risk levels, which is defined as follows: where, F(η b |τ b = D) is the conditional probability distribution function under the high-risk level, F(η b |τ b = A) is the conditional probability distribution function under the medium-risk level, F(η b |τ b = S) is the conditional probability distribution function under the low-risk level, η D and η s , σ S and σ D are the predefined high-risk threshold hyperparameter, low-risk threshold hyperparameter, low-risk standard deviation hyperparameter, and high-risk standard deviation hyperparameter, respectively; S3.

3. Calculate the posterior probability P(τ b | η) of the risk level τ of the background vehicle b given η b using Bayesian inference. The expression is as follows: b | η b ) Among them, P(τ b ) represents the prior probability of risk level τ b ; For simplicity of calculation, assume that the prior probability distributions of different risk levels are uniform and satisfy the normalization condition, and the expression is: S3.

4. Comprehensively consider the risk levels of all background vehicles in the surrounding environment to determine the overall risk distribution of the entire environment. The probability distribution of the risk levels of the surrounding environment is defined as: P(τ env = D) = 1 - ∏ b∈B (1 - P(τ b = D|η b )) (12) P(τ env = S) = ∏ b∈B P(τ b = S|η b ) (13) P(τ env = A) = 1 - P(τ env = S) - P(τ env = D)(14) where P(τ env = D), P(τ env = A), and P(τ env = S) represent the probabilities of the surrounding environment being in three risk levels of high, medium, and low respectively, and are denoted by P(τ b = D|η b ), P(τ b = A|η b ), P(τ b = S|η b ) to represent the risk level probabilities of the background vehicle b.

5. A method for robust adversarial training of autonomous driving based on an adaptive background vehicle model according to claim 4, characterized in that, The specific implementation method of step S4 includes the following steps: S4.

1. Introduce the scoring function Sc(τ) to map the risk level to a numerical score. Assign 2, 1, and 0 to D, A, and S respectively, and calculate the expected risk of the background vehicle and the expected risk of the surrounding environment. The expression is as follows: where, ∈ b is the expected risk of the background vehicle, ∈ env is the expected risk of the surrounding environment, P(τ b |n b ) is the conditional probability of the risk level τ b under the given state n b , Sc(τ b ) is the score value of the risk level τ b , Sc(τ env ) is the score value of the risk level τ env ; S4.

2. According to the expected risk of the candidate background vehicle The expected risk relative to the surrounding environment ∈ env Evaluate whether the candidate background vehicle is suitable for performing adversarial actions and define the activation function as follows: where k is a positive constant that determines the steepness of the function, M l and M u are the upper and lower bounds of the activation function respectively. If the result of the activation function is 1, it determines that the candidate background vehicle performs the candidate adversarial behavior, and if it is 0, it performs the normal behavior.

6. A method for robust adversarial training of autonomous driving based on an adaptive background vehicle model according to claim 5, characterized in that, The specific implementation method of step S5 includes the following steps: S5.

1. Build an adaptive alternating training mechanism, according to the switching indicator C R Adaptively switch the training process between the target vehicle and the background vehicle, and adopt as an indicator, where C c represents the number of training rounds between two accidents; The switching is controlled by the adaptive alternating training threshold H. When C R > H, the background vehicle model training is turned off, and the target vehicle model training is started to enhance its decision-making strategy under adversarial conditions; S5.

2. The goal of setting up the background vehicle model during the training phase is to optimize the policy of the background vehicle model π using the fixed policy of the target vehicle model during the non-training period of the target vehicle to maximize the cumulative discounted reward, and the expression is as follows: β to maximize the cumulative discounted reward, and the expression is: Among them, is the optimal policy of the background vehicle model β, that is, the policy that maximizes the long-term discounted reward; is the expectation under the background vehicle model π β policy; R β is the reward function of the background vehicle model β, indicating the immediate reward obtained when transferring to the state s (t) after executing the action to the state s (t+1) ; γ ∈ [0, 1) is the discount factor, used to weigh the importance of future rewards; s (t) and s (t+1) are the states at times t and t + 1 respectively, and is the action executed by the background vehicle β at time t; is the state transition probability function of the background vehicle model π α under the influence of the target vehicle model π β , indicating the probability of transferring to s (t) after executing the action to s (t+1) ; s0 = s represents the initial state as s; S5.

3. After the background vehicle model is trained, by introducing adversarial behavior into the deep reinforcement learning policy π of the target vehicle model α for decision-making, when the adversarial behavior effectively creates failure cases for the target vehicle, update the weight parameters of the target vehicle model by retraining, while keeping the parameters of the background vehicle model fixed in the training environment. The policy π of the target vehicle model α is to maximize the cumulative discounted reward, and the expression is wherein, is the optimal policy for the target vehicle model α, that is, the policy that maximizes the long-term discounted reward; is the expectation under the target vehicle model π α policy; R α is the reward function of the target vehicle model α, representing the immediate reward obtained when transferring from state s (t) to execute the action and then transferring to state s (t+1) ; γ ∈ [0, 1) is the discount factor, used to weigh the importance of future rewards; s (t) and s (t+1) are the states at times t and t + 1 respectively, and is the action executed by the target vehicle α at time t; is the state transition probability function of the target vehicle model π α under the influence of the background vehicle model π α , representing the probability of transferring from state s (t) to execute the action and then transferring to s (t+1) ; s0 = s represents the initial state as s.

Citation Information

Patent Citations

  • Automatic driving control method based on adversarial reinforcement learning

    CN117826603A

  • Autonomous confrontation and automatic driving function safety verification method for background vehicle

    CN118521878A

  • Defense method and an application against adversarial examples based on feature remapping

    US20220172000A1