Defensive intelligent driving aid decision-making method based on inverse reinforcement learning

Through the method based on inverse reinforcement learning, the typical dangerous scenarios are simulated and reproduced and the defensive driving decision model is trained, which solves the problem that the existing technology is difficult to effectively realize defensive driving in complex traffic scenarios, and realizes high reliability and adaptability of intelligent driving assisted decisions.

CN120135154APending Publication Date: 2025-06-13TONGJI UNIV

Patent Information

Application Number
CN202510495886.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively realize defensive driving in complex traffic scenarios, especially in the problems of human driver response speed, experience differences and distractions, and it is difficult to maintain a high level of defensive driving capabilities at all times.

Method used

The defensive intelligent driving assisted decision-making method based on inverse reinforcement learning is adopted, and a typical dangerous scenario is reproduced, and a decision-making model with defensive driving ability is collected through simulation and reproducible risk scenarios, and a decision-making model with defensive driving ability is trained using the generative reinforcement learning method. The model evaluates the driving risks of the current scenario through risk indicators and performs intelligent driving assistance operations, including active defense driving behavior suggestions and passive risk aversion driving behavior intervention.

Benefits of technology

Effectively learn defensive driving behavior from human drivers, understand and respond to complex and changeable traffic scenarios, improve the reliability of defensive driving assist decision-making, significantly improve the generalization ability and adaptability of the system, and achieve preventive driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120135154A_ABST
    Figure CN120135154A_ABST
Patent Text Reader

Abstract

The invention relates to a defensive intelligent driving aid decision-making method based on inverse reinforcement learning, and the method comprises the steps: simulating and reproducing a typical dangerous scene, and collecting the defensive driving behavior data of a driver; based on a generative adversarial reinforcement learning AIRL method, training by using the collected defensive driving behavior data to obtain a defensive driving decision model with defensive driving ability; and the driving risk of the current scene is evaluated by adopting the risk index, and intelligent driving assistance operation including active defense driving behavior suggestion and passive risk avoiding driving behavior intervention is executed by the defensive driving decision-making model according to the driving risk classification. Compared with the prior art, the method can effectively learn defensive driving behaviors from human drivers, accurately understand and cope with complex and variable traffic scenes, improve the reliability of defensive driving aid decision making, and realize preventive driving safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent driving, and particularly to a defensive intelligent driving assistance decision-making method based on inverse reinforcement learning. Background Art

[0002] In the safety driving system, defensive driving is a driving concept that proactively anticipates potential risks and takes evasive measures in advance to prevent accidents. Its core lies in the driver's ability to continuously observe and analyze the traffic environment, identify potential hazards in advance (such as sudden lane changes by other vehicles, pedestrians crossing illegally, blind spots, etc.), and take reasonable measures (such as adjusting vehicle speed, maintaining a safe distance, actively avoiding, etc.) to reduce the probability of accidents. Defensive driving can significantly reduce traffic accidents caused by driver reaction lags or decision-making mistakes, and has important safety value especially in complex traffic scenarios.

[0003] However, human drivers are limited by problems such as reaction speed, experience differences, and attention dispersion, and it is difficult to always maintain a high level of defensive driving ability. Therefore, embedding defensive driving logic into intelligent driving assistance systems to achieve all-weather and high-precision risk prediction and active defense through intelligent decision-making has become a key direction for improving driving safety.

[0004] Traditionally, the construction of defensive driving decision models adopts rule-driven or imitation learning methods. Among them, the rule-driven method realizes basic risk avoidance through preset traffic rules and safety thresholds. For example, Patent JP 2020147271A proposes a defensive driving strategy generation scheme that determines whether there is a collision risk by detecting multiple obstacle types within the vehicle's perceptible range and combining a collision risk detection method; when there is a collision risk, the vehicle's defensive driving strategy is determined. However, this rule-driven approach depends on rigid rules designed manually, and it is difficult to cover complex and variable long-tail scenarios, such as scenarios where aggressive vehicles cut in, unstructured roads, sudden obstacles, etc. Moreover, overly conservative rules may lead to a decrease in traffic efficiency; although imitation learning can reproduce driving behaviors through expert data, its essence is behavior cloning, lacking a deep understanding of the risk perception and decision-making logic behind expert behaviors, resulting in insufficient generalization ability of the model in scenarios. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a defensive intelligent driving assistance decision-making method based on inverse reinforcement learning, which can effectively learn defensive driving behaviors from human drivers, accurately understand and respond to complex and variable traffic scenarios, and improve the reliability of defensive driving assistance decision-making.

[0006] The object of the present invention can be achieved by the following technical solutions: A defensive intelligent driving assistance decision-making method based on inverse reinforcement learning, comprising the following steps:

[0007] S1. Simulate and reproduce typical dangerous scenarios, and collect defensive driving behavior data of drivers;

[0008] S2. Based on the Adversarial Inverse Reinforcement Learning (AIRL) method, use the collected defensive driving behavior data to train a defensive driving decision-making model with defensive driving capabilities;

[0009] S3. Use a risk index to evaluate the driving risk of the current scenario. According to the driving risk classification, the defensive driving decision-making model performs intelligent driving assistance operations, including active defensive driving behavior suggestions and passive hazard avoidance driving behavior interventions.

[0010] Further, the step S1 includes the following steps:

[0011] S11. Select typical dangerous driving scenarios

[0012] Collect accident data in natural driving scenarios, and screen through expert analysis various scenarios where accidents occur due to insufficient defensive driving capabilities of drivers as typical dangerous scenarios for judging whether drivers have defensive driving capabilities;

[0013] S12. Select expert drivers with defensive driving characteristics;

[0014] S13. Simulate and reproduce typical dangerous driving scenarios, and synchronously collect the driving behavior data of expert drivers in dangerous scenarios;

[0015] S14. Screening of driving behavior data

[0016] In the collected driving behavior data, screen the data where expert drivers have successfully avoided accidents in dangerous scenarios due to taking active defensive driving behaviors, that is, as defensive driving behavior data.

[0017] Further, the step S13 is specifically based on real vehicle field tests or driving simulators to simulate and reproduce typical dangerous driving scenarios.

[0018] Further, the step S2 includes the following steps:

[0019] S21. Based on the collected defensive driving behavior data, construct an expert data set;

[0020] S22. Infer the reward function using the generative adversarial inverse reinforcement learning method based on the expert dataset, and train to obtain a defensive driving decision-making model.

[0021] Further, the specific process of step S21 is as follows:

[0022] Define the state space S and action space A of the defensive driving decision-making model, and convert each piece of collected defensive driving behavior data into a corresponding state-action pair data trajectory to construct an expert dataset.

[0023] Further, the state space S is composed of the road topology, the positions and motion states of each traffic participant, and the safety distance characteristics between the vehicle and other traffic participants;

[0024] The action space A is continuous and includes the acceleration, braking, and steering actions taken to control the vehicle;

[0025] The state-action pair specifically includes the current environmental state s, the action a taken by the expert driver, and the next environmental state s'.

[0026] Further, the specific process of step S22 is as follows:

[0027] S221. Initialize the simulation environment, the generator π, and the discriminator D θ,φ ;

[0028] S222. Use the generator π to interact with the simulation environment to collect the data trajectory τ i =(s 0 ,a 0 ,…,s T ,a T );

[0029] S223. Train the discriminator D based on binary logistic regression θ,φ , to distinguish between expert trajectory data and the generated trajectory data τ i ;

[0030] S224. Update the reward function r based on the discriminator D θ,φ ; θ,φ(s,a,s′) ;

[0031] S225. Based on the reward function r θ,φ(s,a,s′) , use the proximal policy optimization (PPO) reinforcement learning method to update the generator model π;

[0032] S226. Repeat steps S222 to S225 until the parameters of the generator π converge to the optimal, that is, the trained defensive driving decision-making model is obtained.

[0033] Further, in step S224, the reward function r is updated θ,φ(s,a,s′) Specifically:

[0034] r θ,φ(s,a,s′) = logD θ,φ(s,a,s′) - log(1 - D θ,φ(s,a,s′) )

[0035] Further, the risk index in step S3 is specifically an evaluation index based on the potential field theory, that is, the influence of each traffic element in the scene on the vehicle driving is represented by the potential field theory, and then the driving risk is evaluated

[0036] Further, the proactive defensive driving behavior suggestion in step S3 means that when the scene driving risk level exceeds the set potential driving risk threshold, the defensive driving decision-making model correspondingly outputs a defensive driving behavior suggestion:

[0037]

[0038] The passive risk avoidance driving behavior intervention means that when the scene driving risk level exceeds the set actual driving risk threshold, the defensive driving decision-making model correspondingly outputs an instruction to intervene in the driver's operation to force the vehicle to take an obstacle avoidance driving behavior:

[0039]

[0040] Wherein, C potl 、C subs respectively represent the scene sets with potential driving risks and actual driving risks. c ∈ C represents the current driving scene of the vehicle, C is the set of all scenes, and R(c) represents the driving risk level in scene c respectively represent the potential driving risk threshold and the actual driving risk threshold of the scene

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] The present invention simulates and reproduces typical dangerous scenarios, and collects defensive driving behavior data of drivers; then, based on the generative adversarial inverse reinforcement learning method, using the collected defensive driving behavior data, a defensive driving decision-making model with defensive driving ability is trained; furthermore, a risk index is used to evaluate the driving risk of the current scenario, and according to the driving risk classification, the defensive driving decision-making model performs intelligent driving assistance operations, including active defensive driving behavior suggestions and passive risk avoidance driving behavior interventions. Thereby, it can effectively learn defensive driving behaviors from human drivers, the driving style conforms to human driving habits, and it can correspond to remind drivers to take defensive driving measures in advance according to the driving risk of the current scenario, so as to avoid actively causing or being involved in traffic accidents passively, and achieve preventive driving safety.

[0043] The present invention collects accident data in natural driving scenarios, and through expert analysis, screens various scenarios in which accidents occur due to insufficient defensive driving ability of drivers as typical dangerous scenarios for judging whether drivers have defensive driving ability. In addition, after collecting the driving behavior data of expert drivers in dangerous scenarios, continue to screen the driving behavior data, that is, combine standards such as maintaining a safe distance and stable driving to screen the data, which can effectively screen out high-quality defensive driving behavior data and is beneficial to the accuracy of subsequent training of the defensive driving decision-making model.

[0044] The present invention adopts the method of adversarial inverse reinforcement learning (AIRL), which can deeply explore the defensive driving logic of human drivers and train a driving decision-making model with defensive driving ability from the collected defensive driving behavior data. This model can better understand and respond to complex and changeable traffic scenarios, and significantly improve the generalization ability and adaptability of the system.

[0045] The present invention applies a defensive driving decision-making model and performs intelligent driving assistance operations based on the driving risk classification, avoiding excessive interference with driving behaviors and thus affecting the comfort of drivers. At the same time, it can remind drivers to take defensive driving measures in advance to avoid actively causing or being involved in traffic accidents passively, achieving preventive safety and improving driving safety. Brief Description of the Drawings

[0046] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0047] Figure 2 It is a schematic diagram of the inverse reinforcement learning process in the present invention;

[0048] Figure 3 It is a continuous scenario driving risk map when the driver does not adopt the active defensive driving behavior suggestion in the embodiment;

[0049] Figure 4It is the continuous scenario driving risk map when the driver adopts the active defense driving behavior suggestion in the embodiment. Detailed implementation manners

[0050] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0051] Embodiment

[0052] As Figure 1 shown, a defensive intelligent driving assistance decision-making method based on inverse reinforcement learning includes the following steps:

[0053] S1. Simulate and reproduce typical dangerous scenarios, and collect defensive driving behavior data of the driver;

[0054] S2. Based on the generative adversarial inverse reinforcement learning AIRL method, use the collected defensive driving behavior data to train a defensive driving decision-making model with defensive driving ability;

[0055] S3. Use a risk index to evaluate the driving risk of the current scenario. According to the driving risk classification, the defensive driving decision-making model performs intelligent driving assistance operations, including active defense driving behavior suggestions and passive hazard avoidance driving behavior interventions.

[0056] This embodiment applies the above solution, and the main contents are as follows:

[0057] 1. Based on real vehicle field tests or driving simulators, simulate and reproduce typical dangerous scenarios, and synchronously collect defensive driving behavior data of the driver as expert data for subsequent training of the defensive driving decision-making model, specifically as follows:

[0058] 1.1 Selection of typical dangerous driving scenarios

[0059] First, summarize the accident data in natural driving scenarios, and screen out various scenarios in which accidents occur due to insufficient defensive driving ability of the driver through expert analysis as typical dangerous scenarios for judging whether the driver has defensive driving ability;

[0060] Among them, the defensive driving ability refers to the ability of the driver to foresee the potential collision risks brought about by the vehicle's failure to ensure sufficient safety distances from other traffic participants and unstable driving and to take reasonable and effective measures in advance to prevent accidents.

[0061] In this embodiment, the typical dangerous scenarios obtained through screening mainly include accident scenarios caused by blind spots and accident scenarios caused by the uncertainty of the movement of surrounding vehicles. The accident scenarios caused by blind spots refer to the obstacles or vehicles in the front that block the driver's line of sight, resulting in the driver being unable to judge in advance whether there are traffic participants in the blind spot, and thus colliding with the traffic participants suddenly emerging from the blind spot; the accident scenarios caused by the uncertainty of the movement of surrounding vehicles refer to the unexpected behaviors of surrounding vehicles such as sudden lane changes or braking, which cause the driver to be unable to take preventive measures in advance and thus collide with the surrounding vehicles.

[0062] 1.2 Selection of expert drivers

[0063] In this embodiment, a research questionnaire on defensive driving ability is designed, and drivers with the characteristics of defensive driving are preliminarily screened out through the questionnaire as expert drivers;

[0064] Among them, the research questionnaire on defensive driving ability mainly consists of parts such as driver's personal information, experimental matching degree, driving experience, driving habits, and defensive driving ability test. The experimental matching degree and driving experience information are used to screen out skilled drivers who meet the data collection requirements; the driving habit information and defensive driving ability test are used to further screen out drivers with the characteristics of defensive driving such as risk prediction, looking far ahead, taking overall situation into account, leaving room, and attracting attention.

[0065] 1.3 Collection of driving behavior data

[0066] Based on on-road vehicle tests or driving simulators, typical dangerous driving scenarios are simulated and reproduced, and the driving behavior data of expert drivers in the dangerous scenarios are collected synchronously;

[0067] In this embodiment, the method of on-road vehicle tests is adopted, which is carried out in a closed test field. The traffic participants in the scenario include remotely controllable dummy cars, movable dummies, etc. The test vehicle is equipped with in-vehicle sensors such as cameras, radars, and steering wheel force feedback devices, which can collect driver operation data. Among them, the method of reproducing typical dangerous scenarios is as follows: build the road environment corresponding to the dangerous scenario in the closed test field and set controllable traffic participants; then, an expert driver drives the test vehicle into the test scenario. When the vehicle approaches the preset dangerous trigger area, the background system activates dangerous elements such as pedestrians rushing out of the blind spot and sudden lane changes of surrounding vehicles in real time through wireless communication, and records the driver's operations and the vehicle running state through in-vehicle sensors and external monitoring devices at the same time, so as to obtain the defensive driving behavior data under typical dangerous scenarios.

[0068] 1.4 Screening of driving behavior data

[0069] Among the collected driving behavior data, filter out the data where accidents are successfully avoided due to the proactive defensive driving behaviors of expert drivers in dangerous scenarios as defensive driving behavior data, which is used to construct an expert dataset for training a defensive driving decision-making model later.

[0070] II. Based on the generative adversarial inverse reinforcement learning (AIRL), use the collected defensive driving behavior data to train a driving decision-making model with defensive driving capabilities. Specifically as follows:

[0071] 2.1 Construction of expert dataset

[0072] Define the state space S and action space A of the driving decision-making model, and convert each piece of collected defensive driving behavior data into a corresponding state-action pair data trajectory to construct an expert dataset;

[0073] In this embodiment, the state space S includes the relative position information Δv of the host vehicle and the surrounding background vehicles in the scenario s , relative speed information Δv d , relative heading angle information Δe, the safety distance SA between the vehicle and each traffic participant, and the road topology information T;

[0074] The action space A includes the throttle pedal opening α, the brake pedal opening δ, and the steering wheel angle ω;

[0075] The state-action pair specifically includes the current environmental state s, the action a taken by the expert driver, and the next environmental state s'.

[0076] 2.2 Construction of defensive driving decision-making model

[0077] As Figure 2 shown, based on the expert dataset, use the method of generative adversarial inverse reinforcement learning (AIRL) to infer the reward function and train the defensive driving decision-making model. The specific steps are as follows:

[0078] 1) Initialize the Carla simulation environment, the generator π, and the discriminator D θ,φ , and keep the settings of the simulation environment consistent with those corresponding to the collection of expert data on the real vehicle site;

[0079] 2) Use the generator π to interact with the simulation environment to collect the data trajectory τ i =(s 0 ,a 0 ,…,s T ,a T )

[0080] 3) Train the discriminator D based on binary logistic regression θ,φ , used to distinguish expert trajectory data from the generated trajectory data τi ;

[0081] 4), Based on the discriminator D θ,φ Update the reward function r θ,φ(s,a,s ′ ) , and the update method can be expressed by the following formula:

[0082] r θ,φ(s,a,s′) = logD θ,φ(s,a,s′) - log(1 - D θ,φ(s,a,s′) )

[0083] 5), Based on the reward function r θ,φ(s,a,s′) , adopt the Proximal Policy Optimization (PPO) method of reinforcement learning to update the generator model π;

[0084] 6), Repeat steps 2) to 5) until the parameters of the generator π converge to the optimal, which is the trained defensive driving decision-making model.

[0085] III. Evaluate the driving risk of the current scenario using risk indicators, and the defensive driving decision-making model then performs intelligent driving assistance operations according to the driving risk classification, including proactive defensive driving behavior suggestions and passive hazard avoidance driving behavior interventions.

[0086] In this embodiment, the driving risk indicator refers to an evaluation indicator based on the potential field theory. This indicator constructs a driving risk field through the potential field theory to represent the influence of various traffic elements in the scenario on vehicle driving, and then evaluates the driving risk. It can be calculated by the following formula:

[0087] R potl (c)= E j M j r j exp[- k 2 v j cos(θ j )](1 + D rj )

[0088] In the formula: r potl (s) represents the driving risk of the ego vehicle j in the current scenario c; R j is the road condition influence factor at the location of the ego vehicle j(x j , y j ), the x-axis is along the road line direction, and the y-axis is perpendicular to the road line direction; M j is the equivalent mass of the ego vehicle j; v j is the vehicle speed of the ego vehicle j; θ j is the angle between the speed direction of the ego vehicle j and the E j direction; D rj is the driver risk factor of the ego vehicle j; E j is the driving risk field at (x j , y j) The field strength at the position can be expressed by the following formula:

[0089]

[0090] Where: E V_pj , E R_nj , E D_qj Are respectively the field strength vectors of the single kinetic energy field, potential energy field and behavior field of the object at the position of the self-vehicle j, and can be expressed by the following formula:

[0091]

[0092] Where: R i Is the road condition influence factor of the object i(x i , y i ); r ij =(x j -x i , y j -t i ) represents the distance vector between the object i and the self-vehicle j; k1, k2, G are all undetermined constants greater than 0; M i Is the equivalent mass of the object i; the object i moves in the positive x-axis direction, v i Is the speed of the object i: θ i Is the angle between the speed direction of the object i and r ij , with the clockwise direction being positive; Dr i Is the risk factor of the driver i.

[0093] In this embodiment, the active defense driving behavior suggestion means that when the scene driving risk level exceeds the set potential driving risk threshold, the defensive driving decision-making model provides the defensive driving behavior suggestions that the driver can take; the passive risk avoidance driving behavior intervention means that when the scene driving risk level exceeds the set substantial driving risk threshold, the defensive driving decision-making model intervenes in the driver's operation and forces the vehicle to take obstacle avoidance driving behavior. When judging the driving risk of the current scene, it is judged by the following formula:

[0094]

[0095] Where: C potl , C subs Respectively represent the scene sets where there are potential driving risks and substantial driving risks. c ∈ C represents the current driving scene of the vehicle. C is the set of all scenes, and R(c) represents the driving risk level in the scene c. Respectively represent the potential driving risk threshold and substantial driving risk threshold of the scene.

[0096] In this embodiment, the trained defensive driving decision-making model is applied to the vehicle intelligent driving assistance system, such as Figure 3 ,Figure 4 As shown, as the driving risk level evaluated by the risk indicator continuously increases and exceeds the potential driving risk threshold, the intelligent driving assistance system identifies the current scenario as a potential driving risk scenario and starts to actively provide the driver with defensive driving behavior suggestions based on the defensive driving decision-making model, including behaviors such as decelerating and changing lanes.

[0097] As Figure 3 shown, if the driver does not adopt the system's defensive driving behavior suggestions and does not take other safe driving behaviors but keeps the behavior unchanged, the driving risk level will continue to increase and exceed the substantial driving risk threshold. The intelligent driving assistance system identifies the current scenario as a substantial driving risk scenario, starts to enable the defensive driving decision-making model to take over the driver's operation, and forcibly intervenes to make the vehicle take evasive driving behaviors to avoid accidents.

[0098] As Figure 4 shown, if the driver adopts the system's defensive driving behavior suggestions and takes measures in advance to avoid potential collision risks, the driving risk level will start to decrease and gradually fall below the potential driving risk threshold, and the intelligent driving assistance system will exit the activation state. This indicates that the driver has achieved preventive safety with the help of the intelligent driving assistance system.

Claims

1. A defensive intelligent driving assistance decision-making method based on inverse reinforcement learning, characterized in that: The following steps are involved: S1. Simulate and reproduce typical dangerous scenarios and collect the driver's defensive driving behavior data; S2. Based on the generative adversarial reinforcement learning (AIRL) method, the collected defensive driving behavior data is used to train a defensive driving decision model with defensive driving capabilities. S3. Use risk indicators to assess the driving risk of the current scenario. According to the driving risk classification, the defensive driving decision model performs intelligent driving assistance operations, including active defensive driving behavior recommendations and passive risk-avoidance driving behavior interventions.

2. According to claim 1, a defensive intelligent driving assistance decision-making method based on inverse reinforcement learning is characterized in that: The step S1 comprises the following steps: S11. Select typical dangerous driving scenarios Collect accident data from natural driving scenarios, and use expert analysis to screen various scenarios where accidents occur due to drivers' insufficient defensive driving ability as typical dangerous scenarios for judging whether drivers have defensive driving ability; S12. Select expert drivers with defensive driving characteristics; S13, simulating and reproducing typical dangerous driving scenarios, and simultaneously collecting driving behavior data of expert drivers in dangerous scenarios; S14. Driving behavior data screening Among the collected driving behavior data, the data that successfully avoids accidents in dangerous scenarios due to the expert drivers taking active defensive driving behaviors are screened, that is, the defensive driving behavior data.

3. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 2 is characterized in that: The step S13 is specifically based on a real vehicle field test or a driving simulator to simulate and reproduce typical dangerous driving scenarios.

4. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 1 is characterized in that: The step S2 comprises the following steps: S21. Construct an expert dataset based on the collected defensive driving behavior data; S22. Based on the expert dataset, the generative adversarial reinforcement learning method is used to infer the reward function and train the defensive driving decision model.

5. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 4 is characterized in that: The specific process of step S21 is as follows: The state space S and action space A of the defensive driving decision model are defined, and each piece of collected defensive driving behavior data is converted into the corresponding state-action pair data trajectory to construct an expert dataset.

6. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 5 is characterized in that: The state space S is composed of the road topology, the position and motion state of each traffic participant, and the safety distance characteristics between the vehicle and other traffic participants; The action space A is continuous and includes the acceleration, braking, and steering actions taken by the vehicle; The state-action pair specifically includes the current environment state s, the action a taken by the expert driver, and the next environment state s ′ .

7. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 6 is characterized in that: The specific process of step S22 is as follows: S221, initialize the simulation environment, generator π and discriminator D θ,φ ; S222. Use the generator π to interact with the simulation environment and collect data trajectories τ i =(s0,a0,…,s T ,a T ); S223, training discriminator D based on binary logistic regression θ,φ , used to distinguish expert trajectory data And generate trajectory data τ i ; S224, based on discriminator D θ,φ Update the reward function r θ,φ(s,a,s′) ; S225. Based on the reward function r θ,φ(s,a,s′) , using the reinforcement learning method PPO, update the generator model π; S226. Repeat steps S222 to S225 until the generator π parameter converges to the optimal value, that is, the trained defensive driving decision model.

8. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 7 is characterized in that: In step S224, the reward function r is updated. θ,φ(s,a,s′) Specifically: r θ,φ(s,a,s′) =logD θ,φ(s,a,s′) -log(1-D θ,φ(s,a,s′) )。 9. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 1 is characterized in that: The risk index in step S3 is specifically an evaluation index based on potential field theory, that is, the impact of various traffic elements in the scene on vehicle driving is represented by potential field theory, thereby evaluating driving risks.

10. The defensive intelligent driving assistance decision-making method based on inverse reinforcement learning according to claim 1, characterized in that: The active defensive driving behavior suggestion in step S3 means that when the driving risk level of the scenario exceeds the set potential driving risk threshold, the defensive driving decision model outputs the defensive driving behavior suggestion accordingly: Passive risk-avoidance driving behavior intervention means that when the driving risk level of the scenario exceeds the set actual driving risk threshold, the defensive driving decision model will output instructions to intervene in the driver's operation to force the vehicle to take obstacle avoidance driving behavior: Among them, C potl , C subs They represent the scene sets with potential driving risks and actual driving risks respectively, c∈C represents the current driving scene of the vehicle, C is the set of all scenes, R(c) represents the driving risk level under scene c, They represent the potential driving risk threshold and actual driving risk threshold of the scenario respectively.

Citation Information

Patent Citations

  • Method, apparatus, device, storage medium and program for generating defensive driving strategy

    JP2020147271A

Cited By

  • Convergence area main road vehicle driving behavior modeling method based on deep inverse reinforcement learning

    CN120316454A

  • Decision model training method and related equipment

    CN121809584A