Decision vulnerability detection method and device for automatic driving decision model

Through multi-agent reinforcement learning methods, the detection scenario is built in the autonomous driving simulator, and the reward mechanism is used to train the attack decision model, which solves the problem of incomplete detection of decision vulnerabilities of the autonomous driving decision model in the existing technology, and realizes efficient and low-cost decision vulnerability detection.

CN120046150APending Publication Date: 2025-05-27TOYOTA JIDOSHA KK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311594113.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing autonomous driving simulators cannot effectively detect decision vulnerabilities, resulting in frequent wrong decision-making in the real world of the autonomous driving decision model, and the real world road testing is high, low efficiency, and incomplete detection.

Method used

Multi-agent reinforcement learning methods are adopted to build detection scenarios, and the attack decision model is trained through accident responsibility arbitration rewards and diversity accident scenario rewards to generate accident scenarios caused by decision vulnerabilities in the autonomous driving decision model, thereby accurately detecting decision vulnerabilities.

Benefits of technology

Accurate detection of decision vulnerabilities in the decision-making model of autonomous driving is achieved, the detection efficiency and cost-effectiveness are improved, and the high cost and low efficiency problems of real-world road testing are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046150A_ABST
    Figure CN120046150A_ABST
Patent Text Reader

Abstract

The invention discloses a decision vulnerability detection method and device for an automatic driving decision model. The method comprises the following steps: constructing a detection scene for detecting the automatic driving decision model of a target vehicle; based on a multi-agent learning algorithm, training the first attack decision model by using a behavior reward constructed for the attacking vehicle to obtain a corresponding second attack decision model; based on the detection scene, a target vehicle with an automatic driving decision model interacts with an attack vehicle with a second attack decision model, and an accident scene caused by decision vulnerabilities of the automatic driving decision model is generated; and determining a decision vulnerability of the automatic driving decision model from the accident scene based on the second attack decision model. According to the method, the decision-making vulnerability of the automatic driving decision-making model can be accurately detected, so that the robustness of the automatic driving decision-making model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent vehicles, and particularly to a method and device for detecting decision loopholes in an autonomous driving decision-making model. Background Art

[0002] With the development of key technologies for autonomous driving and their supporting hardware, the deployment of autonomous driving technology on real vehicles has entered the implementation stage. As a field closely related to road safety, any decision-making error of an autonomous driving vehicle may cause major personal and property safety accidents. Although most autonomous driving companies deploying real vehicles have conducted large-scale long-distance real-world autonomous driving road tests to verify the safety of their vehicles, in recent years, there have still been frequent traffic accident cases caused by decision-making errors of autonomous driving vehicles. At the same time, the real-world autonomous driving road test data collection is inefficient in terms of time, highly restricted by sites, and extremely costly, further hindering the implementation process of the real vehicle deployment of autonomous driving technology.

[0003] To address the disadvantages of high cost and low efficiency in real-world autonomous driving road tests, simulators developed for autonomous driving are gradually being applied to the testing and data collection of autonomous driving decisions. Autonomous driving simulators such as CARLA and SMARTS support the simulation requirements of autonomous driving at different levels of fidelity, improving the testing efficiency and reducing the testing cost. However, existing autonomous driving simulators mainly simulate traffic flow and vehicle behavior in the real world. Testing autonomous driving decisions in such an environment can improve the operation efficiency, but the occurrence frequency of key accident cases is not much different from that in the real world, and key accident cases related to autonomous driving decision loopholes are still very sparse. Even due to the determinacy of traffic participant modeling design, some key accident cases with harsh conditions cannot be constructed in the simulator, resulting in missed detection of decision loopholes in the autonomous driving decision-making model. Currently, there is no complete process framework for detecting (mining) autonomous driving decision loopholes, nor is there a general method for detecting (mining) decision loopholes in autonomous driving decision-making models. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a method and device for detecting decision loopholes in an autonomous driving decision-making model, which can accurately detect decision loopholes in the autonomous driving decision-making model loaded on an autonomous driving vehicle.

[0005] To achieve the above purpose, the present application provides a method for detecting decision loopholes in an autonomous driving decision-making model, including:

[0006] Construct a detection scenario for detecting an autonomous driving decision model of a target vehicle, where the detection scenario includes at least one of the following: road environment, the autonomous driving decision model, one or more attacking vehicles for attacking the target vehicle, and a first attacking vehicle decision model of the attacking vehicle;

[0007] Based on a multi-agent learning algorithm, use the behavior rewards constructed for the attacking vehicle to train the first attacking decision model to obtain a corresponding second attacking decision model, where the behavior rewards include at least one of the following: accident liability arbitration rewards and diverse accident scenario rewards. The accident liability arbitration rewards are used to determine the accident liability according to traffic rules through the vehicle driving information during the accident process between the target vehicle and the attacking vehicle. The diverse accident scenario rewards are determined according to the dissimilarity between accidents;

[0008] Based on the detection scenario, interact the target vehicle with the autonomous driving decision model and the attacking vehicle with the second attacking decision model to generate an accident scenario caused by a decision vulnerability of the autonomous driving decision model;

[0009] Based on the second attacking decision model, determine the decision vulnerability of the autonomous driving decision model from the accident scenario.

[0010] Optionally, the construction of the detection scenario for detecting the autonomous driving decision model of the target vehicle includes:

[0011] Construct the road environment in a simulator;

[0012] Determine the initial position and destination of the target vehicle, and the initial position of the attacking vehicle.

[0013] Optionally, the input of the first attacking decision model is the state information of the attacking vehicle and the relative position relationship between the attacking vehicle and the target vehicle; the output of the first attacking decision model is the decision action of the attacking vehicle.

[0014] Optionally, the behavior rewards further include distance rewards, which are used to reward the attacking vehicle according to the distance between the target vehicle and the attacking vehicle. The training of the first attacking decision model using the behavior rewards constructed for the attacking vehicle includes:

[0015] Based on the state of the detection scenario, determine the initial decision action of the attacking vehicle;

[0016] Adjust the state of the detection scenario based on the initial decision-making action of the attacking vehicle, and generate the corresponding behavior reward;

[0017] Reward the attacking vehicle based on the accident liability arbitration reward, the diversity accident scenario reward, and / or the distance reward.

[0018] Optionally, the distance reward includes a negative penalty and a positive reward. Training the first attack decision model using the behavior reward constructed for the attacking vehicle includes:

[0019] When the distance between the attacking vehicle and the target vehicle exceeds a first predetermined distance, give the attacking vehicle the negative penalty to encourage the attacking vehicle to approach the target vehicle;

[0020] When the attacking vehicle gradually approaches the target vehicle, the attacking vehicle will receive an increasingly larger positive reward;

[0021] When the attacking vehicle approaches the target vehicle and the distance is less than a second predetermined distance, give the attacking vehicle the negative penalty.

[0022] Optionally, the accident liability arbitration reward includes a collision reward and an aggression penalty. Training the first attack decision model using the behavior reward constructed for the attacking vehicle includes:

[0023] When it is determined that the attacking vehicle collides with the target vehicle, or an environmental vehicle in the detection scenario collides with the target vehicle, give the attacking vehicle the collision reward;

[0024] When the attacking vehicle violates the driving regulation restrictions, give the attacking vehicle the aggression penalty to standardize the driving actions.

[0025] Optionally, training the first attack decision model using the behavior reward constructed for the attacking vehicle includes:

[0026] When it is determined that the attacking vehicle explores a second dangerous scenario different from the first dangerous scenario that has been discovered, give the attacking vehicle the diversity accident scenario reward.

[0027] Optionally, where when there are multiple attacking vehicles, the interaction between the target vehicle with the autonomous driving decision model and the attacking vehicle with the second attack decision model includes:

[0028] In the detection scenario, multiple of the attacking vehicles coordinate with each other to attack the target vehicle, where each of the attacking vehicles is driven by its respective second attack decision model.

[0029] Optionally, determining the decision vulnerability of the autonomous driving decision model from the accident scenario based on the second attack decision model includes:

[0030] Obtaining the entropy of multiple actions output by the attacking vehicle in the accident scenario;

[0031] Based on multiple of the entropies, determining the state action of the attacking vehicle corresponding to the minimum entropy;

[0032] Based on the state action, determining the decision vulnerability of the autonomous driving decision model.

[0033] An embodiment of the present application also provides a decision vulnerability detection device for an autonomous driving decision model, including:

[0034] A construction module configured to construct a detection scenario for detecting an autonomous driving decision model of a target vehicle, where the detection scenario includes at least one of the following: road environment, the autonomous driving decision model, one or more attacking vehicles for attacking the target vehicle, and a first attack decision model possessed by the attacking vehicle;

[0035] A training module configured to train the first attack decision model based on a multi-agent learning algorithm using a behavior reward constructed for the attacking vehicle to obtain a corresponding second attack decision model, where the behavior reward includes at least one of the following: accident liability arbitration reward and diverse accident scenario reward, the accident liability arbitration reward determines the accident liability according to traffic rules through vehicle driving information during the accident process between the target vehicle and the attacking vehicle, and the diverse accident scenario reward is determined according to the dissimilarity between accidents;

[0036] An interaction module configured to, based on the detection scenario, interact the target vehicle with the autonomous driving decision model and the attacking vehicle with the second attack decision model to generate an accident scenario caused by the decision vulnerability of the autonomous driving decision model;

[0037] A determination module configured to determine the decision vulnerability of the autonomous driving decision model from the accident scenario based on the second attack decision model.

[0038] The decision vulnerability detection method of the autonomous driving decision-making model according to the embodiments of the present application can efficiently explore accident scenarios that are more complex than single-agent reinforcement learning through the multi-agent reinforcement learning method, distinguish accident responsibilities through accident liability arbitration rewards, and enable intelligent attacking vehicles to generate diverse accident scenarios caused by vulnerabilities in the autonomous driving decision-making model through diverse accident scenario rewards, thereby accurately identifying the vulnerabilities. On the one hand, this solves the problems of high cost, low efficiency, and incomplete detection in real road tests. On the other hand, it also avoids the problems of weak accident transferability, single scenarios, and lack of liability determination in existing methods. Description of the Drawings

[0039] Figure 1 It is a flowchart of the decision vulnerability detection method according to the embodiments of the present application;

[0040] Figure 2 For the embodiments of the present application Figure 1 It is a flowchart of an embodiment of step S100;

[0041] Figure 3 For the embodiments of the present application Figure 1 It is a flowchart of an embodiment of step S200;

[0042] Figure 4 For the embodiments of the present application Figure 1 It is a flowchart of an embodiment of step S400;

[0043] Figure 5 It is a schematic diagram of the process of rewarding the attacking vehicle according to the embodiments of the present application;

[0044] Figure 6 It is a schematic diagram of the process of the attacking vehicle adjusting its own actions according to the reward according to the embodiments of the present application;

[0045] Figure 7 It is a schematic diagram of the action entropy of the attacking vehicle in the accident scenario of the autonomous driving decision-making model according to the embodiments of the present application.

[0046] Figure 8 It is a structural block diagram of the decision vulnerability detection device according to the embodiments of the present application. Detailed Embodiments

[0047] Reference is made herein to the drawings to describe various solutions and features of the present application.

[0048] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be construed as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.

[0049] The drawings included in and forming a part of the specification illustrate embodiments of the present application, and together with the general description of the present application given above and the detailed description of the embodiments given below are used to explain the principles of the present application.

[0050] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given as non - limiting examples with reference to the drawings.

[0051] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.

[0052] When combined with the drawings, the above - mentioned and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description.

[0053] Specific embodiments of the present application will be described hereinafter with reference to the drawings; however, it should be understood that the embodiments claimed are merely examples of the present application, which can be implemented in various ways. Well - known and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but are merely used as a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.

[0054] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments", each of which may refer to one or more of the same or different embodiments according to the present application.

[0055] A method for detecting decision - making loopholes in an autonomous driving decision - making model according to an embodiment of the present application, the method includes: building a detection scenario of the autonomous driving decision - making model in a simulator, where the detection scenario includes road structures, a target vehicle loaded with the autonomous driving decision - making model, an attacking vehicle, other environmental vehicles, etc. Based on the multi - agent reinforcement learning method, designing an accident liability arbitration reward and a diversity accident scenario reward, and training to obtain an attacking vehicle decision - making model for controlling the attacking vehicle. Among them, the attacking vehicle interacts with the target vehicle loaded with the autonomous driving decision - making model, generating multiple accident scenarios caused by the decision - making loopholes of the autonomous driving decision - making model, and then detecting the decision - making loopholes of the autonomous driving decision - making model from the accident scenarios.

[0056] The following describes this method in detail with reference to the drawings. Figure 1 is a flowchart of the decision - making loophole detection method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0057] S100. Construct a detection scenario for detecting an autonomous driving decision-making model for a target vehicle, where the detection scenario includes at least one of the following: road environment, the autonomous driving decision-making model, one or more attacking vehicles for attacking the target vehicle, and a first attack decision-making model of the attacking vehicle.

[0058] In one embodiment, the autonomous driving decision-making model is used to control the actions of the target vehicle. The autonomous driving decision-making model includes a well-designed and trained autonomous driving control strategy, and the target vehicle loaded with the autonomous driving decision-making model can achieve autonomous driving.

[0059] In this embodiment, a detection scenario for detecting an autonomous driving decision-making model for a target vehicle can be constructed in a simulator, thereby virtualizing the test scenario, without the need to actually construct it in reality, thus improving the detection efficiency and reducing the detection cost.

[0060] The detection scenario may include a road environment, which may include road elements such as lanes, lane boundaries, traffic lights, crosswalks, etc. The detection scenario also includes the autonomous driving decision-making model as the object of decision-making vulnerability detection. The autonomous driving decision-making model is constructed based on a deep neural network and is used to control the driving of the target vehicle, and is also the detection object in the detection scenario. The detection scenario also includes one or more attacking vehicles for attacking the target vehicle and a first attack decision-making model of the attacking vehicle. The first attack decision-making model is the initial decision-making model of the attacking vehicle, which can be constructed based on historical data and / or empirical data, or can be pre-constructed based on known relevant data. Among them, in this embodiment, the method will train the first attack decision-making model as the initial decision-making model to form a second attack decision-making model.

[0061] The attacking vehicle, as an agent, can drive under the control of the first attack decision-making model. The goal of the attacking vehicle is to interact with the target vehicle in the test scenario to cause an accident. When there are multiple attacking vehicles, they can coordinate with each other under the control of their respective first attack decision-making models to attack the target vehicle.

[0062] S200. Based on a multi-agent learning algorithm, use the behavior rewards constructed for the attacking vehicle to train the first attack decision-making model to obtain a corresponding second attack decision-making model, where the behavior rewards include at least one of the following: accident liability arbitration reward and diverse accident scenario reward. The accident liability arbitration reward determines the accident liability according to traffic rules through the vehicle driving information during the accident process between the target vehicle and the attacking vehicle. The diverse accident scenario reward is determined according to the dissimilarity between accidents.

[0063] In one embodiment, in combination with Figure 6, The multi-agent learning algorithm is an algorithm that learns through trial and error, and is also a method for solving the problem of how agents make decisions during their interaction with a Markov environment. It can efficiently explore accident scenarios that are more complex than single-agent reinforcement learning. Among them, the object that conducts learning or makes decisions is called an agent (corresponding to the attacking vehicle in this embodiment). And all other things that interact with the agent are called the environment. The agent selects actions according to the state of the environment and interacts with the environment. The environment responds to the agent's actions and transfers to a new state, while feeding back a certain reward to the agent. The agent continuously adjusts its actions according to the reward.

[0064] In one embodiment, the multi-agent learning algorithm can be modeled based on the Markov Decision Process (MDP), where an MDP consists of the following five parts:

[0065] First, the state space S, which represents the set of all possible states of the environment, and \(s\) t represents the state of the environment at time step \(t\).

[0066] Second, the action space A, which represents the set of all executable actions of the agent, and \(a\) t represents the action of the agent at time step \(t\).

[0067] Third, the state transition probability \(P(s'|s,a)\), which represents the probability that the environment state \(s\) transfers to state \(s'\) after the agent executes action \(a\).

[0068] Fourth, the reward function \(r(s,a,s')\), which represents the reward obtained by the agent after transferring from the environment state \(s\) by executing action \(a\) to state \(s'\), and \(r\) t represents the reward at time step \(t\).

[0069] Fifth, the discount factor \(\gamma\). The discount factor represents the discount coefficient of future rewards, and its value is usually taken as \((0,1)\).

[0070] The policy \(\pi\) of the agent is a function from the state space to the action space, that is, \(\pi:S\rightarrow A\), and \(\pi(a\) t |s t ) represents the probability of choosing action \(a\) t at state \(s\) t . The agent generates an MDP trajectory \(\tau\) through multiple interactions with the environment according to the policy \(\pi\) as follows:

[0071] s 0 ,a 0 ,r 1 ,s 1 ,a 1 ,r2 ,s 2 ,a 2 ,r 3 ,s 3 ,…

[0072] Given a strategy π, the cumulative reward received by the trajectory τ of an interaction process between the agent and the environment is the total reward G(τ), as shown in the following formula 1:

[0073]

[0074] The goal of the multi-agent learning algorithm is to find an optimal strategy π so that the expected total reward obtained by the agent is maximum.

[0075] Combination Figure 5 In this embodiment, based on the multi-agent learning algorithm, at least the initial first attack decision model of the attack vehicle is trained by using the accident responsibility arbitration reward and the diverse accident scene reward. During the training process, the attack vehicle and the target vehicle can be interactively operated, and then the behavior reward for the attack vehicle can be constructed according to the interactive operation, and the behavior reward at least includes the accident responsibility arbitration reward and the diverse accident scene reward. Among them, on the one hand, during the interaction between the attack vehicle and the target vehicle, according to the traffic rules, the accident responsibility is determined by the speed, acceleration, displacement, direction, etc. of the attack vehicle before and after the accident, and then the attack vehicle (agent) is given the accident responsibility arbitration reward. On the other hand, according to the difference between the current accident between the attack vehicle and the target vehicle and the known accident, the attack vehicle is given a diverse accident scene reward, wherein the difference is the degree of difference between different accidents, such as two accidents are completely different, indicating that the difference is large. In addition, the diverse accident scene reward includes positive rewards and negative rewards (penalties) for the attack vehicle. Thus, the attack vehicle can be given a more flexible reward for diverse accident scenes.

[0076] After training the first attack decision model, a more complete second attack decision model is obtained, which will enable the attacking vehicle to use the second attack decision model to make decisions on attack behaviors, and thus interact with the target vehicle more intelligently in the detection scenario.

[0077] S300, based on the detection scenario, the target vehicle having the autonomous driving decision model interacts with the attack vehicle having the second attack decision model to generate an accident scenario caused by a decision vulnerability of the autonomous driving decision model.

[0078] Exemplarily, under the influence of road elements such as lanes, lane boundaries, traffic lights, and crosswalks in the detection scenario, the target vehicle is driven based on the autonomous driving decision model, and the attacking vehicle is driven based on the second vehicle decision model, so that the attacking vehicle has an impact on the driving process of the target vehicle, or an accident occurs. Then, an accident scenario is generated based on the interaction result, and this accident scenario is caused by the decision-making vulnerability of the autonomous driving decision model. For example, when the target vehicle is driving based on the autonomous driving decision model, it fails to avoid the attack of the attacking vehicle in time, resulting in an accident.

[0079] The following is illustrated by way of multiple specific embodiments: (1), when the target vehicle Apollo changes lanes to the right, the attacking vehicle in the right front suddenly decelerates. The target vehicle Apollo fails to change its driving direction in time and collides with the attacking vehicle on the side.

[0080] (2), the attacking vehicle in the middle lane pretends to change lanes to the right, misleading the target vehicle Apollo to change from the left lane to the middle lane and collide with it.

[0081] (3), the attacking vehicle on the left side of the target vehicle Apollo continuously changes lanes to the right side of the target vehicle Apollo, and then changes lanes in front of the target vehicle Apollo. The target vehicle Apollo cannot accurately predict its trajectory and collides with the vehicle.

[0082] (4), the attacking vehicle in front suddenly brakes after the target vehicle Apollo completes changing lanes to the right. The target vehicle Apollo fails to decelerate in time and collides with the attacking vehicle in front.

[0083] (5), the target vehicle Apollo first continuously changes lanes from the left lane to the right lane, and then changes back to the middle lane. However, the target vehicle Apollo steers too early and collides with the attacking vehicle in the middle lane on the side.

[0084] (6), the attacking vehicle inserts from the left lane, and the target vehicle Apollo does not have enough distance to decelerate and rear-ends it.

[0085] (7), an attacking vehicle follows the target vehicle Apollo from the right, preventing the target vehicle Apollo from changing lanes to the right. Another attacking vehicle suddenly inserts from the left, and the target vehicle Apollo does not have enough distance to decelerate and collides with it.

[0086] (8), two attacking vehicles drive side by side in front of the target vehicle Apollo. The target vehicle Apollo misjudges the distance between the two attacking vehicles and believes it is sufficient to overtake. An accident occurs when the target vehicle Apollo accelerates and tries to pass between the two attacking vehicles.

[0087] (9), Two attacking vehicles collided in front of the target vehicle Apollo. The target vehicle Apollo did not have enough distance to brake and a rear-end collision occurred.

[0088] (10), One attacking vehicle followed the target vehicle Apollo from the right, preventing the target vehicle Apollo from changing lanes to the right. Another attacking vehicle suddenly braked. The target vehicle Apollo did not have enough distance to decelerate and collided with it.

[0089] (11), The target vehicle Apollo changed lanes to the right. The attacking vehicle on the left side of the target vehicle Apollo continuously changed lanes in front of the target vehicle Apollo. Another attacking vehicle approached the target vehicle Apollo from behind. The target vehicle Apollo did not have enough distance to decelerate, resulting in its rear-end collision with the vehicle in front.

[0090] (12), The target vehicle Apollo changed from the middle lane to the right lane. The attacking vehicle in the left lane changed lanes to the right and drove on the lane line, causing the target vehicle Apollo to misjudge that there was enough space to return to the middle lane. A side collision occurred.

[0091] During the above-mentioned various different types of interactions between the target vehicle and the attacking vehicles, corresponding accident scenarios are respectively generated. The accidents in these accident scenarios can be considered as accidents caused by vulnerabilities in the autonomous driving decision-making model of the target vehicle.

[0092] S400, Based on the second attack decision-making model, determine the decision-making vulnerabilities of the autonomous driving decision-making model from the accident scenarios.

[0093] In one embodiment, since the accidents in the accident scenarios are accidents caused by vulnerabilities in the autonomous driving decision-making model of the target vehicle. Therefore, the entropy of the actions input by the attacking vehicles in the accident scenarios recorded in the second attack decision-making model can be analyzed, the state actions corresponding to the smaller entropy values are selected, and the decision-making vulnerabilities of the autonomous driving decision-making model are determined based on these state actions.

[0094] In a specific embodiment, the action probabilities of the attacking vehicle in different states during an accident trajectory can be obtained from the accident scenario, and the action entropy can be calculated. Based on its magnitude, it can be determined whether the current state-action pair is a critical state-action pair. Since the attack strategy has learned how to generate accident scenarios, the larger the entropy, the more random the actions of the attacker, and the behavior of the attacking vehicle has no tendency, indicating that the actions of the attacking vehicle in the current state have little impact on generating accident scenarios. On the contrary, the smaller the entropy, the more inclined the attacking vehicle is to perform a certain action, which means that the attacking vehicle should perform a certain action in this state to generate an accident scenario, that is, this state-action is a critical state-action pair. Therefore, based on the state-action of the attacking vehicle corresponding to the minimum entropy, the decision-making loopholes of the autonomous driving decision-making model can be determined more accurately and realistically. Of course, the determination of such decision-making loopholes can also be solved by other methods. Accurately determining the decision-making loopholes of the autonomous driving decision-making model can help improve the robustness of the autonomous driving decision-making model.

[0095] The method for detecting decision-making loopholes of the autonomous driving decision-making model in the embodiments of the present application, through the multi-agent reinforcement learning method, efficiently explores accident scenarios that are more complex than single-agent reinforcement learning, distinguishes accident responsibilities through accident liability arbitration rewards, and enables intelligent attacking vehicles to generate diverse accident scenarios caused by loopholes in the autonomous driving decision-making model through diverse accident scenario rewards, thereby more accurately mining such loopholes. On the one hand, this solves the problems of high cost, low efficiency, and incomplete detection in real road tests, and on the other hand, it also avoids the problems of weak accident transferability, single scenario, and lack of liability determination in existing methods.

[0096] In an embodiment of the present application, the construction of a detection scenario for detecting the autonomous driving decision-making model of a target vehicle is as Figure 2 shown, including:

[0097] S110, constructing the road environment in the simulator;

[0098] S120, determining the initial position and destination of the target vehicle, and the initial position of the attacking vehicle.

[0099] Exemplarily, the simulator can be constructed based on software and hardware. It is a substitute for the target system, can complete external functions, and can present behaviors and goals consistent with the target system to the outside.

[0100] The simulator in this embodiment can simulate the occurrence process of accidents. This includes constructing a road environment in the simulator, such as constructing road elements such as lanes, lane boundaries, traffic lights, and crosswalks for the attacking vehicle and the target vehicle to drive and interact. Thus, the target vehicle and the attacking vehicle are set in the constructed road environment.

[0101] In addition, the initial positions of the target vehicle and multiple intelligent attacking vehicles are respectively set in the emulator, and a destination is set for the target vehicle of autonomous driving, so that the target vehicle has a certain movement route, while other environmental vehicles travel along a fixed route.

[0102] In an embodiment of the present application, the input of the first attack decision model is the state information of the attacking vehicle and the relative position relationship between the attacking vehicle and the target vehicle; the output of the first attack decision model is the decision-making action of the attacking vehicle.

[0103] Exemplarily, the first attack decision model is the training target. During the training process, the state information of the attacking vehicle and the relative position relationship between the attacking vehicle and the target vehicle can be input into the first attack decision model. Among them, the state information of the attacking vehicle includes information such as the position, speed, and acceleration of the attacking vehicle. And the relative position relationship between the attacking vehicle and the target vehicle can include the relative distance and relative speed between the attacking vehicle and the target vehicle. This will enable the first attack decision model to control the driving of the attacking vehicle based on the input information, making it easier to interact with the target vehicle. In addition, the output of the first attack decision model is the decision-making action of the attacking vehicle, and this decision-making action can be the current action of the attacking vehicle. Then, the attacking vehicle interacts with the target vehicle based on the decision-making action, including having an accident.

[0104] In an embodiment of the present application, the behavior reward further includes a distance reward, and the distance reward is used to reward the attacking vehicle according to the distance between the target vehicle and the attacking vehicle; training the first attack decision model by using the behavior reward constructed for the attacking vehicle, as Figure 3 shown, includes:

[0105] S210, determining the initial decision-making action of the attacking vehicle based on the state of the detection scenario;

[0106] S220, adjusting the state of the detection scenario based on the initial decision-making action of the attacking vehicle to generate the corresponding behavior reward;

[0107] S230, rewarding the attacking vehicle based on the accident liability arbitration reward, the diversity accident scenario reward, and / or the distance reward.

[0108] Exemplarily, the target vehicle is controlled by a well-trained or designed autonomous driving decision-making model and deployed in a simulated detection scenario, and is regarded as part of the environment together with lanes, traffic lights, etc. Multiple attacking vehicles (multi-agent) actively explore accident scenarios (such as car accidents) caused by incorrect decisions of the tested autonomous driving decision-making model. In one embodiment, environment vehicles operating based on rules are added to the detection scenario simultaneously to better simulate real-world driving scenarios.

[0109] During the training process of the first attack decision-making model, multiple attacking vehicles (multi-agent) obtain observation information o such as the position and speed of the target vehicle according to the state s of the detection scenario t and output an initial decision-making action a according to the observation information t , and interact with the detected target vehicle. The detection scenario (environment) is adjusted to a new state s according to the transition probability p(s t ∣s t+1 ,a t ). The detection scenario (environment) transmits the adjusted observation o t and the behavior reward r t+1 to the multi-agent. t+1 and the behavior reward r t to the multi-agent.

[0110] The behavior reward includes accident liability arbitration reward, diverse accident scenario reward, and distance reward. For example, speed reward for the multi-agent to drive normally, collision penalty, liability arbitration reward, and scenario difference reward. Compared with the interaction between a single agent and the target vehicle, the multi-agent learning algorithm is more efficient, and through cooperation among agents, more complex accident scenarios can be discovered. The liability arbitration reward is related to traffic rules and decoupled from specific scenarios, and can discover reasonable accident scenarios. The diverse accident scenario reward (scenario difference reward) encourages agents to discover diverse accident scenarios by comparing new scenarios with historical accident scenarios. The distance reward is used to reward the attacking vehicle according to the distance between the target vehicle and the attacking vehicle, and when the distance is within a suitable range, the distance reward for the attacking vehicle is larger.

[0111] In one embodiment of the present application, the distance reward includes negative value penalty and positive value reward. Training the first attack decision-making model using the behavior reward constructed for the attacking vehicle includes the following steps:

[0112] When the distance between the attacking vehicle and the target vehicle exceeds a first predetermined distance, give the attacking vehicle the negative value penalty to encourage the attacking vehicle to approach the target vehicle;

[0113] When the attacking vehicle gradually approaches the target vehicle, the attacking vehicle will receive an increasingly larger positive reward;

[0114] When the attacking vehicle approaches the target vehicle and the vehicle distance is less than a second predetermined distance, a negative penalty is given to the attacking vehicle.

[0115] Exemplarily, the distance reward includes a negative penalty and a positive reward. When the distance between the attacking vehicle and the target vehicle is within an appropriate range, a positive reward can be given to the attacking vehicle; when the distance between the attacking vehicle and the target vehicle exceeds the appropriate range, a negative penalty can be given to the attacking vehicle.

[0116] Specifically, when the attacking vehicle is far from the detected target vehicle and exceeds a first predetermined distance, the attacking vehicle will receive a negative penalty to encourage the attacking vehicle to approach the detected target vehicle, which is prone to accidents; when the attacking vehicle gradually approaches the detected target vehicle, the attacking vehicle will receive an increasingly larger positive reward, thereby increasing the reward effect; when the attacking vehicle is too close to the detected vehicle and less than a second predetermined distance, a negative penalty will be received to prevent the attacking vehicle from over-exploring accident scenarios such as directly colliding with the target vehicle, which are not the responsibility of the detected vehicle or are difficult to avoid.

[0117] In an embodiment of the present application, the accident liability arbitration reward includes a collision reward and an aggressiveness penalty. Training the first attack decision model by using the behavior reward constructed for the attacking vehicle includes:

[0118] When it is determined that the attacking vehicle collides with the target vehicle, or an environmental vehicle in the detection scenario collides with the target vehicle, the attacking vehicle is given the collision reward;

[0119] When the attacking vehicle violates the driving regulation restrictions, the attacking vehicle is given the aggressiveness penalty to standardize the driving actions.

[0120] Exemplarily, vulnerabilities in the autonomous driving decision-making model can cause the target vehicle to take inappropriate actions (e.g., unsafe lane-changing behavior or risky overtaking behavior), resulting in accidents. However, since some accident scenarios may not be caused by vulnerabilities in the detected autonomous driving decision-making model, if the accident is inevitable or not caused by the detected autonomous driving decision-making model, it is of no value for improving the autonomous driving decision-making model. Therefore, in this embodiment, the vulnerability mining process pays more attention to accident scenarios caused by errors in the autonomous driving decision-making model. Since the multi-agent learning algorithm is a reward-driven algorithm, in order to distinguish between dangerous scenarios for which the autonomous driving decision-making model is responsible and those for which it is not, this application sets an accident liability arbitration reward, which determines the liability of dangerous scenarios based on common sense of human judgment, and punishes scenarios for which the attacking vehicle is responsible rather than the autonomous driving decision-making model. The accident liability arbitration reward adjusts the distribution of accident liability, and thus can regulate the behavior of the attacking vehicle. The accident liability arbitration reward includes: a collision reward r c and an aggressiveness penalty r agg . When the detected target vehicle collides with an environmental vehicle or with an attacking vehicle, the attacking vehicle will receive a sparse collision reward r c . And when the attacking vehicle violates driving regulation limits, such as violating acceleration limits and lane-changing behavior limits, it will receive an aggressiveness penalty r agg to regulate its own actions. The accident liability arbitration reward r HAR can be expressed as shown in the following equations (2) and (3):

[0121]

[0122]

[0123] r HAR = r c + r agg

[0124] Equation (3)

[0125] where v ˙ att is the acceleration of the attacking vehicle, δ is the acceleration threshold, a att represents the action of the attacking vehicle, and a lc and a rc are the left lane-changing and right lane-changing actions respectively.

[0126] The accident liability arbitration reward is not limited to specific dangerous scenarios, so new dangerous scenarios can be mined. By punishing the aggressive actions of the attack strategy, it can prevent the attack strategy from converging to trivial accident scenarios such as direct collisions, and thus can discover valuable accident scenarios for which the target vehicle of autonomous driving is responsible, that is, vulnerabilities in the tested autonomous driving decision-making model.

[0127] In one embodiment of the present application, training the first attack decision model by using the behavior reward constructed for the attack vehicle includes:

[0128] When it is determined that the attack vehicle explores a second dangerous scenario different from the first dangerous scenario that has been discovered, giving the diversity accident scenario reward to the attack vehicle.

[0129] Exemplarily, due to the characteristic that the reinforcement learning algorithm is prone to overfitting itself, the attack strategy of the attack vehicle often converges to a specific type of scenario, although there are still other vulnerabilities in the autonomous driving decision model of the target vehicle being tested at this time. Therefore, the present application sets the diversity accident scenario reward to encourage the attack vehicle to discover as many second dangerous scenarios different from the first dangerous scenario that has been discovered as possible.

[0130] Specifically, the set S represents all the discovered dangerous scenarios. The i-th frame in an accident scenario trajectory is defined as [x i1 , y i1 , x i2 , y i2 ,..., x in , y in , where n is the number of attack vehicles, and (x j , y j ) is the position of the j-th attack vehicle relative to the target vehicle. The scenario s i in the set S is the last frame in the accident trajectory. The dissimilarity d(s 1 , s 2 ) between two scenarios s 1 and s 2 is defined as the l 2 norm. Considering that there is no interaction between the attack vehicle and the target vehicle when the relative distance is too large, the present application uses a logical function to map the relative position to the interval (0, 1) before calculating the dissimilarity, and the formula is as shown in Equation 4 below:

[0131]

[0132] Therefore, the complete definition of the dissimilarity d(s 1 , s 2 ) is as shown in Equation 5 below:

[0133]

[0134] Then, the scenario difference reward r SDIR corresponding to the scenario s′ is defined as shown in Equation 6 below:

[0135]

[0136]

[0137] The diversity accident scenario reward is the minimum dissimilarity between the second dangerous scenario and all the first dangerous scenarios that have been discovered. Therefore, it encourages the attacking vehicle to explore new second dangerous scenarios that are different from all the first dangerous scenarios that have been discovered.

[0138] In an embodiment of the present application, when there are multiple attacking vehicles, the interaction between the target vehicle with the autonomous driving decision model and the attacking vehicle with the second attack decision model includes:

[0139] In the detection scenario, the target vehicle is attacked through the coordination of multiple attacking vehicles, where each attacking vehicle is driven by its respective second attack decision model.

[0140] Exemplarily, in the detection scenario, multiple attacking vehicles can be set up to form a multi-agent to attack the target vehicle, so as to more accurately simulate the real autonomous driving scenario. Specifically, for multiple attacking vehicles, each attacking vehicle is driven by its respective second attack decision model, so that each attacking vehicle can independently or coordinately attack the target vehicle, such as the various attack methods described above, thus achieving a higher authenticity interaction effect with the target vehicle.

[0141] In an embodiment of the present application, based on the second attack decision model, the decision loopholes of the autonomous driving decision model are determined from the accident scenario, as Figure 4 shown, including:

[0142] S410, obtaining the entropy of multiple actions output by the attacking vehicle in the accident scenario;

[0143] S420, based on multiple entropies, determining the state action of the attacking vehicle corresponding to the minimum entropy;

[0144] S430, based on the state action, determining the decision loopholes of the autonomous driving decision model.

[0145] Exemplarily, entropy is a measure of the uncertainty of a random variable. In this embodiment, the decision loopholes of the autonomous driving decision model can be determined based on the entropy of multiple actions output by the attacking vehicle.

[0146] Specifically, obtain the action probabilities of the attacking vehicle in different states during an accident trajectory from the accident scenario, calculate the action entropy, and determine whether the current state-action pair is a critical state-action pair based on its magnitude. Since the second attack decision model has learned how to generate accident scenarios, the larger the entropy, the more random the attacker's actions, and the behavior of the attacking vehicle has no tendency, indicating that the actions of the attacking vehicle have little impact on generating the accident scenario in the current state. On the contrary, the smaller the entropy, the more inclined the attacking vehicle is to perform a certain action, which means that the attacking vehicle should perform a certain action in this state to generate the accident scenario, that is, this state-action is a critical state-action pair. Therefore, based on the state-action of the attacking vehicle corresponding to the minimum entropy, the decision vulnerability of the autonomous driving decision model can be determined more accurately and realistically.

[0147] For example, as Figure 7 shown, it shows the action entropy of the attacking vehicle in the accident scenario. In the trajectory at the beginning stage, the action entropy output by the second attack decision model is very large, and all vehicles are driving normally along the road. When the target vehicle of the autonomous driving reaches about 15 meters behind the attacking vehicle above in the figure, the action entropy output by the attacking vehicle becomes smaller, and it performs the action of changing lanes to the right with a probability close to 1. At this time, the action entropy of the attacking vehicle is relatively large, and it takes the action of driving normally in the lane. After the attacking vehicle above in the figure changes lanes in front of the target vehicle, its action entropy becomes larger, and it drives normally in the lane. After the attacking vehicle above in the figure completes the lane-changing action, the action entropy of the attacking vehicle below in the figure becomes smaller, and it starts to decelerate with a probability close to 1 and approaches the target vehicle in a serpentine motion trajectory. Finally, in this case, the target vehicle takes an inappropriate action of changing lanes to the right and collides with the attacking vehicle below in the figure, resulting in an accident scenario for which the autonomous driving decision model is responsible. From the analysis of the action entropy of the trajectory of this scenario, two critical state-action pairs with a time sequence can be obtained: the first critical state action is that when the attacking vehicle is 15 meters in front of the left front of the target vehicle, it changes lanes from the left to the front of the same lane as the target vehicle; the other critical state action is that when this attacking vehicle completes the above lane-changing action, another attacking vehicle approaches the target vehicle from the right. And the critical state actions extracted by the action entropy are consistent with human intuition.

[0148] Based on the same inventive concept, an embodiment of the present application also provides a decision vulnerability detection device for an autonomous driving decision model, as Figure 8 shown, including:

[0149] A construction module configured to construct a detection scenario for detecting the autonomous driving decision model of the target vehicle, where the detection scenario includes at least one of the following: road environment, the autonomous driving decision model, one or more attacking vehicles for attacking the target vehicle, and the first attack decision model possessed by the attacking vehicle.

[0150] In one embodiment, an autonomous driving decision-making model is used to control the actions of a target vehicle. The autonomous driving decision-making model includes a well-designed and trained autonomous driving control strategy. The target vehicle loaded with this autonomous driving decision-making model can achieve autonomous driving.

[0151] In this embodiment, the construction module can construct a detection scenario for detecting the autonomous driving decision-making model of the target vehicle in a simulator, thereby virtualizing the test scenario, without the need to physically construct it in reality, thus improving the detection efficiency and reducing the detection cost.

[0152] The detection scenario may include a road environment, which may include road elements such as lanes, lane boundaries, traffic lights, crosswalks, etc. The detection scenario also includes the autonomous driving decision-making model that is the object of decision-making vulnerability detection. The autonomous driving decision-making model is constructed based on a deep neural network and is used to control the driving of the target vehicle, and is also the detection object in the detection scenario. The detection scenario also includes one or more attack vehicles for attacking the target vehicle and a first attack decision-making model possessed by the attack vehicle. The first attack decision-making model is the initial model of the attack vehicle, which can be constructed based on historical data and / or empirical data, or can be pre-constructed based on known relevant data.

[0153] The attack vehicle, as an agent, can drive under the control of the first attack decision-making model. The goal of the attack vehicle is to interact with the target vehicle in the test scenario to cause an accident. When there are multiple attack vehicles, they can coordinate with each other under the control of their respective first attack vehicle decision-making models to attack the target vehicle.

[0154] A training module, configured to train the first attack decision-making model based on a multi-agent learning algorithm using the behavior rewards constructed for the attack vehicle to obtain a corresponding second attack decision-making model, where the behavior rewards include at least one of the following: accident liability arbitration rewards and diverse accident scenario rewards. The accident liability arbitration rewards are used to determine the accident liability according to traffic rules through the vehicle driving information during the accident process between the target vehicle and the attack vehicle. The diverse accident scenario rewards are determined according to the dissimilarity between accidents.

[0155] In one embodiment, the multi-agent learning algorithm is an algorithm that learns through trial and error, and is also a method for solving the problem of how an agent makes decisions during interaction with a Markov environment. It can efficiently explore accident scenarios that are more complex than single-agent reinforcement learning. Among them, the object that conducts learning or makes decisions is called an agent (corresponding to the attacking vehicle in this embodiment). And all other things that interact with the agent are called the environment. The agent selects actions according to the state of the environment and interacts with the environment. The environment responds to the agent's actions and transfers to a new state, while feeding back a certain reward to the agent. The agent continuously adjusts its actions according to the reward.

[0156] In one embodiment, the multi-agent learning algorithm can be modeled based on the Markov Decision Process (MDP), where an MDP consists of the following five parts:

[0157] First, the state space S, which represents the set of all possible states of the environment, and s t represents the state of the environment at time step t.

[0158] Second, the action space A, which represents the set of all executable actions of the agent, and a t represents the action of the agent at time step r.

[0159] Third, the state transition probability P(s′|s,a), which represents the probability that the environment state s transfers to state s′ after the agent executes action a.

[0160] Fourth, the reward function r(s,a,s′), which represents the reward obtained by the agent after transferring from the environment state s by executing action a to state s′, and r t represents the reward at time step t.

[0161] Fifth, the discount factor γ. The discount factor represents the discount coefficient of future rewards, and its value is usually taken as (0,1).

[0162] The policy π of the agent is a function from the state space to the action space, that is, π∶S→A, and π(a t |s t ) represents the probability of selecting action a t at state s t . The agent generates an MDP trajectory τ through multiple interactions with the environment according to the policy π as follows:

[0163] s 0 ,a 0 ,r 1 ,s 1 ,a 1 ,r2 ,s 2 ,a 2 ,w 3 ,s 3 ,…

[0164] The cumulative reward received by the trajectory τ of one interaction process between the given policy π agent and the environment is the total return G(τ), as shown in Equation 7 below:

[0165]

[0166] The goal of the multi-agent learning algorithm is to find an optimal policy π such that the expected value of the total return obtained by the agent is maximized.

[0167] In this embodiment, the training module trains the initial first attack vehicle decision model of the attack vehicle based on the multi-agent learning algorithm, at least using the accident liability arbitration reward and the diversity accident scenario reward. During the training process, the attack vehicle can be interacted with the target vehicle, and then a behavior reward for the attack vehicle is constructed according to the interaction operation. The behavior reward at least includes the accident liability arbitration reward and the diversity accident scenario reward. Among them, during the interaction process between the attack vehicle and the target vehicle, on the one hand, according to the traffic rules, the accident liability is judged based on the speed, acceleration, displacement, orientation, etc. of the attack vehicle before and after the accident, and then the accident liability arbitration reward is given to the attack vehicle (agent). On the other hand, according to the accident currently occurring between the attack vehicle and the target vehicle, the diversity accident scenario reward is given to the attack vehicle based on the dissimilarity from the known accidents. Among them, the dissimilarity is the degree of difference between different accidents. If two accidents are completely different, it indicates a large dissimilarity. In addition, the diversity accident scenario reward includes a positive reward and a negative reward (penalty) for the attack vehicle. Thus, the diversity accident scenario reward can be given to the attack vehicle more flexibly.

[0168] After the training module trains the first attack decision model, a more perfect second attack decision model is obtained. This will enable the attack vehicle to use the second attack decision model to make decisions on attack behaviors, and then interact with the target vehicle more intelligently in the detection scenario.

[0169] The interaction module is configured to, based on the detection scenario, interact the target vehicle with the autonomous driving decision model and the attack vehicle with the second attack decision model to generate an accident scenario caused by the decision vulnerability of the autonomous driving decision model.

[0170] Exemplarily, under the influence of road elements such as lanes, lane boundaries, traffic lights, and crosswalks in the detection scenario, the interaction module drives the target vehicle based on the autonomous driving decision model and drives the attacking vehicle based on the second attack decision model, so that the attacking vehicle has an impact on the driving actions of the target vehicle during the driving process, or an accident occurs. Then, an accident scenario is generated based on the interaction result, and this accident scenario is caused by the decision-making loopholes of the autonomous driving decision model. For example, when the target vehicle is driving based on the autonomous driving decision model, it fails to avoid the attack of the attacking vehicle in time, resulting in an accident.

[0171] The following uses multiple specific embodiments to illustrate the occurring accidents. (1), when the target vehicle Apollo changes lanes to the right, the attacking vehicle in the right front suddenly decelerates. The target vehicle Apollo fails to change the driving direction in time and collides with the attacking vehicle sidewise.

[0172] (2), the attacking vehicle in the middle lane pretends to change lanes to the right, misleading the target vehicle Apollo to change from the left lane to the middle lane and collide with it.

[0173] (3), the attacking vehicle on the left side of the target vehicle Apollo continuously changes lanes to the right side of the target vehicle Apollo, and then changes lanes in front of the target vehicle Apollo. The target vehicle Apollo cannot accurately predict its trajectory and collides with the vehicle.

[0174] (4), the attacking vehicle in front suddenly brakes after the target vehicle Apollo completes changing lanes to the right. The target vehicle Apollo fails to decelerate in time and collides with the attacking vehicle in front.

[0175] (5), the target vehicle Apollo first continuously changes lanes from the left lane to the right lane, and then changes lanes back to the middle lane. However, the target vehicle Apollo turns too early and collides with the attacking vehicle in the middle lane sidewise.

[0176] (6), the attacking vehicle inserts from the left lane, and the target vehicle Apollo does not have enough distance to decelerate and rear-ends it.

[0177] (7), an attacking vehicle follows the target vehicle Apollo from the right, preventing the target vehicle Apollo from changing lanes to the right. Another attacking vehicle suddenly inserts from the left, and the target vehicle Apollo does not have enough distance to decelerate and collides with it.

[0178] (8), two attacking vehicles drive side by side in front of the target vehicle Apollo. The target vehicle Apollo misjudges the distance between the two attacking vehicles and believes it is sufficient to overtake. An accident occurs when the target vehicle Apollo accelerates and attempts to pass between the two attacking vehicles.

[0179] (9), Two attacking vehicles collided in front of the target vehicle Apollo. The target vehicle Apollo did not have enough distance to brake and a rear-end collision occurred.

[0180] (10), One attacking vehicle followed the target vehicle Apollo from the right, preventing the target vehicle Apollo from changing lanes to the right. Another attacking vehicle suddenly braked. The target vehicle Apollo did not have enough distance to decelerate and collided with it.

[0181] (11), The target vehicle Apollo changed lanes to the right. The attacking vehicle on the left side of the target vehicle Apollo continuously changed lanes in front of the target vehicle Apollo. Another attacking vehicle approached the target vehicle Apollo from behind. The target vehicle Apollo did not have enough distance to decelerate, resulting in it rear-ending the vehicle in front.

[0182] (12), The target vehicle Apollo changed from the middle lane to the right lane. The attacking vehicle in the left lane changed lanes to the right and drove on the lane line, causing the target vehicle Apollo to misjudge that there was enough space to return to the middle lane. A side collision occurred.

[0183] During the above-mentioned various types of interactions between the target vehicle and the attacking vehicles, the interaction module respectively generates corresponding accident scenarios, and the accidents in these accident scenarios can be considered as accidents caused by vulnerabilities in the autonomous driving decision-making model of the target vehicle.

[0184] A determination module, configured to determine a decision-making vulnerability of the autonomous driving decision-making model from the accident scenarios based on the second attack decision-making model.

[0185] In one embodiment, since the accidents in the accident scenarios are accidents caused by vulnerabilities in the autonomous driving decision-making model of the target vehicle. Therefore, the determination module can analyze the entropy of the actions input by the attacking vehicles in the accident scenarios recorded in the second attack decision-making model, select the state actions corresponding to the smaller entropy values, and determine the decision-making vulnerabilities of the autonomous driving decision-making model based on these state actions.

[0186] In a specific embodiment, the action probabilities of the attacking vehicle in different states along a section of the accident trajectory can be obtained from the accident scenario, and the action entropy can be calculated. Based on its magnitude, it can be determined whether the current state-action pair is a critical state-action pair. Since the attack strategy has learned how to generate accident scenarios, the greater the entropy, the more random the actions of the attacker, and the behavior of the attacking vehicle has no tendency, indicating that the actions of the attacking vehicle in the current state have little impact on generating accident scenarios. On the contrary, the smaller the entropy, the more inclined the attacking vehicle is to perform a certain action, which means that the attacking vehicle should perform a certain action in this state to generate an accident scenario, that is, this state-action is a critical state-action pair. Therefore, the state-action corresponding to the smaller entropy value can be selected, and based on this state-action, the decision-making vulnerability of the autonomous driving decision-making model can be determined. Of course, the determination of this decision-making vulnerability can also be solved by other methods. Accurately determining the decision-making vulnerability of the autonomous driving decision-making model can help improve the robustness of the autonomous driving decision-making model.

[0187] In an embodiment of the present application, the construction module is further configured to:

[0188] Construct the road environment in the simulator;

[0189] Determine the initial position and destination of the target vehicle, as well as the initial position of the attacking vehicle.

[0190] In an embodiment of the present application, the input of the first attack decision-making model is the state information of the attacking vehicle and the relative position relationship between the attacking vehicle and the target vehicle; the output of the first attack decision-making model is the decision-making action of the attacking vehicle.

[0191] In an embodiment of the present application, the behavior reward further includes a distance reward, and the distance reward is used to reward the attacking vehicle according to the distance between the target vehicle and the attacking vehicle; the training module is further configured to:

[0192] Based on the state of the detection scenario, determine the initial decision-making action of the attacking vehicle;

[0193] Based on the initial decision-making action of the attacking vehicle, adjust the state of the detection scenario to generate the corresponding behavior reward;

[0194] Reward the attacking vehicle based on the accident liability arbitration reward, the diversity accident scenario reward, and / or the distance reward.

[0195] In an embodiment of the present application, the distance reward includes negative value punishment and positive value reward, and the training module is further configured to:

[0196] When the distance between the attacking vehicle and the target vehicle exceeds a first predetermined distance, a negative penalty is imposed on the attacking vehicle to encourage the attacking vehicle to approach the target vehicle;

[0197] When the attacking vehicle gradually approaches the target vehicle, the attacking vehicle will receive an increasingly larger positive reward;

[0198] When the attacking vehicle approaches the target vehicle and the distance is less than a second predetermined distance, a negative penalty is imposed on the attacking vehicle.

[0199] In an embodiment of the present application, the accident liability arbitration reward includes a collision reward and an aggressiveness penalty, and the training module is further configured to:

[0200] When it is determined that the attacking vehicle collides with the target vehicle, or an environmental vehicle in the detection scenario collides with the target vehicle, the attacking vehicle is given the collision reward;

[0201] When the attacking vehicle violates the driving regulation restrictions, the attacking vehicle is given the aggressiveness penalty to standardize the driving actions.

[0202] In an embodiment of the present application, the training module is further configured to:

[0203] When it is determined that the attacking vehicle explores a second dangerous scenario different from the first dangerous scenario that has been discovered, the attacking vehicle is given the diversity accident scenario reward.

[0204] In an embodiment of the present application, when there are multiple attacking vehicles, the interaction module is further configured to:

[0205] In the detection scenario, the target vehicle is attacked by multiple attacking vehicles in a coordinated manner, where each attacking vehicle is driven by its respective second attack decision model.

[0206] In an embodiment of the present application, the determination module is further configured to:

[0207] Obtain the entropy of multiple actions output by the attacking vehicle in the accident scenario;

[0208] Based on multiple entropies, determine the state action of the attacking vehicle corresponding to the minimum entropy;

[0209] Based on the state action, determine the decision loopholes of the autonomous driving decision model.

[0210] An embodiment of the present application also provides an electronic device, including a processor and a memory. An executable program is stored in the memory, and the processor executes the executable program to perform the steps of the method as described above.

[0211] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium carries one or more computer programs, and when the one or more computer programs are executed by a processor, the steps of the method as described above are implemented.

[0212] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an electronic device, a computer-readable storage medium, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0213] The above-mentioned processor can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above-mentioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0214] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0215] The above-mentioned readable storage medium can be a magnetic disk, an optical disk, a DVD, a USB, a read-only memory (ROM), or a random access memory (RAM), etc. The present application does not limit the specific form of the storage medium.

[0216] The above embodiments are only exemplary embodiments of the present application and are not intended to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.

Claims

1. A decision vulnerability detection method for an autonomous driving decision model. It is characterized in that include: Constructing a detection scenario for detecting the autonomous driving decision model of a target vehicle, wherein the detection scenario includes at least one of the following: a road environment, the autonomous driving decision model, one or more attack vehicles for attacking the target vehicle, and a first attack decision model of the attack vehicle; Based on a multi-agent learning algorithm, the first attack decision model is trained using a behavior reward constructed for the attack vehicle to obtain a corresponding second attack decision model, wherein the behavior reward includes at least one of the following: an accident responsibility arbitration reward and a diversity accident scenario reward, wherein the accident responsibility arbitration reward is based on traffic rules and determines the accident responsibility through the vehicle driving information during the accident between the target vehicle and the attack vehicle, and the diversity accident scenario reward is based on the differences between the accidents. Based on the detection scenario, the target vehicle having the autonomous driving decision model interacts with the attack vehicle having the second attack decision model to generate an accident scenario caused by a decision vulnerability of the autonomous driving decision model; Based on the second attack decision model, a decision vulnerability of the autonomous driving decision model is determined from the accident scenario.

2. The method according to claim 1, It is characterized in that The construction of a detection scenario for detecting the automatic driving decision model of the target vehicle includes: constructing the road environment in a simulator; The initial position and destination of the target vehicle and the initial position of the attacking vehicle are determined.

3. The method according to claim 1, It is characterized in that in, The input of the first attack decision model is the state information of the attack vehicle and the relative position relationship between the attack vehicle and the target vehicle; the output of the first attack decision model is the decision action of the attack vehicle.

4. The method according to claim 1, It is characterized in that The behavior reward also includes a distance reward, and the distance reward is used to reward the attacking vehicle according to the distance between the target vehicle and the attacking vehicle; The step of training the first attack decision model by using the behavior reward constructed for the attack vehicle includes: Determining an initial decision action of the attacking vehicle based on the state of the detection scene; Based on the initial decision action of the attacking vehicle, adjusting the state of the detection scene and generating the corresponding behavior reward; The attack vehicle is rewarded based on the accident responsibility arbitration reward, the diverse accident scenario reward and / or the distance reward.

5. The method according to claim 4, It is characterized in that The distance reward includes a negative penalty and a positive reward, and the first attack decision model is trained by using the behavior reward constructed for the attack vehicle, including: When the distance between the attacking vehicle and the target vehicle exceeds a first predetermined distance, giving the attacking vehicle the negative penalty to encourage the attacking vehicle to approach the target vehicle; When the attacking vehicle gradually approaches the target vehicle, the attacking vehicle will receive a gradually increasing positive reward; When the attacking vehicle approaches the target vehicle and the distance between the vehicles is less than a second predetermined distance, the negative penalty is given to the attacking vehicle.

6. The method according to claim 1, It is characterized in that The accident liability arbitration reward includes a collision reward and an aggressive penalty, and the first attack decision model is trained by using the behavior reward constructed for the attacking vehicle, including: When it is determined that the attacking vehicle collides with the target vehicle, or when an environment vehicle in the detection scene collides with the target vehicle, giving the attacking vehicle the collision reward; When the attacking vehicle violates the driving regulations, the attacking vehicle is given the aggressive penalty to regulate the driving action.

7. The method according to claim 1, It is characterized in that The step of training the first attack decision model by using the behavior reward constructed for the attack vehicle includes: When it is determined that the attacking vehicle has explored a second dangerous scene different from the discovered first dangerous scene, the attacking vehicle is rewarded with the diverse accident scenes.

8. The method according to claim 1, It is characterized in that in, When there are multiple attacking vehicles, the target vehicle having the autonomous driving decision model interacts with the attacking vehicle having the second attack decision model, including: In the detection scenario, the target vehicle is attacked by a plurality of attack vehicles in coordination with each other, wherein each of the attack vehicles is driven by its own second attack decision model.

9. The method according to claim 1, It is characterized in that The determining, based on the second attack vehicle decision model, a decision vulnerability of the autonomous driving decision model from the accident scenario includes: Obtaining entropies of multiple actions output by the attacking vehicle in the accident scenario; Based on the plurality of entropies, determining a state action of the attacking vehicle corresponding to a minimum entropy; Based on the state action, a decision vulnerability of the autonomous driving decision model is determined.

10. A decision vulnerability detection device for an autonomous driving decision model, It is characterized in that include: A construction module configured to construct a detection scenario for detecting the autonomous driving decision model of a target vehicle, wherein the detection scenario includes at least one of the following: a road environment, the autonomous driving decision model, one or more attack vehicles for attacking the target vehicle, and a first attack decision model of the attack vehicle; A training module, configured to train the first attack decision model based on a multi-agent reinforcement learning algorithm using a behavior reward constructed for the attack vehicle to obtain a corresponding second attack decision model, wherein the behavior reward includes at least one of the following: an accident responsibility arbitration reward and a diversity accident scenario reward, wherein the accident responsibility arbitration reward is based on traffic rules and determines the accident responsibility through vehicle driving information during the accident between the target vehicle and the attack vehicle, and the diversity accident scenario reward is based on the differences between the accidents. an interaction module configured to interact the target vehicle having the autonomous driving decision model with the attack vehicle having the second attack decision model based on the detection scenario to generate an accident scenario caused by a decision vulnerability of the autonomous driving decision model; A determination module is configured to determine a decision vulnerability of the autonomous driving decision model from the accident scenario based on the second attack decision model.