Method and system for checking automated driving functions by reinforcement learning
Through reinforcement learning, the scene is generated and reward function is set, and the learning is accelerated using rules-based models and prior knowledge, the coverage and efficiency problems of autonomous driving function verification are solved, and the rapid and effective autonomous driving function inspection is achieved.
Patent Information
- Application Number
- CN202080073227.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-07
- Filing Date
- 2020-11-03
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2040-11-03
AI Technical Summary
The prior art is difficult to effectively verify and verify the safety strategy of autonomous driving functions. Traditional testing solutions cannot cover all possible traffic conditions, and sample-based analysis methods cannot check the integrity of autonomous driving functions end-to-end.
The reinforcement learning method is adopted to accelerate the learning process by generating scenes and setting reward functions, using rules-based models to speed up the learning process, forging scenarios that violate driving functions to expose weaknesses, and speed up learning by combining prior knowledge and inequality constraints.
The automatic driving function is realized quickly and effectively, and can cover a variety of traffic conditions, expose potential weaknesses, and improve verification efficiency and accuracy.
Smart Images

Figure CN114616157B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and system for checking an automated driving function by reinforcement learning, and in particular to generating scenarios that violate specifications of the automated driving function. Background Art
[0002] Driving assistance systems for automated driving are becoming increasingly important. Automated driving can be performed at different levels of automation. Exemplary levels of automation are assisted driving, partially automated driving, highly automated driving, or fully automated driving. These levels of automation are defined by the Federal Highway Research Institute (BASt) (see the BASt publication "Forschungkompakt," edition 11 / 2012). For example, a vehicle with Level 4 autonomous driving can operate completely autonomously in urban areas.
[0003] The main challenge in the development of autonomous driving functions is rigorous validation and verification to achieve compliance with safety policies and a sufficient level of customer confidence. Traditional testing approaches are only scalable for autonomous driving because a large number of real-world trips are required for each release.
[0004] One possible approach for validating and evaluating autonomous vehicles that must handle a wide range of possible traffic situations lies in virtual simulation environments. To obtain a convincing evaluation of autonomous driving functions from simulations, the simulated environment must be sufficiently realistic. Furthermore, the permitted behavior (specifications) of the autonomous vehicle must be automatically checked, and the implemented test scenarios must cover all typical situations as well as rare but realistic driving situations.
[0005] While some approaches can satisfy the first two requirements, fulfilling them is a non-trivial task due to the high dimensionality and non-convexity of the relevant parameter space. Data-driven approaches offer some remedy, but analysis of large amounts of real-world data does not guarantee that all relevant scenarios are included and tested. Consequently, most existing approaches rely on sample-based checks, possibly using analytical models. However, these approaches cannot cover the entire end-to-end driving function, from sensor data processing to the resulting actuator signals, and must be completely reimplemented whenever the system changes. Summary of the Invention
[0006] The present disclosure aims to provide a method for checking automated driving functions using reinforcement learning, a storage medium for implementing the method, and a system for checking automated driving functions using reinforcement learning. Reinforcement learning allows for rapid and efficient checking of automated driving functions. Furthermore, the present disclosure aims to effectively falsify automated driving functions to expose weaknesses in the automated driving functions.
[0007] This object is achieved by the subject matter of the independent claim. Advantageous embodiments are described in the dependent claims.
[0008] According to an independent aspect of the present disclosure, a method for checking an automated driving function using reinforcement learning is provided. The method includes: providing at least one specification for the automated driving function; generating a scenario, wherein the scenario is indicated by a first parameter set; and determining a reward function such that a reward is higher when a simulated scenario does not meet the at least one specification than when the simulated scenario meets the at least one specification. For example, the reward function can be determined using a rule-based model.
[0009] According to the present invention, a reward function is determined that relates the trajectories of all objects in the scene. In particular, the RL agent learns to generate scenarios that maximize reward and reflect violations of the driving function specifications. Therefore, incorporating existing prior knowledge into the training process accelerates learning. This allows the automated driving function to be effectively falsified, thereby exposing weaknesses in the automated driving function.
[0010] Preferably, the rule-based model describes a vehicle controller for an automated driving function. In this case, the controller is a (simplified) model of the behavior of a vehicle operating with the automated driving function.
[0011] Preferably, the method further comprises generating a second parameter set indicating modifications to the first parameter set. This may be done by an adversarial agent.
[0012] Preferably, the method further comprises:
[0013] Use a rule-based model in simulation to estimate the value of the reward function for a specific scenario R est ;
[0014] Generates the value corresponding to the third parameter set a t+1 Another scenario, in which based on the second parameter group a nn and parameter group a est , determine the third parameter group a t+1 , parameter group a est Make the estimate R of the rule-based model est Maximize; and
[0015] Find the reward function so that the value of the reward function for the scene in the simulation is estimated R est Below the actual value R of the reward function, the reward R is higher.
[0016] Preferably, an inequality constraint that excludes a specific scenario is used to generate additional scenarios corresponding to the third parameter set. The inequality constraint can be defined as:
[0017] |a nn -d est |<a Schwelle
[0018] The generation of the further scene corresponding to the third parameter set preferably takes place using a projection of the parameter set onto a set of specific scenes.
[0019] According to another aspect of the present disclosure, a system for checking an automated driving function by reinforcement learning is provided. The system includes a processor unit configured to implement a method for checking an automated driving function by reinforcement learning according to the embodiments described herein.
[0020] The system is particularly configured to implement the method described herein. The method can implement aspects of the system described herein.
[0021] The method according to the invention can also be simulated in a HIL (Hardware-in-the-Loop) environment.
[0022] According to a further independent aspect, a software (SW) program is described. The SW program can be arranged to be executed on one or more processors and thereby implement the method described herein.
[0023] According to a further independent aspect, a storage medium is described. The storage medium may include a SW program that is configured to be executed on one or more processors and thereby implement the method described herein.
[0024] In the present context, the term "automated driving" can be understood as driving with automated longitudinal or lateral guidance or autonomous driving with automated longitudinal and lateral guidance. Automated driving can include, for example, extended driving on highways or limited driving within the context of parking or maneuvering. The term "automated driving" encompasses automated driving with any degree of automation. Exemplary degrees of automation are assisted driving, partially automated driving, highly automated driving, or fully automated driving. These degrees of automation are defined by the Federal Highway Research Institute (BASt) (see the BASt publication "Forschung kompakt," edition 11 / 2012).
[0025] In assisted driving, the driver continuously controls the longitudinal or lateral steering, while the system takes over other functions within certain limits. In partially automated driving (TAF), the system takes over longitudinal and lateral steering for certain periods of time and / or in specific situations, requiring the driver to continuously monitor the system, as in assisted driving. In highly automated driving (HAF), the system takes over longitudinal and lateral steering for certain periods of time, without the driver having to constantly monitor the system; the driver does not necessarily have to be able to take over vehicle control within certain periods of time. In fully automated driving (VAF), the system can automatically handle driving in all situations for specific applications; a driver is no longer required for these applications.
[0026] The four degrees of automation described above correspond to SAE Levels 1 to 4 of the SAE J3016 standard (SAE Society of Automotive Engineers). For example, Highly Automated Driving (HAF) corresponds to Level 3 of the SAE J3016 standard. Furthermore, SAE J3016 specifies SAE Level 5 as the highest level of automation not included in the definition of BASt. SAE Level 5 corresponds to autonomous driving, in which the system can automatically handle all situations during the entire driving process, just as a human driver would; a driver is generally no longer required. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Embodiments of the present disclosure are shown in the accompanying drawings and described in more detail below.
[0028] Figure 1 Schematically illustrates a driving assistance system for automated driving according to an embodiment of the present disclosure,
[0029] Figure 2 A general schematic diagram showing a reinforcement learning scheme;
[0030] Figure 3 A flow chart showing a method for checking an automated driving function according to an embodiment of the present disclosure;
[0031] Figure 4 A schematic diagram illustrating a method for checking an automated driving function according to an embodiment of the present disclosure is shown; and
[0032] Figure 5 A schematic diagram is shown for checking an automated driving function according to another specific embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] Unless otherwise stated, the same reference numerals are used below for identical and identically acting elements.
[0034] Figure 1A driving assistance system for automated driving according to an embodiment of the present disclosure is schematically shown.
[0035] Vehicle 100 includes a driver assistance system 110 for automated driving. During automated driving, vehicle 100 is automatically guided longitudinally and transversely. Driver assistance system 110 thus takes over vehicle guidance. To this end, driver assistance system 110 controls drive 20, transmission 22, hydraulic service brake 24, and steering 26 via an intermediate unit (not shown).
[0036] To plan and execute automated driving, driver assistance system 110 receives environmental information from an environmental sensor system that observes the vehicle's surroundings. The vehicle may include, in particular, at least one environmental sensor 12 configured to record environmental data indicative of the vehicle's surroundings. The at least one environmental sensor 12 may include, for example, a LiDAR system, one or more radar systems, and / or one or more cameras.
[0037] The object of the present disclosure is to learn how to efficiently generate scenarios that distort the function based on automatically verifiable specifications and a continuous virtual simulation environment for an autonomous or automated driving function.
[0038] In one example, consider the ACC (Adaptive Cruise Control) function. The ACC function is configured to maintain a safe distance from the vehicle traveling in front. By means of a parameter defined as t h = time interval t of h / v h , the ACC request can be expressed as follows:
[0039] - Two possible modes: rated speed mode and time interval mode;
[0040] In rated speed mode, the speed v set or expected by the driver should be observed. d , that is, v d ∈[v d,min ;v d,max ];
[0041] In time interval mode, the lead time t h , that is, t h ∈[t h,min ;t h,max ].
[0042] When V d ≤h / t d When the system is in rated speed mode, otherwise the system is in time interval mode. In addition, the acceleration of the vehicle must always meet a c ∈[a c,min ;a c,max].
[0043] According to embodiments of the present disclosure, reinforcement learning (RL) is used, and in particular, an adversarial agent based on RL (represented in the figure by Agent). The RL agent learns to generate scenarios that maximize a specific reward. Because the agent's goal is to falsify the driving function, the reward function is designed such that the agent receives a high reward when the scenario results in a violation of the specification and a low reward when the autonomous driving function operates according to the specification.
[0044] The agent repeatedly observes the state of a system s, which includes all relevant variables for a given specification. Based on the state, the agent performs an action a according to its learned policy and receives the corresponding reward R(s, a). The action consists of a finite set of scenario parameters. Over time, the agent modifies its policy to maximize its reward.
[0045] The output of the RL agent is a set of scenario parameters a, which includes, for example, the initial vehicle speed, the desired vehicle, the initial time interval, and the speed segment v f The vehicle speed distribution of the finite time series encoding, where t i ∈t0, t1, …, t n It starts with an initial parameter set a0 and calculates the corresponding initial environment state s0. State s t Contains all variables relevant for checking compliance with the specification, such as minimum and maximum acceleration, minimum and maximum distance to the vehicle ahead, or minimum and maximum time progress, minimum and maximum speed, etc. All of the above specification statements can then be directly detected or numerically approximated by inequalities of the form A[s;a]-b≤0.
[0046] The input of the RL-based agent is the state of the environment s at time t t And the output is the modified scenario parameter a for the next run t+1 For example, the reward function is chosen so that R(s, a) = Σ x max(0, (exp(x)-1)) where x represents the value of any row on the left side of the inequality A[s;a]-b≤0 for the specification. This ensures that the reward is large only when the agent has found a scenario that violates the specification. Figure 2 A general schematic diagram showing this situation.
[0047] Typical RL approaches come at the cost of slowness and high variability. Learning complex tasks can require millions of iterations, and each iteration can be expensive. Furthermore, the variability between learning episodes can be very high, meaning that some episodes of the RL algorithm succeed while others fail due to randomness in initialization and scanning. This high variability in learning can be a significant obstacle to the application of RL. This problem becomes even more pronounced in large parameter spaces.
[0048] The above problems can be alleviated by introducing prior knowledge about the process, which can be obtained by the inequality g(s t , a t )≤0 (this inequality excludes scenarios with minor violations of the specification) is appropriately modeled, i.e., ensuring that the vehicle starts in a harmless (safe) state. This inequality is integrated into the learning process as a conditioning expression in the reward function or as an output constraint on the neural network used to focus learning progress. Any continuous variable RL method, such as policy gradient methods or actuator key methods, can be used for the RL agent.
[0049] Although the above approach can exclude many parameterizations that slightly violate the specification, a large number of sessions, which may last up to several days, are still required until the RL agent generates interesting scenarios. Therefore, more prior knowledge can be integrated to accelerate the learning process.
[0050] Figure 3 A flow chart of a method 300 for checking an automated driving function by reinforcement learning according to an embodiment of the present disclosure is shown.
[0051] The method 300 includes providing at least one specification for an automated driving function in box 310; generating a scenario in box 320, wherein the scenario is indicated by a first parameter set; and determining a reward function in box 330, such that a reward is higher if the scenario in the simulation does not meet the at least one specification than if the scenario in the simulation meets the at least one specification, wherein the reward function is determined using a rule-based model.
[0052] Independent of the actual algorithms used in the autonomous or automated vehicle, it is assumed that the vehicle is controlled by a conventional (rule-based) control system and the driving dynamics are described by a simple analytical model, all of which are expressed by the difference equation x k+1 =f k (x k , s t , a t ) can be detected, where x k Represents the state of the vehicle at the implementation time. On this basis, for the current environment state s tIt can be formulated as the following optimization problem:
[0053]
[0054] x k+1 =f k (x k , a est , s t )
[0055] Used to determine the new parameter group a est , to provide an estimate of the maximum reward R est,max If the optimization problem is not convex (which is usually the case), convex relaxation or other approximation methods can be used.
[0056] Then, the RL agent obtains the state s in parallel t and RL agent rewards
[0057] R nn =|R(s t , a t )-R est | n , n∈{1,2}
[0058] And generate a new parameter group a nn In this way, the RL agent can learn only the differences between the rule-based control behavior and the actual system rather than the entire system, and make corresponding modifications to the nn Finally, a new set of parameters for the next implementation is determined as a s+t =a est +a nn In order to avoid initialization in an unsafe state, the above method can be used to pass the inequality g(s t , a est )≤0 to approximate prior knowledge.
[0059] Figure 4 and Figure 5 Two possible schematic diagrams according to embodiments of the present disclosure are shown to achieve this.
[0060] The method includes generating a second parameter set indicating a modification to the first parameter set, and generating a further scenario corresponding to a third parameter set, wherein the third parameter set is determined based on the second parameter set and using a rule-based model.
[0061] In some embodiments, generating additional scenarios corresponding to the third parameter set is performed using inequality constraints that exclude certain scenarios. Figure 5 In particular, Figure 5 The box G in represents an exemplary inequality constraint of the following type:
[0062] |a nn -a est |<a Schwelle
[0063] However, the present disclosure is not limited to inequality constraints, and a generalized optimization problem may be used, which can be described as follows:
[0064]
[0065]
[0066]
[0067] Here, appropriate control input is selected according to the specific scenario category For example, to prevent a collision with a vehicle traveling ahead.
[0068] According to the present invention, a rule-based model is used, for example, to determine the reward function. Specifically, the RL agent learns to generate scenarios that maximize the reward and reflect violations of the driving function specifications. Therefore, incorporating existing prior knowledge into the training process accelerates learning. This allows for the efficient falsification of automated driving functions, thereby exposing weaknesses within them.
Claims
1. A method (300) for checking an automated driving function by reinforcement learning, comprising: providing (310) at least one specification for an automated driving function; generating (320) a scenario, wherein the scenario is indicated by a first parameter set; evaluating (330) a reward function for the scenario, wherein the reward function causes a higher reward R if the scenario in the simulation does not meet the at least one criterion than if the scenario in the simulation of the automated driving function meets the at least one criterion; using a rule-based model (RBM) in a simulation to estimate a value of the reward function for the scenario, wherein the rule-based model is a simplified model of the behavior of the vehicle having the automated driving function; generating a second parameter set by an adversarial agent, wherein the second parameter set indicates modifications to the first parameter set and a corresponding modified scenario, wherein the adversarial agent obtains an RL agent reward and generates the second parameter set based on the RL agent reward, wherein the RL agent reward corresponds to an absolute value of a difference between a reward value and a reward value estimate; determining a third set of parameters based on the second set of parameters and using the rule-based model; and A further scenario is generated corresponding to the third parameter set.
2. A storage medium comprising a software program, said software program being arranged to be executed on one or more processors and thereby implementing the method (300) according to claim 1.
3. A system for checking an automated driving function by reinforcement learning, comprising a processor unit configured to implement the method (300) for checking an automated driving function by reinforcement learning according to claim 1.