A security scenario acceleration test method and system based on adversarial reinforcement learning

By employing adversarial reinforcement learning, we acquire the test objects and test case library for autonomous driving, design acquisition functions and adversarial learning, and generate trajectory distributions. This solves the problems of low efficiency and low confidence in existing autonomous driving simulation testing platforms, and achieves efficient and high-confidence scenario construction and safety testing.

CN115292154BActive Publication Date: 2026-05-08SUZHOU GUANRUI AUTOMOBILE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU GUANRUI AUTOMOBILE TECH CO LTD
Filing Date
2022-05-24
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing autonomous driving simulation testing platforms suffer from low testing efficiency and insufficient confidence, making it difficult to construct highly complex and random traffic environments. They also lack efficient and high-confidence scenario construction methods, failing to meet testing objectives at different levels and failing to address the high safety issues of autonomous vehicles.

Method used

By adopting an adversarial reinforcement learning approach, we acquire autonomous driving test objects and a pre-set test case library, design the acquisition function of the machine learning agent simulation model, search for a suitable test set, use adversarial learning to conduct tests, generate trajectory distributions, and adjust the test set according to the trajectory distributions to achieve accelerated testing of driving scenarios with different safety levels.

Benefits of technology

It improves the safety and efficiency of autonomous vehicle testing, enables high-confidence scenario construction and rapid search, and meets the testing objectives of different levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115292154B_ABST
    Figure CN115292154B_ABST
Patent Text Reader

Abstract

The application relates to a kind of security scene acceleration test method and system based on adversarial reinforcement learning, method includes: obtaining automatic driving measured object and preset test case library;Based on the agent simulation model of machine learning, design the collection function that can balance development and exploration;Based on the automatic driving measured object, according to the test case library is searched according to the collection function, obtains the test set suitable for the automatic driving measured object;The test set includes multiple driving scenes;Based on the method of adversarial learning, according to the driving scene and the automatic driving measured object are tested, obtain trajectory distribution;According to the trajectory distribution, the test set is adjusted in time, to realize the acceleration test of different safety driving scenes.The application can improve the safety degree and test efficiency of automatic driving vehicle test.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle autonomous driving testing technology, and in particular to a method and system for accelerating safety scenario testing based on adversarial reinforcement learning. Background Technology

[0002] In recent years, various autonomous driving simulation and testing software programs have emerged in large numbers. In particular, many internet companies entering the autonomous driving industry have developed their own simulation software to support the research and development of autonomous driving technology. Typical examples include Baidu Apollo, Tencent TAD Sim, Microsoft AirSim, Waymo's CarCraft, Toyota / Intel's CARLAO, and Siemens PreScan. These simulation software programs offer a wide variety of functions and have different usage methods.

[0003] Current autonomous driving simulation testing platforms suffer from low testing efficiency and insufficient confidence. The vehicle driving environment is highly complex and random, and existing test scenarios are unable to reflect the real high complexity and randomness of traffic environments. They have failed to develop a multi-level, multi-difficulty level high-confidence scenario construction method for autonomous driving systems, and also lack a fast scenario search method for corresponding test targets. As a result, testing efficiency is low and confidence is insufficient. It is necessary to build an efficient high-confidence scenario construction theory and accelerated testing method that can adapt to different levels of requirements.

[0004] Achieving efficient matching between test objects and test cases to create personalized test methods that meet the characteristics of heterogeneous test objects is the core of accelerating testing. Reinforcement learning has been used in scenario-accelerated testing, but current commonly used learning methods have not solved the high safety requirements of autonomous vehicles. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for accelerating security scenario testing based on adversarial reinforcement learning.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] A method for accelerating security scenario testing based on adversarial reinforcement learning includes:

[0008] Obtain the autonomous driving test object and the preset test case library;

[0009] Based on a machine learning-based agent simulation model, a data acquisition function that balances development and exploration is designed.

[0010] Based on the autonomous driving test object, the test case library is searched according to the acquisition function to obtain a test set suitable for the autonomous driving test object; the test set includes multiple driving scenarios.

[0011] Based on the adversarial learning method, tests are performed according to the driving scenario and the autonomous driving test object to obtain the trajectory distribution;

[0012] The test set is adjusted in a timely manner according to the trajectory distribution to achieve acceleration testing in driving scenarios with different safety levels.

[0013] Preferably, the method for determining the test case library is as follows:

[0014] Based on the preset test objectives for autonomous driving and the general test scenario data format, the scenario parameters are divided into environmental parameters, road condition parameters, and object parameters; the test objectives include the design operating domain, dynamic driving tasks, and traffic regulations.

[0015] Determine the set of parameter constraints based on the requirements of the test scenario, and obtain the set of uncovered combinations in the test scenario;

[0016] The scene parameters are evaluated to obtain a set of scene parameter combinations;

[0017] Based on the parameter constraint set, parameter values ​​are selected from the scene parameter combination set to obtain the parameter value combination set;

[0018] The value selection process is iterated repeatedly to resolve the parameter space, which includes the uncovered combination set.

[0019] Based on the combined testing algorithm, multiple test cases are generated according to the parameter space; the test case library consists of multiple test cases.

[0020] Preferably, determining the parameter constraint set according to the requirements of the test scenario includes:

[0021] The requirements in the test scenario are transformed into multiple constraints;

[0022] The initial constraint set is obtained based on the constraints;

[0023] The initial constraint set is simplified to obtain the simplest constraint set;

[0024] The implicit constraints are searched in the simplest constraint set, and the found implicit constraints are added to the initial constraint set to obtain the final parameter constraint set.

[0025] Preferably, the machine learning-based agent simulation model is designed with a data acquisition function that balances development and exploration, including:

[0026] Acquire simulation test scenario data;

[0027] The simulation test scenario data is used as training data to train a preset multilayer perceptron neural network to obtain a trained test case matching proxy model.

[0028] The network parameters of the trained test case matching proxy model are iteratively updated by the hyperparameter configuration function to obtain a test case matching proxy model with optimal generalization ability; the hyperparameter configuration function corresponding to the test case matching proxy model with optimal generalization ability is the acquisition function; the network parameters include the number of hidden layers, the number of hidden layer neurons, the learning rate, the exponential decay rate, and the training data batch size.

[0029] Preferably, the step of searching the test case library based on the autonomous driving test object according to the acquisition function to obtain a test set suitable for the autonomous driving test object includes:

[0030] Based on the hyperparameter optimization algorithm, the optimal hyperparameter configuration is obtained according to the acquisition function;

[0031] The search direction for input sample points is determined based on the optimal hyperparameter configuration.

[0032] The test case library is searched according to the search direction of the input sample points to obtain the test set of the autonomous driving test object.

[0033] Preferably, the adversarial learning-based method, which performs tests based on the driving scenario and the autonomous driving test object to obtain the trajectory distribution, includes:

[0034] Each Nash equilibrium point is determined through reinforcement learning interactive adversarial training; each Nash equilibrium point corresponds to a different type of driving scenario; the autonomous driving test object is modified and evolved according to the different types of driving scenarios;

[0035] New Nash equilibrium points are determined through adversarial reinforcement learning;

[0036] Test the test cases of the modified and evolved autonomous driving test object based on the new Nash equilibrium point to generate the trajectory distribution based on adversarial learning.

[0037] Preferably, the step of adjusting the test set in a timely manner according to the trajectory distribution to achieve acceleration testing in driving scenarios with different safety levels includes:

[0038] Based on importance sampling theory, an unbiased estimate of the performance of the tested autonomous driving system in a natural driving environment is established;

[0039] The importance sampling weights are obtained based on the probability distribution of the trajectory distribution;

[0040] During the testing process, the probability of occurrence of risk events in the test set is increased based on the importance sampling weight and the unbiased estimation to achieve high-confidence accelerated testing.

[0041] A security scenario acceleration testing system based on adversarial reinforcement learning includes:

[0042] The acquisition module is used to acquire the autonomous driving test object and the preset test case library;

[0043] The design module is used for machine learning-based agent simulation models to design acquisition functions that balance development and exploration.

[0044] The search module is used to search the test case library based on the autonomous driving test object and according to the acquisition function to obtain a test set suitable for the autonomous driving test object; the test set includes multiple driving scenarios.

[0045] The testing module is used to perform tests based on the driving scenario and the autonomous driving test object using an adversarial learning method to obtain the trajectory distribution.

[0046] An adjustment module is used to adjust the test set in a timely manner according to the trajectory distribution in order to achieve acceleration testing in driving scenarios with different safety levels.

[0047] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0048] This invention provides a method and system for accelerating safety scenario testing based on adversarial reinforcement learning. The method includes: acquiring an autonomous driving test object and a preset test case library; designing a data acquisition function that balances development and exploration based on a machine learning-based agent simulation model; searching the test case library according to the data acquisition function based on the autonomous driving test object to obtain a test set suitable for the autonomous driving test object; the test set includes multiple driving scenarios; performing tests based on the driving scenarios and the autonomous driving test object using an adversarial learning method to obtain a trajectory distribution; and adjusting the test set in a timely manner according to the trajectory distribution to achieve accelerated testing of driving scenarios with different safety levels. This invention can improve the safety level and testing efficiency of autonomous vehicle testing. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 A flowchart of a security scenario acceleration testing method based on adversarial reinforcement learning in an embodiment of the present invention;

[0051] Figure 2 The module connection diagram of the security scenario acceleration testing system based on adversarial reinforcement learning in the embodiments provided by the present invention is shown. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] The purpose of this invention is to provide a method and system for accelerating safety scenario testing based on adversarial reinforcement learning, which can improve the safety level and testing efficiency of autonomous vehicle testing.

[0054] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] Figure 1 The flowchart of the security scenario acceleration testing method based on adversarial reinforcement learning in the embodiments provided by the present invention is as follows: Figure 1 As shown, this invention provides a method for accelerating security scenario testing based on adversarial reinforcement learning, comprising:

[0056] Step 100: Obtain the autonomous driving test object and the preset test case library;

[0057] Step 200: Based on the machine learning-based agent simulation model, design a data acquisition function that balances development and exploration;

[0058] Step 300: Based on the autonomous driving test object, search the test case library according to the acquisition function to obtain a test set suitable for the autonomous driving test object; the test set includes multiple driving scenarios;

[0059] Step 400: Using an adversarial learning method, perform tests based on the driving scenario and the autonomous driving test object to obtain the trajectory distribution;

[0060] Step 500: Adjust the test set in a timely manner according to the trajectory distribution to achieve acceleration testing in driving scenarios with different safety levels.

[0061] Optionally, in this embodiment, based on the test objectives of autonomous driving in terms of design and operation domain, dynamic driving tasks, and traffic regulations, the scenario element hierarchy is deconstructed, a set of working condition test scenario schemes is constructed according to the design and operation domain, and a test case library corresponding to the test requirements is generated through quantification and iterative evolution.

[0062] Preferably, the method for determining the test case library is as follows:

[0063] Based on the preset test objectives for autonomous driving and the general test scenario data format, the scenario parameters are divided into environmental parameters, road condition parameters, and object parameters; the test objectives include the design operating domain, dynamic driving tasks, and traffic regulations.

[0064] Determine the set of parameter constraints based on the requirements of the test scenario, and obtain the set of uncovered combinations in the test scenario;

[0065] The scene parameters are evaluated to obtain a set of scene parameter combinations;

[0066] Based on the parameter constraint set, parameter values ​​are selected from the scene parameter combination set to obtain the parameter value combination set;

[0067] The value selection process is iterated repeatedly to resolve the parameter space, which includes the uncovered combination set.

[0068] Based on the combined testing algorithm, multiple test cases are generated according to the parameter space; the test case library consists of multiple test cases.

[0069] Specifically, this embodiment is based on the test objectives of autonomous driving in terms of design operating domain, dynamic driving tasks, and traffic regulations, as shown in Table 1. Referring to existing research and common test scenario data formats, the scenario parameters are divided into three categories: environmental parameters, road condition parameters, and object parameters. Since there are constraints between the parameters in the scenario, it is necessary to calculate all constraints between the parameters in the current scenario to avoid generating test scenarios that do not conform to reality. Specifically, the requirements in the scenario are transformed into different types of constraints, and then all constraints are transformed into prohibitive constraints to obtain an initial constraint set. The initial constraint set is simplified to obtain the simplest constraint set. Implicit constraints are searched under the simplest constraint set, and these implicit constraints are incorporated into the original constraint set, ultimately reducing the size of the constraint set.

[0070] Table 1

[0071]

[0072] Furthermore, before generating test cases using the combinatorial testing algorithm, it is necessary to generate the set of uncovered combinations for the current test scenario. This involves three processes: scenario parameter combination, parameter value combination, and random sorting of uncovered combinations. Values ​​are assigned to all scenario parameters. When combining scenario parameters, parameter constraints do not need to be considered; simply listing all combinations is sufficient. After combining scenario parameters, parameter value combination is required, at which point the influence of parameter constraints must be considered. Combinations appearing in the constraints must be excluded from all parameter value combinations; otherwise, it will never be possible to include all uncovered combinations in the test case set. This generation process is repeated, and the generated uncovered combinations always follow a specific order. The combinatorial testing algorithm can generate a large number of test cases in the parsed parameter space. To avoid affecting the probabilistic controllability of the combinatorial testing algorithm, a Bayesian network model of the scenario is constructed to calculate the probability of test cases, ensuring that the probability of generating test cases is controllable.

[0073] Preferably, determining the parameter constraint set according to the requirements of the test scenario includes:

[0074] The requirements in the test scenario are transformed into multiple constraints;

[0075] The initial constraint set is obtained based on the constraints;

[0076] The initial constraint set is simplified to obtain the simplest constraint set;

[0077] The implicit constraints are searched in the simplest constraint set, and the found implicit constraints are added to the initial constraint set to obtain the final parameter constraint set.

[0078] Preferably, step 200 specifically includes:

[0079] Acquire simulation test scenario data;

[0080] The simulation test scenario data is used as training data to train a preset multilayer perceptron neural network to obtain a trained test case matching proxy model.

[0081] The network parameters of the trained test case matching proxy model are iteratively updated by the hyperparameter configuration function to obtain a test case matching proxy model with optimal generalization ability; the hyperparameter configuration function corresponding to the test case matching proxy model with optimal generalization ability is the acquisition function; the network parameters include the number of hidden layers, the number of hidden layer neurons, the learning rate, the exponential decay rate, and the training data batch size.

[0082] Furthermore, this embodiment, considering the characteristics of autonomous driving system test data, uses the simulation test scenario data as training data for the neural network based on the MLP proxy model. It establishes a test case matching proxy model based on a safety scenario domain, and designs the number of hidden layers k and the number of hidden layer neurons n in the MLP network. k The learning rate η and exponential decay rate β of the Adam algorithm t And the training data batch size B, etc., are configured through hyperparameters. The convergence result of the training network is affected, and the best generalization ability of the surrogate model can be obtained by adjusting the MLP structure.

[0083] Specifically, MLP is the most commonly used neural network framework in constructing deep regression surrogate models. An MLP mainly consists of three parts: an input layer, hidden layers, and an output layer, where the number of hidden layers... Generally greater than or equal to 2. Similar to traditional neural network models, MLP uses the activation functions of hidden layer neurons to map data features. By superimposing the mappings between multiple hidden layer neurons, it learns the complex mapping relationships between multiple inputs and multiple outputs of data.

[0084] The input layer neurons in an MLP are primarily used for data buffering. After the data enters the hidden layer neurons, it undergoes weighted processing, as follows:

[0085]

[0086] in, It is the first The first layer One output, It is the first The first layer One output, It is the connection of the first The first layer Outputs and The first layer The weights of each output, It is the first Layer bias.

[0087] Then, the data is transformed using an activation function:

[0088]

[0089] In previous studies, various data transformation functions have been used to construct MLP neural networks, such as the Identity function, the sigmoid function, etc. Functions include ReLU, LeakyReLU, ELU, Binary, ExponentialLinearUnit, and SoftSign. Among these, ReLU is one of the most commonly used activation functions because it is computationally fast and avoids the gradient vanishing or exploding issues that occur with the sigmoid function during network training. Therefore, the MLP deep proxy model construction process in this embodiment mainly uses the ReLU activation function.

[0090] The construction of the MLP deep proxy model is mainly based on the BP (Back Propagation) algorithm. First, partial derivatives are calculated layer by layer using the chain rule to obtain the error of each layer. Then, the parameters of each layer are corrected as follows:

[0091]

[0092]

[0093] in, and These are the partial derivatives of the error of each layer with respect to the weight and bias of that layer, respectively. It is the learning rate.

[0094] Furthermore, after updating the network, forward and backward propagation are re-executed. Through this iterative learning strategy, the model's generalization ability is continuously improved until the performance of the surrogate model converges.

[0095] Preferably, step 300 specifically includes:

[0096] Based on the hyperparameter optimization algorithm, the optimal hyperparameter configuration is obtained according to the acquisition function;

[0097] The search direction for input sample points is determined based on the optimal hyperparameter configuration.

[0098] The test case library is searched according to the search direction of the input sample points to obtain the test set of the autonomous driving test object.

[0099] Furthermore, this embodiment improves the efficiency of single-sample calculation by adaptively adjusting the target search through the acquisition function, accurately locating the search direction of sample points, and thus adaptively and quickly searching for a new test set suitable for the test object of autonomous driving.

[0100] Specifically, by designing an adaptive construction method for the surrogate model of the hyperparameter optimization algorithm, each given set of hyperparameter configurations... After the trained network converges, a deep proxy model will be obtained. A model that will achieve uninterrupted generalization ability.

[0101] Furthermore, let Given a set of hyperparameter configurations The mathematical expression for the deep surrogate model obtained by training a deep neural network and converging on the training set is:

[0102] ;

[0103] in, It is the parameter space of the neural network. It is the loss of the surrogate model on the training set. This is the parameter matrix of the model. For the training dataset, This is the j-th input.

[0104] make For the proxy model In the validation set The loss on, and its relationship with the optimal hyperparameter configuration The mathematical expression for the relationship is:

[0105] ;

[0106] in, It is the hyperparameter space of neural networks. This is the i-th input.

[0107] For a given validation set, cross-validation can be used to obtain... This process is also called hyperparameter optimization. Typically, exhaustive verification is difficult. All possible configurations. Therefore, the formula for the optimal hyperparameter configuration is rewritten as:

[0108] ;

[0109] in, It is the hyperparameter configuration for evaluation. Quantity, It is all configurations The best among them. When obtained This represents a deep proxy model that has achieved near-optimal generalization ability.

[0110] Specifically, according to the rewritten equation, the process of hyperparameter configuration optimization and evaluation is also the process of constructing a proxy model. By automatically obtaining the optimal hyperparameter configuration based on hyperparameter optimization algorithms, the search direction of input sample points is determined, thereby adaptively and quickly searching for a new test set suitable for the autonomous driving test object.

[0111] Preferably, step 400 specifically includes:

[0112] Each Nash equilibrium point is determined through reinforcement learning interactive adversarial training; each Nash equilibrium point corresponds to a different type of driving scenario; the autonomous driving test object is modified and evolved according to the different types of driving scenarios;

[0113] New Nash equilibrium points are determined through adversarial reinforcement learning;

[0114] Test the test cases of the modified and evolved autonomous driving test object based on the new Nash equilibrium point to generate the trajectory distribution based on adversarial learning.

[0115] Furthermore, this embodiment uses adversarial reinforcement learning training to find Nash equilibrium points. Multiple Nash equilibrium points may exist in high-dimensional driving scenarios, corresponding to different types of driving scenarios. The autonomous driving test subject is modified and evolved according to these different types of driving scenarios. New Nash equilibrium points are quickly found through adversarial reinforcement learning, and then adaptively and quickly matched with suitable test cases for testing, generating trajectory distributions based on adversarial learning.

[0116] Specifically, this embodiment constructs a strategy π that mimics expert behavior, and a discriminator for distinguishing between expert trajectories and virtual trajectories. In this context, policy π corresponds to the generator G in generative adversarial imitation learning, where π is responsible for interacting with the environment to generate virtual trajectories τ, while the discriminator... It is a model trained by supervised learning.

[0117] Given the environment state s and the action a of the autonomous driving test object, the next state s′ given by the environment can be expressed in the following form:

[0118]

[0119] in for The parameters. Let and In this embodiment, the reward functions for the execution policy of the tested object and the environment policy are obtained respectively, and can be used to improve the generator through reinforcement learning. and Conduct training. Because of simultaneous learning... and Both strategies would result in an excessively large search space for the algorithm, making it difficult to search and optimize, and ultimately preventing the acquisition of a high-performance strategy. Therefore, this embodiment will use an environment-based strategy. Execution strategy with the tested object Combined into a joint strategy It can be described by the following formula:

[0120]

[0121] Through joint strategies Optimization can be performed simultaneously on and Learn from the experts. and the initial state distribution of the environment As input. In each iteration, first let and Interact to generate virtual simulation trajectories Then, the expert trajectory... and virtual trajectory As a discriminant The input is used to minimize the loss function through supervised learning. The discriminator is optimized using this method. Then, the discriminator is used to evaluate the state-action-state triplet values ​​given by the discriminator in the virtual trajectory. As a joint strategy The reward function is used to update the joint policy using reinforcement learning methods. After repeating the above iterative process until convergence, take the value from the joint strategy. This embodiment uses the environment transfer model as the objective and employs it to construct a virtual environment for training. In other words, this embodiment can learn a trajectory discriminator generated through interaction with the environment model. Adversarial strategies that yield lower evaluation values This is used to find states where the environment model's learning of the environment transition function is inaccurate. It's easy to see that learning adversarial policies... The process can also be solved using reinforcement learning. The reward function corresponding to the adversarial policy required in this embodiment is:

[0122]

[0123] Then use this reward function to construct a Markov process. The environment transfer function in this Markov decision process This refers to the environmental model obtained earlier, namely... The reward function is maximized using reinforcement learning methods. After formulating the strategy, we can obtain a state-adversarial strategy that can find the current environment transition model but has poor accuracy in mimicking it. .

[0124] Given an adversary At that time, an environment model π can be provided. E

[0125] Design the following reward function:

[0126]

[0127] That is, it requires the environment model to be in adversarial strategy. After performing the action, a more precise next state is provided. Similarly, reinforcement learning is used to find the maximum reward function. After determining the strategy, we can obtain a method that can find the adversarial strategy. A more accurate environmental model that mimics visited states .

[0128] In this context, the state and reward of the autonomous driving test object are both under the combined behavior of multiple influencing factors. Its state and reward are also subject to joint behavior. Therefore, the process of finding the optimal joint strategy is also the process of finding the Nash equilibrium of the game.

[0129] According to the Bellman equation, the Nash Q function of the measured object i is defined as the joint policy at state s, where the measured object... Nash The function is defined as follows: for state Joint actions of the place When all tested objects execute the Nash equilibrium strategy, the tested objects The sum of the current return and the future return is:

[0130]

[0131] in, To combine with Nash's strategy, For the object being tested In state Division and joint action The return, The total discounted return for all other participants under the given state when implementing the Nash equilibrium strategy.

[0132] Based on the above definition, iterative approximation can be performed directly by adopting a Nash equilibrium strategy that maximizes the return. The Dardai equation is:

[0133]

[0134] in The object being tested Nash equilibrium point in the new state, Nash The learning algorithm updates and learns using the above formula.

[0135] Adversarial reinforcement learning is used to quickly find new Nash equilibrium points, and then adaptively and rapidly match suitable test cases for the autonomous driving test object. To evaluate the safety performance of the test object, the probability distribution of vehicle trajectories under different driving environments is crucial, requiring the generation of trajectory distributions based on adversarial learning. The natural environment of the test object is represented by a series of parameters, which are pre-determined by the operational design domain (e.g., road type, weather conditions, etc.) and potentially changing variables (e.g., the acceleration of background vehicles). The variables can be represented as:

[0136] ;

[0137] Where X i,j Let X represent the variables (e.g., position and speed) of the i-th background vehicle at time step j, N represent the number of background vehicles of interest, T represent the total number of time steps, and X represent the feasible space of the variables. The variable values ​​are sampled according to the optimal joint distribution of the variables, denoted as x∼P(x). Since P(x) has extremely high dimension, the problem is simplified by utilizing the spatiotemporal independence relationships between the variables. Assuming the Markov property, the joint distribution can be simplified by decomposition as follows:

[0138] ;

[0139] Here, the states and actions at time steps k=0,…,T are denoted as:

[0140] ;

[0141] ;

[0142] Where s0 represents the state of the object being measured (e.g., position and velocity), s i (i=1,…,N) represents the state of the i-th background vehicle, u i Let P(u(k)|s(k)) represent the i-th background vehicle's action (e.g., longitudinal acceleration). To simplify P(u(k)|s(k)), assuming all background vehicles simultaneously and independently choose their actions, it can be computed in a decomposition manner:

[0143] ;

[0144] P(u i (k)|s(k)) is further simplified by assuming space independence, let N i Let P(u) represent all vehicles that are dependent on the i-th background vehicle. i (k)|s(k)) can be approximated as Finally, the empirical probabilities of state-action pairs are calculated. Obtain the probability distribution of the trajectory.

[0145] Preferably, step 500 specifically includes:

[0146] Based on importance sampling theory, an unbiased estimate of the performance of the tested autonomous driving system in a natural driving environment is established;

[0147] The importance sampling weights are obtained based on the probability distribution of the trajectory distribution;

[0148] During the testing process, the probability of occurrence of risk events in the test set is increased based on the importance sampling weight and the unbiased estimation to achieve high-confidence accelerated testing.

[0149] Specifically, based on importance sampling theory, an unbiased estimate of the performance of the tested autonomous driving system in a natural driving environment is established. The probability distribution of the trajectory is used as an importance function to calculate the importance sampling weight. During the test, the probability of low-probability risk events is increased, and the variance is reduced while ensuring unbiasedness, thereby achieving high-confidence accelerated testing.

[0150] Specifically: Based on adversarial natural driving environment testing, if the low-probability risk events in the test subject's testing process are represented as A, the driving intelligence of the test subject can be measured in the following ways:

[0151] ;

[0152] Where x represents the variable of the driving environment, X represents its feasible region, n represents the number of tests, m represents the number of events A during the test, and q(x) is the importance function.

[0153] By introducing an importance function, the testing priority of key scenarios can be increased, thereby improving evaluation efficiency. The performance estimation equation is then obtained as follows:

[0154] ;

[0155] Where T i This represents the total time step of the i-th test. The test will terminate if event A occurs or the predetermined distance traveled is reached. Let T... i , c Let be the set of critical moments in the i-th test, and finally, the performance estimation equation can be obtained as:

[0156] ;

[0157] in It is the simulated weight (likelihood ratio) recorded during the test. T(A|x iThe probability of low-probability events is estimated by calculating the number of incidents that occur during the test. Based on this equation, the probability of low-probability events occurring in the tested object can be estimated from the test results, reducing variance while ensuring unbiasedness and completing high-confidence accelerated testing.

[0158] Figure 2 The module connection diagram of the security scenario acceleration testing system based on adversarial reinforcement learning in the embodiments provided by the present invention is as follows: Figure 2 As shown, this embodiment also provides a security scenario acceleration testing system based on adversarial reinforcement learning, including:

[0159] The acquisition module is used to acquire the autonomous driving test object and the preset test case library;

[0160] The design module is used for machine learning-based agent simulation models to design acquisition functions that balance development and exploration.

[0161] The search module is used to search the test case library based on the autonomous driving test object and according to the acquisition function to obtain a test set suitable for the autonomous driving test object; the test set includes multiple driving scenarios.

[0162] The testing module is used to perform tests based on the driving scenario and the autonomous driving test object using an adversarial learning method to obtain the trajectory distribution.

[0163] An adjustment module is used to adjust the test set in a timely manner according to the trajectory distribution in order to achieve acceleration testing in driving scenarios with different safety levels.

[0164] The beneficial effects of this invention are as follows:

[0165] This invention enables accelerated testing of driving scenarios with varying safety levels, improving the safety and efficiency of autonomous vehicle testing.

[0166] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0167] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for accelerating security scenario testing based on adversarial reinforcement learning, characterized in that, include: Obtain the autonomous driving test object and the preset test case library; Based on a machine learning-based agent simulation model, a data acquisition function that balances development and exploration is designed. The machine learning-based agent simulation model is designed with a data acquisition function that balances development and exploration. This function includes: acquiring simulation test scenario data; using the simulation test scenario data as training data to train a pre-defined multilayer perceptron neural network to obtain a trained test case matching agent model; iteratively updating the network parameters of the trained test case matching agent model through a hyperparameter configuration function to obtain a test case matching agent model with optimal generalization ability; the hyperparameter configuration function corresponding to the test case matching agent model with optimal generalization ability is the data acquisition function; the network parameters include the number of hidden layers, the number of hidden layer neurons, the learning rate, the exponential decay rate, and the training data batch size. Based on the autonomous driving test object, the test case library is searched according to the acquisition function to obtain a test set suitable for the autonomous driving test object; the test set includes multiple driving scenarios. Based on the adversarial learning method, tests are performed according to the driving scenario and the autonomous driving test object to obtain the trajectory distribution; The test set is adjusted in a timely manner according to the trajectory distribution to achieve accelerated testing of driving scenarios with different safety levels. This adjustment includes: establishing an unbiased estimate of the performance of the tested autonomous driving system in a natural driving environment based on importance sampling theory; obtaining importance sampling weights based on the probability distribution of the trajectory distribution; and increasing the probability of risk events occurring in the test set during testing based on the importance sampling weights and the unbiased estimate to achieve high-confidence accelerated testing.

2. The method for accelerating security scenario testing based on adversarial reinforcement learning according to claim 1, characterized in that, The method for determining the test case library is as follows: Based on the preset test objectives for autonomous driving and the general test scenario data format, the scenario parameters are divided into environmental parameters, road condition parameters, and object parameters; the test objectives include the design operating domain, dynamic driving tasks, and traffic regulations. Determine the set of parameter constraints based on the requirements of the test scenario, and obtain the set of uncovered combinations in the test scenario; The values ​​of all the scene parameters are taken to obtain a set of scene parameter combinations; Based on the parameter constraint set, parameter values ​​are selected from the scene parameter combination set to obtain the parameter value combination set; The value selection process is iterated repeatedly to resolve the parameter space; The parameter space includes the uncovered set of combinations; Based on the combined testing algorithm, multiple test cases are generated according to the parameter space; The test case library consists of multiple test cases.

3. The accelerated testing method for security scenarios based on adversarial reinforcement learning according to claim 2, characterized in that, The process of determining the parameter constraint set based on the requirements of the test scenario includes: The requirements in the test scenario are transformed into multiple constraints; The initial constraint set is obtained based on the constraints; The initial constraint set is simplified to obtain the simplest constraint set; The implicit constraints are searched in the simplest constraint set, and the found implicit constraints are added to the initial constraint set to obtain the final parameter constraint set.

4. The accelerated testing method for security scenarios based on adversarial reinforcement learning according to claim 1, characterized in that, The step of searching the test case library based on the autonomous driving test object according to the acquisition function to obtain a test set suitable for the autonomous driving test object includes: Based on the hyperparameter optimization algorithm, the optimal hyperparameter configuration is obtained according to the acquisition function; The search direction for input sample points is determined based on the optimal hyperparameter configuration. The test case library is searched according to the search direction of the input sample points to obtain the test set of the autonomous driving test object.

5. The accelerated testing method for security scenarios based on adversarial reinforcement learning according to claim 1, characterized in that, The adversarial learning-based method performs tests based on the driving scenario and the autonomous driving test object to obtain a trajectory distribution, including: Each Nash equilibrium point is determined through reinforcement learning interactive adversarial training; each Nash equilibrium point corresponds to a different type of driving scenario; the autonomous driving test object is modified and evolved according to the different types of driving scenarios; New Nash equilibrium points are determined through adversarial reinforcement learning; Test the test cases of the modified and evolved autonomous driving test object based on the new Nash equilibrium point to generate the trajectory distribution based on adversarial learning.

6. A security scenario acceleration testing system based on adversarial reinforcement learning, characterized in that, include: The acquisition module is used to acquire the autonomous driving test object and the preset test case library; The design module is used for machine learning-based agent simulation models to design acquisition functions that balance development and exploration. The machine learning-based agent simulation model is designed with a data acquisition function that balances development and exploration. This function includes: acquiring simulation test scenario data; using the simulation test scenario data as training data to train a pre-defined multilayer perceptron neural network to obtain a trained test case matching agent model; iteratively updating the network parameters of the trained test case matching agent model through a hyperparameter configuration function to obtain a test case matching agent model with optimal generalization ability; the hyperparameter configuration function corresponding to the test case matching agent model with optimal generalization ability is the data acquisition function; the network parameters include the number of hidden layers, the number of hidden layer neurons, the learning rate, the exponential decay rate, and the training data batch size. The search module is used to search the test case library based on the autonomous driving test object and according to the acquisition function to obtain a test set suitable for the autonomous driving test object; the test set includes multiple driving scenarios. The testing module is used to perform tests based on the driving scenario and the autonomous driving test object using an adversarial learning method to obtain the trajectory distribution. An adjustment module is used to adjust the test set in a timely manner according to the trajectory distribution to achieve accelerated testing of driving scenarios with different safety levels. This adjustment includes: establishing an unbiased estimate of the performance of the tested autonomous driving system in a natural driving environment based on importance sampling theory; obtaining importance sampling weights based on the probability distribution of the trajectory distribution; and increasing the probability of risk events occurring in the test set during testing based on the importance sampling weights and the unbiased estimate to achieve high-confidence accelerated testing.

Citation Information

Patent Citations

  • Test method, device and system for automatic driving vehicle

    CN110160804A

  • Driving scenario machine learning network and driving environment simulation

    US20200353943A1