An automatic driving unknown dangerous scene generation method and electronic equipment
By employing a joint learning method of criticality and adversarial degree in intensive deep reinforcement learning in autonomous driving systems, the generation of unknown dangerous scenarios is optimized, solving the problem of difficulty in generating unknown dangerous scenarios in existing technologies and achieving efficient and accurate safety testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2025-08-26
- Publication Date
- 2026-07-21
AI Technical Summary
Existing deep reinforcement learning-based methods for generating autonomous driving scenarios have a criticality close to zero when faced with unknown and dangerous scenarios, failing to effectively reflect the importance of the scenario. This makes it difficult to efficiently generate and evaluate unknown and dangerous scenarios, affecting the accuracy and efficiency of autonomous driving safety testing.
We employ a joint learning method for criticality and adversarial degree based on intensive deep reinforcement learning. By optimizing the criticality and adversarial degree of state and action, we generate unknown dangerous scenarios for autonomous driving. Combined with a defensive importance sampling strategy, we optimize the sample distribution and improve testing efficiency and accuracy.
It significantly improves the ability to generate unknown and dangerous scenarios, reduces computational costs, optimizes accident rate estimation, and enhances the efficiency of autonomous driving testing and the reliability of safety assessment.
Smart Images

Figure CN121143113B_ABST
Abstract
Description
Technical Field
[0001] This article relates to autonomous driving technology, and in particular to a method and electronic device for generating unknown and dangerous scenarios for autonomous driving. Background Technology
[0002] Currently, autonomous vehicles still face safety challenges in complex driving environments. Traditional safety assessment methods struggle to comprehensively cover potential hazardous scenarios, limiting the accuracy of accident rate predictions. Therefore, efficiently identifying and assessing potentially dangerous scenarios that autonomous driving systems may encounter has become a significant challenge before the large-scale deployment of autonomous vehicles.
[0003] In related technologies, the safety verification of autonomous driving systems can be optimized by generating highly efficient key test scenarios in test scenario generation based on deep reinforcement learning. Deep reinforcement learning utilizes proximal policy optimization algorithms to train adversarial algorithms, thereby dynamically adjusting the importance sampling strategy to increase the sampling probability of key scenarios while maintaining policy stability.
[0004] However, existing scene generation methods based on deep reinforcement learning only involve learning adversarial degree. When facing unknown and dangerous scenes, the criticality often approaches zero, which cannot effectively reflect the importance of the scene. Summary of the Invention
[0005] This application provides a method and electronic device for generating unknown dangerous scenarios for autonomous driving, which can effectively reflect the importance of the scenario by jointly learning criticality and adversarial degree based on deep reinforcement learning.
[0006] This application provides a method for generating unknown dangerous scenarios for autonomous driving, the method comprising: Obtain a first importance sampling strategy; wherein, the first importance sampling strategy is an optimized importance sampling strategy obtained by imitating the second importance sampling strategy and simultaneously optimizing the state-action criticality and adversarial degree; the second importance sampling strategy is an importance sampling strategy obtained based on deep reinforcement learning and adversarial degree learning for the target type scene; The system utilizes the obtained first-importance sampling strategy and generates unknown dangerous scenarios for autonomous driving based on environmental perception information.
[0007] This application also provides an electronic device, including: a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the method for generating unknown dangerous scenarios for autonomous driving as described above by running the program in the memory.
[0008] This application embodiment includes obtaining a first importance sampling strategy; wherein, the first importance sampling strategy is an optimized importance sampling strategy obtained by imitating a second importance sampling strategy and simultaneously optimizing the criticality and adversarial nature of state actions; the second importance sampling strategy is an importance sampling strategy obtained based on deep reinforcement learning and adversarial learning for target type scenarios; using the obtained first importance sampling strategy and based on environmental perception information, an unknown dangerous scenario for autonomous driving is generated. Because a criticality and adversarial learning method based on deep reinforcement learning is adopted, the sampling strategy is optimized by simultaneously combining criticality and adversarial nature, thus improving the ability to generate target type scenarios.
[0009] Other features and advantages of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the application. Other advantages of this application can be realized and obtained by means of the embodiments described in the description and the accompanying drawings. Attached Figure Description
[0010] The accompanying drawings are used to provide an understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0011] Figure 1 This is a flowchart illustrating a method for generating unknown hazardous scenarios for autonomous driving, according to an embodiment of this application. Figure 2 This is a schematic diagram of the learning process for an adversarial degree in a related technology; Figure 3 This is a schematic diagram illustrating the learning process of adversarial degree and criticality in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the acquisition process of a first importance sampling strategy according to an embodiment of this application; Figure 5 This is a schematic diagram illustrating the acquisition process of another first importance sampling strategy according to an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an autonomous driving unknown dangerous scenario generation device according to an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0012] This application describes several embodiments, but these descriptions are exemplary and not limiting, and it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of the embodiments described herein. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with, or may replace, any feature or element of any other embodiment.
[0013] This application includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements disclosed in this application can also be combined with any conventional features or elements to form unique inventive solutions. Any feature or element of any embodiment can also be combined with features or elements from other inventive solutions to form another unique inventive solution. Therefore, it should be understood that any feature shown and / or discussed in this application can be implemented individually or in any suitable combination. Therefore, the embodiments are not limited except by the limitations imposed by the appended claims and their equivalents. Furthermore, various modifications and changes can be made within the scope of the appended claims.
[0014] Furthermore, in describing representative embodiments, the specification may have presented methods and / or processes as a specific sequence of steps. However, the method or process should not be limited to the specific order of steps described herein, to the extent that it does not depend on such a specific order. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims concerning the method and / or process should not be limited to the steps performed in the written order, and those skilled in the art will readily understand that these orders can be varied and still remain within the spirit and scope of the embodiments of this application.
[0015] Currently, autonomous vehicles still face safety challenges in complex driving environments. Traditional safety assessment methods struggle to comprehensively cover potential hazardous scenarios, limiting the accuracy of accident rate predictions. Therefore, efficiently identifying and assessing potentially dangerous scenarios that autonomous driving systems may encounter has become a significant challenge before the large-scale deployment of autonomous vehicles.
[0016] Because safety-critical scenarios, especially unknown hazardous scenarios, are extremely rare in natural driving environments, traditional Monte Carlo sampling-based testing methods struggle to efficiently collect accident samples. Furthermore, existing reinforcement learning methods are ineffective at generating unknown hazardous scenarios, impacting the reliability of autonomous driving safety testing. Therefore, effectively predicting and generating unknown hazardous scenarios has become a core challenge in the current field of autonomous driving safety testing.
[0017] Therefore, embodiments of this disclosure provide a method for generating unknown dangerous scenarios for autonomous driving, such as... Figure 1 As shown, it includes: Step 100: Obtain the first importance sampling strategy; wherein, the first importance sampling strategy is the optimized importance sampling strategy obtained by imitating the second importance sampling strategy and simultaneously optimizing the criticality of state and action and the adversarial degree; the second importance sampling strategy is the importance sampling strategy obtained by deep reinforcement learning and adversarial degree learning for the target type scene.
[0018] Step 110: Utilize the obtained first importance sampling strategy and generate unknown dangerous scenarios for autonomous driving based on environmental perception information.
[0019] The framework of reinforcement learning typically includes five basic elements: agent, environment, state, action, and reward. The agent learns the optimal policy through interaction with the environment. At each time step, it obtains the current state from the environment, selects an action according to the policy, and executes it. After execution, the environment provides a reward based on the agent's action and updates the state. The agent's goal is to maximize the cumulative reward through trial and error and policy optimization. The core challenge of reinforcement learning lies in balancing exploration (trying new actions) with exploitation (selecting the best existing actions) and continuously improving the policy in a dynamic environment.
[0020] Dense deep reinforcement learning is a variant of reinforcement learning that typically refers to training agents in large-scale, complex environments using dense reward signals to guide the learning process. In this approach, the agent can acquire more information through frequent feedback, thereby accelerating learning and optimizing decision-making strategies. By combining deep neural networks, dense deep reinforcement learning can handle higher-dimensional state spaces and complex environmental tasks. Dense deep reinforcement learning has already achieved significant results in the field of autonomous driving.
[0021] In the field of autonomous driving, the testing metrics for autonomous vehicles (e.g., accident rate) are denoted as... ,in Indicates an event of interest (e.g., a collision between vehicles). Indicates the test scenario The natural probability distribution, For the event Indicator functions, Monte Carlo sampling is a numerical computation method based on random sampling. Its principle is to randomly generate a large number of sample points and then use statistical analysis of these samples to approximate the solution to the problem. For a random variable... Expected value The estimation formula for the Monte Carlo method is: (1) in, From The samples are independently and randomly drawn from a probability distribution. Based on the Monte Carlo sampling principle, the test index can be estimated as: (2) in, Independent and identically probable distributions The testing scenario for autonomous vehicles is that the probability of accidents is extremely low, which means that Monte Carlo sampling-based testing methods require a very large number of tests to obtain an estimate of the accident rate. The time and economic costs of conducting a large number of tests are unacceptable.
[0022] Importance sampling methods can alleviate this problem. Their basic idea is to increase the sampling probability of critical scenarios prone to errors, thereby improving testing efficiency. However, importance sampling methods suffer from the curse of dimensionality, making them unsuitable for high-dimensional testing environments. Therefore, sparse adversarial sampling was proposed, which only samples the critical variables in the test scenario based on their importance, while other variables follow their original natural probability distribution. Based on sparse adversarial sampling, the test metric can be estimated as follows: (3) in, The importance sampling function constructed for the sparse adversarial sampling method is defined as: The update strategy after importance sampling It is the initial state distribution. It is related to autonomous vehicle strategy The relevant state transition probabilities. Scenario It can be defined as: in Indicates time state, It is in time The actions of autonomous vehicles, It is a time span. It is a collection of all feasible scenarios.
[0023] In adversarial learning based on dense deep reinforcement learning, the importance sampling strategy for key states is as follows: (4) in, For action, For state, Basic sampling strategy, For the state value function, For the degree of resistance, Keyness. State-action keyness is obtained based on surrogate model training. Utilizing intensive deep reinforcement learning to counter adversarial systems Training is then performed. Specifically, this method uses a proximal policy optimization algorithm based on dense deep reinforcement learning to train the adversarial system. Its network structure can be as follows Figure 2 As shown: In test scenario generation based on intensive deep reinforcement learning, the safety verification of autonomous driving systems can be optimized by generating highly efficient key test scenarios. Intensive deep reinforcement learning utilizes proximal policy optimization algorithms to train adversarial algorithms. This allows for dynamic adjustment of the importance sampling strategy. This increases the sampling probability in key scenarios while maintaining policy stability.
[0024] However, existing scene generation methods based on dense deep reinforcement learning only involve learning adversarial relationships. When facing unknown and dangerous scenes, the criticality often approaches zero, failing to effectively reflect the importance of the scene. Relying solely on adversarial learning is insufficient to effectively guide scene discovery and evaluation, and because criticality is not considered... The non-optimality of this could affect the importance sampling strategy. This reduces the efficiency of accelerated testing.
[0025] To address this issue, this application proposes a joint learning method for criticality and adversarial degree based on dense deep reinforcement learning in its method for generating unknown dangerous scenarios for autonomous driving. Simultaneously, it optimizes... and ,like Figure 3 As shown. Its core idea is: (1) Keyness learning: Identifying which state-action pairs Most importantly, it can trigger incidents; (2) Adversarial learning: Construct the most challenging test scenarios to improve the effectiveness of testing. In the adversarial learning method based on dense deep reinforcement learning, adversarial degree Measuring a state sAdversarial testing. Based on an importance sampling strategy, the sampling distribution of test scenarios can be optimized to improve testing efficiency and reduce catastrophic failures.
[0026] The method for generating unknown hazardous scenarios for autonomous driving provided in this application embodiment obtains a first importance sampling strategy. The first importance sampling strategy is an optimized importance sampling strategy obtained by imitating and learning a second importance sampling strategy, simultaneously optimizing state / action criticality and adversarial factors. The second importance sampling strategy is an importance sampling strategy obtained based on deep reinforcement learning and adversarial learning for target-type scenarios. The obtained first importance sampling strategy is used to generate unknown hazardous scenarios for autonomous driving based on environmental perception information. Because a joint learning method for criticality and adversarial factors based on deep reinforcement learning is adopted, the sampling strategy is optimized by combining criticality and adversarial factors, thus optimizing the sample distribution, significantly improving testing efficiency, reducing computational costs, and enhancing the ability to generate target-type scenarios.
[0027] Existing scene generation methods based on dense deep reinforcement learning only involve learning adversarial relationships. However, when facing unknown and dangerous scenarios, the criticality often approaches zero, failing to effectively reflect the importance of the scenario. Relying solely on adversarial learning is insufficient to effectively guide scenario discovery and evaluation. The method for generating unknown and dangerous scenarios for autonomous driving provided in this application proposes a method based on dense deep reinforcement learning. It employs joint learning of criticality and adversarial relationships, increasing the density of training data and thus optimizing accident rate estimation. Furthermore, by combining defensive importance sampling, it effectively improves the efficiency and accuracy of autonomous driving testing, providing a feasible accelerated testing scheme for the safety assessment of autonomous vehicles.
[0028] In one exemplary instance, such as Figure 4 As shown, the first importance sampling strategy is obtained in the following way: Step 200: Initialize the imitation learning strategy; wherein, the imitation learning strategy is used to imitate the second importance sampling strategy, and both the imitation learning strategy and the second importance sampling strategy include: criticality and adversarial degree; Step 210: Construct a first reward function; wherein the first reward function is used to reward the criticality in the imitation learning strategy that approximates the criticality in the second importance sampling strategy, and the adversarial degree in the imitation learning strategy approximates the adversarial degree in the second importance sampling strategy; Step 220: Obtain multiple initial states from the state space as the first initial state. The imitation learning policy is updated using the obtained first initial states and the constructed first reward function to obtain the third importance sampling policy; Step 230: Obtain the first importance sampling strategy based on the obtained third importance sampling strategy.
[0029] In practical applications, multiple initial states can be randomly selected from the state space as the first initial state.
[0030] In one exemplary instance, obtaining a first importance sampling strategy based on the obtained third importance sampling strategy includes: The obtained third-importance sampling strategy is used as the first-importance sampling strategy; or, The obtained third-importance sampling strategy is then retrained, and the trained third-importance sampling strategy is used as the first-importance sampling strategy.
[0031] There are two methods to obtain the first importance sampling strategy based on the obtained third importance sampling strategy: one is to directly use the obtained third importance sampling strategy as the first importance sampling strategy, and the other is to retrain the obtained third importance sampling strategy to obtain the first importance sampling strategy.
[0032] In one exemplary instance, the imitation learning policy is updated using multiple obtained first initial states and a constructed first reward function to obtain a third importance sampling policy, including: From the multiple first initial states obtained, arbitrarily select a first initial state, use the initialized imitation learning policy as the imitation learning policy in the selected first initial state, and perform the following first policy update operation; The first strategy update operation includes: The imitation learning policy in the selected first initial state is iteratively updated according to the first reward function to obtain the final updated imitation learning policy in the selected first initial state. If the final updated imitation learning policy converges, the final updated imitation learning policy obtained in the selected first initial state is used as the third importance sampling policy. If the final updated imitation learning policy does not converge, another first initial state is selected from the multiple obtained first initial states, and the final updated imitation learning policy obtained in the previously selected first initial state is used as the imitation learning policy in the newly selected first initial state. The first policy update operation continues until the updated imitation learning policy obtained in the newly selected first initial state converges, and the final updated imitation learning policy obtained in the newly selected first initial state is used as the third importance sampling policy.
[0033] The final updated imitation learning strategy obtained under the selected first initial state refers to the final imitation learning strategy obtained by continuously updating the imitation learning strategy under different times using the selected first initial state.
[0034] If the imitation learning policy obtained using the selected first initial state converges, no other first initial states will be used to update the imitation learning policy. However, if the imitation learning policy obtained using the selected first initial state does not converge, a new first initial state needs to be selected, and the imitation learning policy will continue to be updated using the new selected first initial state.
[0035] In one exemplary instance, the imitation learning policy in the selected first initial state is iteratively updated according to a first reward function to obtain the final updated imitation learning policy in the selected first initial state, including: The initial time in the selected first initial state is taken as the current time in the current first initial state, the imitation learning policy in the selected first initial state is taken as the imitation learning policy in the current time in the selected first initial state, and the following second policy update operation is performed. The second strategy update operation includes: Based on the selected first initial state and the current time step's imitation learning policy, actions are sampled and executed from the action space to transition the selected first initial state from the current time step to the next time step. The reward value of the first reward function for this state transition process is calculated. The reward value of the first reward function for this state transition process is optimized, and the state-action criticality and adversarial degree in the current time step's imitation learning policy are updated in reverse based on the optimized reward value of the first reward function to obtain the updated imitation learning policy. The next time step replaces the current time step as the new current time step, and the updated imitation learning policy is used as the new imitation learning policy for the current time step under the selected first initial state. The second policy update operation is continued until the updated imitation learning policy has converged or the current time step is the maximum time step.
[0036] In one exemplary instance, the first reward function is designed as follows: ; in, Let be the first reward function at time t. and These represent the criticality and adversarial nature of the state-action pair in the second importance sampling strategy. and Let $t$ represent the state-action criticality and adversarial degree in the imitation learning strategy at time $t$. $1$ represents the 1-norm.
[0037] In one exemplary instance, the obtained third-importance sampling strategy is retrained, and the trained third-importance sampling strategy is used as the first-importance sampling strategy, such as... Figure 5 As shown, it includes: Step 300: Construct a second reward function; wherein, the second reward function is used to reward the third importance sampling strategy for fast convergence; Step 310: Obtain multiple initial states from the state space as second initial states. ; Step 320: Randomly select one second initial state from the multiple obtained second initial states, use the third importance sampling strategy as the strategy in the selected second initial state, and perform the following third strategy update operation; The third strategy update operation includes: The third importance sampling strategy in the selected second initial state is iteratively updated according to the second reward function to obtain the final updated third importance sampling strategy in the selected second initial state. If the final updated third importance sampling strategy converges, the strategy obtained in the selected first initial state and applied to the second initial state is... The updated third importance sampling strategy is used as the first importance sampling strategy. If the final updated third importance sampling strategy fails to converge, another second initial state is selected from the multiple obtained second initial states. The final updated third importance sampling strategy obtained in the previously selected second initial state is used as the updated third importance sampling strategy in the newly selected second initial state. The third strategy update operation continues until the updated third importance sampling strategy obtained in the newly selected second initial state converges. The final updated third strategy update operation obtained in the newly selected first initial state is then used as the first importance sampling strategy.
[0038] In practical applications, multiple initial states can be randomly selected from the state space as second initial states.
[0039] The final updated imitation learning strategy obtained under the selected second initial state refers to the final imitation learning strategy obtained by continuously updating the imitation learning strategy under different times using the selected second initial state.
[0040] If the imitation learning policy obtained using the selected second initial state converges, no other second initial states will be used to update the imitation learning policy. However, if the imitation learning policy obtained using the selected second initial state does not converge, a new second initial state needs to be selected, and the imitation learning policy will continue to be updated using the new selected second initial state.
[0041] In one exemplary instance, the third importance sampling policy in the selected second initial state is iteratively updated according to the second reward function to obtain the finally updated third importance sampling policy in the selected second initial state, including: The initial time of the selected second initial state is taken as the current time of the current second initial state, the third importance sampling strategy of the selected second initial state is taken as the third importance sampling strategy of the current time of the selected second initial state, and the following fourth strategy update operation is performed. The fourth strategy update operation includes: Based on the selected second initial state and the third importance sampling policy at the current time, actions are sampled from the action space and executed to transition the selected second initial state from the current time to the next time. The reward value of the second reward function for this state transition process is calculated. The reward value of the obtained second reward function for this state transition process is optimized, and the state-action criticality and adversarial degree in the third importance sampling policy at the current time are updated in reverse based on the optimized reward value of the second reward function to obtain the updated third importance sampling policy. The next time is replaced by the current time as the new current time, and the updated third importance sampling policy is used as the third importance sampling policy for the new current time under the selected second initial state. The fourth policy update operation is continued until the updated third importance sampling policy has converged or the current time is the maximum time.
[0042] In one exemplary instance, the second reward function is designed as follows: ; ,and ; , ; in, Let be the second reward function at time t. , These represent the criticality and adversarial nature of the state-action pair in the first-importance sampling strategy at time t. The first importance sampling strategy at time t, The third most important sampling strategy at time t, Let be the state value function at time t. The weight of the third most important sample at time t.
[0043] , .
[0044] The model training process involved in the method for generating unknown dangerous scenarios for autonomous driving provided in this embodiment is divided into two stages: Phase 1: Imitation Learning (IL) Imitation learning is performed on the importance sampling policy obtained by the adversarial learning method based on dense deep reinforcement learning. Therefore, the reward function is set as follows: in, and This refers to the state-action criticality and adversarial degree in adversarial learning methods based on dense deep reinforcement learning.
[0045] Phase Two: Importance Sampling Strategy Optimization The importance sampling strategy obtained from Phase 1 training is further trained to improve the performance of accelerated testing. Therefore, the reward function is set as follows: ; ,and ; , ; in, Let be the second reward function at time t. , These represent the criticality and adversarial nature of the state-action pair in the first-importance sampling strategy at time t. The first importance sampling strategy at time t, The third most important sampling strategy at time t, Let be the state value function at time t. The weight of the third most important sample at time t.
[0046] This ensures effective training on critical and adversarial states during reinforcement learning, improving testing efficiency.
[0047] The method for generating unknown dangerous scenarios for autonomous driving provided in this application is a joint learning method for criticality and adversarial degree based on dense deep reinforcement learning for generating unknown dangerous scenarios. This method makes the key information in the test data more concentrated and avoids the data sparsity problem in related methods. At the same time, it combines defensive importance sampling to optimize the sample distribution, significantly improves testing efficiency and reduces computational costs.
[0048] The method for generating unknown dangerous scenarios for autonomous driving provided in this application includes the following steps for implementing a criticality and adversarial learning method based on dense deep reinforcement learning for generating unknown dangerous scenarios: 1. Obtain the natural probability distribution of the driving environment , (in table form), recording the state space Action space The longest test time is .
[0049] 2. Collect training data: a collection of all state transition arrays. and the set of state transition arrays for key states. 2.1 Constructing a behavioral strategy for collecting data (controlling the driving behavior of autonomous vehicles): in, .
[0050] 2.2. Set the reward function as follows: in The weights representing the importance sampling, and 2.3, for ( (Maximum number of tests) 2.3.1 Initial sampling state ; 2.3.2, for ( (For the longest test time) a. Based on behavioral strategies Sampling action ; b. Execution of actions Update the next status Calculate and record the reward for the state. ; c. Judgment Is it a critical state? If so, then transfer the state transition array. Recorded at the same time and Otherwise, only record in middle; d. Judgment Is it an accident state? If so, break out of the loop. 2.3.3 Finally, we obtain the set of all state transition arrays. and the set of state transition arrays for key states. .
[0051] 3. Based on intensive deep reinforcement learning training : Randomly initialized adversarial neural network Set the maximum number of training rounds. Each batch is set to have a batch size B. The data is divided into batches of B data, and the last batch may contain fewer than B data.
[0052] for epoch=1, 2…., N: for batch (each batch contains B state transition arrays) in : 3.1. Based on the rewards for each step Calculate the advantage function ; 3.2 Calculate the loss of the Proximal Policy Optimization (PPO) algorithm: in, It is a strategic loss. It is a value function loss. It is an entropy reward. It is the probability ratio of the new strategy to the old strategy. This refers to the clipping range of the probability ratio, i.e., the clipping range, which is determined by the parameters. Limiting the upper and lower limits of the probability ratio between the new and old strategies prevents the strategy performance from becoming unstable due to excessively large update steps. It is a value function estimate. It is an estimated return. It's a hyperparameter; 3.3 Optimize policy parameters using gradient descent method .
[0053] 4. Output .
[0054] 5. Importance sampling strategy after training ( , To learn by imitation: Randomly initialized imitation learning policy neural network ( , (All are outputs of the neural network), and the reward function for imitation learning is set as (where 1 represents the 1-norm): Set the batch data size B, and... The data is divided into batches of B data, and the last batch may contain fewer than B data. while Not convergent (convergence: sufficiently close to 0): for batch (each batch contains B state transition arrays) in : 5.1 Observation Strategy Network Output , ; 5.2 Calculate the corresponding ; 5.3. Based on the rewards for each step Calculate the advantage function ; 5.4 Calculate the loss of the Proximal Policy Optimization (PPO) algorithm (same as step 3 above). 5.5 Determining the Reward Function of the Policy Network Has it converged? 5.6 Optimize policy parameters using gradient descent. ; 5.7 Output the policy network obtained through imitation learning .
[0055] 6. Importance sampling strategy obtained during training Further training will be conducted. Set the batch data size B, and... The data is divided into batches of B data, and the last batch may contain fewer than B data.
[0056] for epoch=1, 2…., N: for batch (each batch contains B state transition arrays) in : 6.1 Based on the rewards for each step Calculate the advantage function ; 6.2 Calculate the loss of the Proximal Policy Optimization (PPO) algorithm (same as step 3 above); 6.3 Optimize policy parameters using gradient descent method .
[0057] 7. Output the final policy network. .
[0058] This application proposes a joint learning method for criticality and adversarial degree based on dense deep reinforcement learning for generating unknown dangerous scenarios. This method improves the ability to generate unknown dangerous scenarios, reduces the underestimation of the risks of unknown dangerous scenarios, optimizes sampling strategies, improves testing efficiency, and reduces the impact of sparsity catastrophism.
[0059] This application proposes a joint learning method for criticality and adversarial degree based on dense deep reinforcement learning for generating unknown hazardous scenarios in the field of autonomous driving safety testing. This method improves testing efficiency and reduces testing costs by optimizing the importance sampling strategy to mitigate the catastrophic impact of sparsity in testing, thereby making autonomous driving safety assessments more efficient, robust, and accurate.
[0060] Corresponding to the above-mentioned method for generating unknown dangerous scenarios for autonomous driving, this application also provides an apparatus for generating unknown dangerous scenarios for autonomous driving. Figure 6 This is a schematic diagram of the structure of a device for generating unknown dangerous scenarios for autonomous driving, provided in an embodiment of this application. Figure 6 As shown, the apparatus for generating unknown dangerous scenarios for autonomous driving provided in this application embodiment includes: The policy acquisition unit 400 is used to acquire a first importance sampling policy; wherein, the first importance sampling policy is an optimized importance sampling policy obtained by imitating the second importance sampling policy and simultaneously optimizing the state-action criticality and adversarial degree; the second importance sampling policy is an importance sampling policy obtained based on deep reinforcement learning and adversarial degree learning for the target type scene; The scene generation unit 410 is used to generate unknown dangerous scenarios for autonomous driving by utilizing the obtained first importance sampling strategy and based on environmental perception information.
[0061] In one exemplary instance, the apparatus for generating unknown dangerous scenarios for autonomous driving provided in this application embodiment further includes: a strategy generation unit 420, which is used to obtain a first importance sampling strategy in the following manner: The imitation learning strategy is initialized; wherein, the imitation learning strategy is used to imitate the second importance sampling strategy, and both the imitation learning strategy and the second importance sampling strategy include: criticality and adversarial degree; Construct a first reward function; wherein the first reward function is used to reward the criticality in the imitation learning strategy that approximates the criticality in the second importance sampling strategy, and the adversarial degree in the imitation learning strategy approximates the adversarial degree in the second importance sampling strategy; Obtain multiple initial states from the state space as the first initial state. The imitation learning policy is updated using the obtained first initial states and the constructed first reward function to obtain the third importance sampling policy; The first importance sampling strategy is derived from the obtained third importance sampling strategy.
[0062] In one exemplary instance, the policy generation unit 420 is used to: The obtained third-importance sampling strategy is used as the first-importance sampling strategy; or, The obtained third-importance sampling strategy is then retrained, and the trained third-importance sampling strategy is used as the first-importance sampling strategy.
[0063] In one exemplary instance, the policy generation unit 420 is used to: From the multiple first initial states obtained, arbitrarily select a first initial state, use the initialized imitation learning policy as the imitation learning policy in the selected first initial state, and perform the following first policy update operation; The first strategy update operation includes: The imitation learning policy in the selected first initial state is iteratively updated according to the first reward function to obtain the final updated imitation learning policy in the selected first initial state. If the final updated imitation learning policy converges, the final updated imitation learning policy obtained in the selected first initial state is used as the third importance sampling policy. If the final updated imitation learning policy does not converge, another first initial state is selected from the multiple obtained first initial states, and the final updated imitation learning policy obtained in the previously selected first initial state is used as the imitation learning policy in the newly selected first initial state. The first policy update operation continues until the updated imitation learning policy obtained in the newly selected first initial state converges, and the final updated imitation learning policy obtained in the newly selected first initial state is used as the third importance sampling policy.
[0064] In one exemplary instance, the policy generation unit 420 is used to: The initial time in the selected first initial state is taken as the current time in the current first initial state, the imitation learning policy in the selected first initial state is taken as the imitation learning policy in the current time in the selected first initial state, and the following second policy update operation is performed. The second strategy update operation includes: Based on the selected first initial state and the current time step's imitation learning policy, actions are sampled and executed from the action space to transition the selected first initial state from the current time step to the next time step. The reward value of the first reward function for this state transition process is calculated. The reward value of the first reward function for this state transition process is optimized, and the state-action criticality and adversarial degree in the current time step's imitation learning policy are updated in reverse based on the optimized reward value of the first reward function to obtain the updated imitation learning policy. The next time step replaces the current time step as the new current time step, and the updated imitation learning policy is used as the new imitation learning policy for the current time step under the selected first initial state. The second policy update operation is continued until the updated imitation learning policy has converged or the current time step is the maximum time step.
[0065] In one exemplary instance, the first reward function is designed as follows: ; in, Let be the first reward function at time t. and These represent the criticality and adversarial nature of the state-action pair in the second importance sampling strategy. and These represent the criticality and adversarial nature of the state-action relationship in the imitation learning strategy at time t.
[0066] In one exemplary instance, the policy generation unit 420 is used to: Construct a second reward function; whereby the second reward function is used to reward the third importance sampling strategy for fast convergence; Obtain multiple initial states from the state space as the second initial state. ; From the multiple obtained second initial states, arbitrarily select one second initial state, use the third importance sampling policy as the policy in the selected second initial state, and perform the following third policy update operation; The third strategy update operation includes: The third importance sampling strategy in the selected second initial state is iteratively updated according to the second reward function to obtain the final updated third importance sampling strategy in the selected second initial state. If the final updated third importance sampling strategy converges, the strategy obtained in the selected first initial state and applied to the second initial state is... The updated third importance sampling strategy is used as the first importance sampling strategy. If the final updated third importance sampling strategy fails to converge, another second initial state is selected from the multiple obtained second initial states. The final updated third importance sampling strategy obtained in the previously selected second initial state is used as the updated third importance sampling strategy in the newly selected second initial state. The third strategy update operation continues until the updated third importance sampling strategy obtained in the newly selected second initial state converges. The final updated third strategy update operation obtained in the newly selected first initial state is then used as the first importance sampling strategy.
[0067] In one exemplary instance, the policy generation unit 420 is used to: The initial time of the selected second initial state is taken as the current time of the current second initial state, the third importance sampling strategy of the selected second initial state is taken as the third importance sampling strategy of the current time of the selected second initial state, and the following fourth strategy update operation is performed. The fourth strategy update operation includes: Based on the selected second initial state and the third importance sampling policy at the current time, actions are sampled from the action space and executed to transition the selected second initial state from the current time to the next time. The reward value of the second reward function for this state transition process is calculated. The reward value of the obtained second reward function for this state transition process is optimized, and the state-action criticality and adversarial degree in the third importance sampling policy at the current time are updated in reverse based on the optimized reward value of the second reward function to obtain the updated third importance sampling policy. The next time is replaced by the current time as the new current time, and the updated third importance sampling policy is used as the third importance sampling policy for the new current time under the selected second initial state. The fourth policy update operation is continued until the updated third importance sampling policy has converged or the current time is the maximum time.
[0068] In one exemplary instance, the second reward function is designed as follows: ; ,and ; , ; in, Let be the second reward function at time t. , These represent the criticality and adversarial nature of the state-action pair in the first-importance sampling strategy at time t. The first importance sampling strategy at time t, The third most important sampling strategy at time t, Let be the state value function at time t.
[0069] The apparatus for generating unknown dangerous scenarios for autonomous driving provided in this embodiment belongs to the same concept as the method for generating unknown dangerous scenarios for autonomous driving provided in the above embodiments of this application. It can execute the method for generating unknown dangerous scenarios for autonomous driving provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method for generating unknown dangerous scenarios for autonomous driving. Technical details not described in detail in this embodiment can be found in the specific processing content of the method for generating unknown dangerous scenarios for autonomous driving provided in the above embodiments of this application, and will not be repeated here.
[0070] This application also provides an electronic device, such as... Figure 7 As shown, it includes: memory 500 and processor 510; The memory 500 is connected to the processor 510 and is used to store programs; The processor 510 is used to implement the method for generating unknown dangerous scenarios for autonomous driving as described in any of the above embodiments by running the program in the memory 500.
[0071] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 520, an input device 530, and an output device 540.
[0072] The processor 510, memory 500, communication interface 520, input device 530, and output device 540 are interconnected via a bus. Among them: A bus can include a pathway for transmitting information between various components of a computer system.
[0073] The processor 510 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0074] The processor 510 may include a main processor, as well as a baseband chip, modem, etc.
[0075] The memory 500 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 500 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0076] Input device 530 may include a device for receiving data and information input by a user, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.
[0077] Output device 540 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0078] The communication interface 520 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0079] The processor 510 executes the program stored in the memory 500 and calls other devices, which can be used to implement each step of any of the autonomous driving unknown dangerous scenarios provided in the above embodiments of this application.
[0080] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the method for generating unknown hazardous scenarios for autonomous driving according to various embodiments of this application as described in any of the foregoing embodiments of this specification.
[0081] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0082] Furthermore, this application also provides a storage medium storing a computer program, which, when executed by a processor, implements the method for generating unknown dangerous scenarios for autonomous driving as described in any of the above embodiments.
[0083] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term "computer storage medium" includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
Claims
1. A method for generating autonomous driving scenarios, characterized in that, include: The imitation learning strategy is initialized; wherein the imitation learning strategy is used to imitate the second importance sampling strategy, and both the imitation learning strategy and the second importance sampling strategy include: criticality and adversarial degree; Construct a first reward function; the first reward function is used to reward the criticality in the imitation learning strategy that approximates the criticality in the second importance sampling strategy, and the adversarial degree in the imitation learning strategy approximates the adversarial degree in the second importance sampling strategy; Obtain multiple initial states from the state space as the first initial state. From the multiple first initial states obtained, arbitrarily select a first initial state, use the initialized imitation learning policy as the imitation learning policy in the selected first initial state, and perform the following first policy update operation to obtain the first importance sampling policy; The first importance sampling strategy is derived from the obtained third importance sampling strategy; The obtained first-importance sampling strategy is used to generate autonomous driving scenarios based on environmental perception information; The first policy update operation includes: The initial time of the selected first initial state is taken as the current time of the current first initial state, and the imitation learning policy of the selected first initial state is taken as the imitation learning policy of the current time of the selected first initial state. The following second policy update operation is performed to obtain the final updated imitation learning policy of the selected first initial state. If the final updated imitation learning policy converges, the final updated imitation learning policy obtained in the selected first initial state is taken as the third importance sampling policy. If the final updated imitation learning policy does not converge, another first initial state is selected from the multiple obtained first initial states. The final updated imitation learning policy obtained in the previously selected first initial state is taken as the imitation learning policy of the newly selected first initial state. The first policy update operation is continued until the updated imitation learning policy obtained in the newly selected first initial state converges. The final updated imitation learning policy obtained in the newly selected first initial state is taken as the third importance sampling policy. The second strategy update operation includes: Based on the selected first initial state and the current time step's imitation learning policy, actions are sampled and executed from the action space. The selected first initial state is then transitioned from the current time step to the next time step, and the reward value of the first reward function for this state transition process is calculated. The reward value of the first reward function for this state transition process is optimized, and the state-action criticality and adversarial degree in the current time step's imitation learning policy are updated in reverse based on the optimized reward value of the first reward function, resulting in an updated imitation learning policy. The next time step replaces the current time step as the new current time step, and the updated imitation learning policy is used as the new imitation learning policy for the current time step under the selected first initial state. The second policy update operation is continued until the updated imitation learning policy has converged or the current time step is the maximum time step.
2. The method according to claim 1, characterized in that, The step of obtaining the first importance sampling strategy based on the obtained third importance sampling strategy includes: The obtained third importance sampling strategy is used as the first importance sampling strategy; or, The obtained third importance sampling strategy is then trained again, and the trained third importance sampling strategy is used as the first importance sampling strategy.
3. The method according to claim 1, characterized in that, The first reward function is designed as follows: ; in, Let be the first reward function at time t. and These represent the criticality and adversarial nature of the state-action pair in the second importance sampling strategy. and These represent the criticality and adversarial nature of the state-action relationship in the imitation learning strategy at time t.
4. The method according to claim 2, characterized in that, The step of retraining the obtained third importance sampling strategy and using the trained third importance sampling strategy as the first importance sampling strategy includes: Construct a second reward function; wherein the second reward function is used to reward the third importance sampling strategy for fast convergence; Obtain multiple initial states from the state space as the second initial state. ; From the multiple obtained second initial states, arbitrarily select one second initial state, use the third importance sampling policy as the policy in the selected second initial state, and perform the following third policy update operation; The third strategy update operation includes: The third importance sampling strategy in the selected second initial state is iteratively updated according to the second reward function to obtain the final updated third importance sampling strategy in the selected second initial state. If the final updated third importance sampling strategy converges, the strategy obtained in the selected first initial state and applied to the second initial state is... The updated third importance sampling strategy is used as the first importance sampling strategy. If the final updated third importance sampling strategy fails to converge, another second initial state is selected from the multiple obtained second initial states. The final updated third importance sampling strategy obtained in the previously selected second initial state is used as the updated third importance sampling strategy in the newly selected second initial state. The third strategy update operation continues to be executed until the updated third importance sampling strategy obtained in the newly selected second initial state converges. The final updated third strategy update operation obtained in the newly selected first initial state is then used as the first importance sampling strategy.
5. The method according to claim 4, characterized in that, The step of iteratively updating the third importance sampling strategy in the selected second initial state according to the second reward function to obtain the finally updated third importance sampling strategy in the selected second initial state includes: The initial time of the selected second initial state is taken as the current time of the current second initial state, the third importance sampling strategy of the selected second initial state is taken as the third importance sampling strategy of the current time of the selected second initial state, and the following fourth strategy update operation is performed. The fourth strategy update operation includes: Based on the selected second initial state and the third importance sampling strategy at the current time, actions are sampled from the action space and executed to transition the selected second initial state from the current time to the next time. The reward value of the second reward function for this state transition process is calculated. The reward value of the obtained second reward function for this state transition process is optimized, and the state-action criticality and adversarial degree in the third importance sampling strategy at the current time are updated in reverse based on the optimized reward value of the second reward function to obtain the updated third importance sampling strategy. The next time is replaced by the current time as the new current time, and the updated third importance sampling strategy is used as the third importance sampling strategy for the new current time under the selected second initial state. The fourth strategy update operation is continued until the updated third importance sampling strategy has converged or the current time is the maximum time.
6. The method according to claim 4 or 5, characterized in that, The second reward function is designed as follows: ; ,and , ; in, Let be the second reward function at time t. , These represent the criticality and adversarial nature of the state-action pair in the first-importance sampling strategy at time t. The first importance sampling strategy at time t, The third most important sampling strategy at time t, Let be the state value function at time t. The weight of the third most important sample at time t.
7. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method for generating autonomous driving scenarios as described in any one of claims 1-6 by running a program in the memory.
Citation Information
Patent Citations
Deep reinforcement learning strategy protection defense method and device based on imitation learning and attention mechanism
CN113255936A
Multi-agent confrontation scene situation assessment method and device based on field energy
CN114492059A