Multi-agent-based interactive automatic driving simulation test scene generation method
The method uses multi-agent reinforcement learning to enhance simulation testing by creating interactive driving strategies for background vehicles, addressing the lack of dynamic interaction in existing methods and improving the realism and efficiency of automatic driving simulations.
Patent Information
- Application Number
- CN202510471420.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
AI Technical Summary
The existing autonomous driving simulation test scenarios lack dynamic interactivity, cannot fully simulate complex traffic environments, and rely on large amounts of data and high computing costs.
A multi-agent reinforcement learning framework is adopted to design interactive driving strategies for background vehicles, and a background vehicle is trained through Level-K game theory and deep neural networks to achieve dynamic interaction with the vehicle being tested, and a multi-level interactive scenario is generated.
It improves the interactivity and complexity of the test scenarios, reduces data dependence and computing costs, generates more realistic dynamic traffic scenarios, and improves the testing efficiency and effectiveness of the autonomous driving system.
Smart Images

Figure CN120316010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving, and particularly to a method for generating an interactive autonomous driving simulation test scenario based on multi - agents. Background Art
[0002] With the continuous development of autonomous driving technology, testing based on simulation scenarios has become one of the core means to verify and evaluate the functions and performance of autonomous driving systems. Existing simulation test methods mainly rely on static data sets, preset trajectories, or background vehicle driving models with fixed rules to construct test scenarios. However, although these methods can evaluate the behavior and performance of autonomous driving systems to a certain extent, due to the lack of dynamic interactivity of background vehicles, test scenarios usually cannot fully simulate the dynamic changes in complex traffic environments, thus limiting the authenticity and effectiveness of test results.
[0003] In addition, dynamic interactive test scenarios can better simulate the interaction characteristics in real - world traffic environments by adjusting the behavior of background traffic participants in real time to generate a continuous and complex - changing test environment, so as to more comprehensively evaluate the performance of autonomous driving vehicles under complex traffic conditions. However, existing methods for constructing dynamic interactive scenarios still have limitations in the number of interactive background vehicles, and the lack of effective simulation of interactive test scenarios also restricts their application effects in autonomous driving tests. Summary of the Invention
[0004] In a first aspect, the present application provides a method for generating an interactive autonomous driving simulation test scenario, which designs an interactive driving strategy for background vehicles, enabling the background vehicles to interact with the vehicle under test according to the real - time simulation environment.
[0005] The method includes: constructing a dynamic driving model based on image data and metric data for simulating the driving decision-making process of a vehicle as a background vehicle, where the image data includes images collected by the vehicle, the metric data includes the state information of the vehicle and the current action commands of the vehicle, the dynamic driving model generates action information and value information through an image encoder, a metric encoder, and at least one neural network model, the action information is used to indicate control actions corresponding to the driving state of the vehicle, and the value information is used to evaluate the long-term effect of the current driving decision of the vehicle; determining and training driving strategies at multiple interaction levels based on the hierarchical training idea and the dynamic driving model, where the vehicle as the background vehicle is trained in the training environments at the multiple interaction levels and is controlled by a fixed rule model or the dynamic driving model to obtain the driving strategies at the multiple interaction levels; constructing a simulation test environment including the vehicle under test and the vehicle as the background vehicle; configuring the driving strategies at the multiple interaction levels for the vehicle as the background vehicle based on the dynamic driving model and the driving strategies at the multiple levels, and jointly training the vehicle as the background vehicle and the vehicle under test based on the multi-agent reinforcement learning algorithm to generate an autonomous driving simulation test scenario with interaction between the background vehicle and the vehicle under test.
[0006] In a possible implementation manner of the first aspect in combination with the first aspect, the image encoder is a convolutional neural network and is used to extract image features from the image data to output a high-dimensional feature representation, and the metric encoder is used to generate a feature representation of the current state of the vehicle based on the metric data; the at least one neural network model includes a long short-term memory network LSTM for processing the time-series dependencies in driving decisions and a fully connected layer for outputting vehicle control commands and action predictions based on the fused features.
[0007] In a possible implementation manner of the first aspect in combination with the first aspect, the state information of the vehicle includes one or more of the vehicle's own speed, distance from the vehicle in front, relative speed, and current position, and the current action commands of the vehicle include one or more of the throttle, brake, and steering wheel angle.
[0008] In a possible implementation manner of the first aspect in combination with the first aspect, the driving strategies at multiple interaction levels are generated based on game theory and a reinforcement learning framework.
[0009] In a possible implementation manner of the first aspect in combination with the first aspect, the driving strategies at multiple interaction levels are generated based on the Level-K game theory, where the multiple interaction levels include: non-interactive where the vehicle as the background vehicle drives based on fixed rules; semi-interactive where the vehicle as the background vehicle gradually introduces the dynamic driving model through an alternating control method; fully interactive where the vehicle as the background vehicle is completely controlled by the dynamic driving model.
[0010] In combination with the first aspect, in a possible implementation manner of the first aspect, the construction of the simulation test environment of the vehicle including the vehicle under test and the vehicle as the background vehicle includes: configuring the interaction level ratio of the background vehicle according to the test target; setting the static scene parameters in the simulation test environment, and the static scene parameters include road conditions, weather, and traffic signals.
[0011] In combination with the first aspect, in a possible implementation manner of the first aspect, based on the multi-agent reinforcement learning algorithm, the vehicle as the background vehicle is jointly trained, including: initializing the Actor-Critic network and the target network for each background vehicle; storing interaction data through the experience replay pool and sampling small batches of samples to optimize the network parameters; guiding policy optimization through the reward function.
[0012] In combination with the first aspect, in a possible implementation manner of the first aspect, the reward function includes driving speed, vehicle spacing, events, and distance to the destination.
[0013] Through the above method, an interactive driving strategy is designed for the background vehicle, enabling the background vehicle to interact with the vehicle under test according to the real-time simulation environment. Moreover, driving strategies for background vehicles with multiple interaction levels are constructed (such as non-interactive, semi-interactive, fully interactive). This adjustable interactivity enables the simulation of various complex traffic scenarios. In addition, the multi-agent framework based on reinforcement learning has low dependence on data and good interpretability. During the simulation process, by self-training the decision-making strategy of the background vehicle, it is possible to generate more flexible and diverse interaction scenarios while achieving low data dependence. This makes the test process more efficient, not only reducing the need for large-scale data sets but also significantly reducing the computing and training costs.
[0014] In the second aspect, an interactive autonomous driving simulation test system is provided, including: a dynamic driving model construction module for generating a dynamic driving model that simulates the driving decision-making process of the vehicle as the background vehicle; a hierarchical driving strategy training module for determining and training driving strategies with different interaction levels; a multi-agent reinforcement learning module for training the vehicle as the background vehicle in training environments with multiple interaction levels, and the vehicle as the background vehicle is controlled by a fixed rule model or the dynamic driving model; an interactive test scenario generation module for configuring and generating an interactive autonomous driving simulation test scenario.
[0015] In the third aspect, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the method described in the first aspect or any one of the first aspect is implemented. Description of the Drawings
[0016] Figure 1Module diagram of the interactive autonomous driving simulation test scenario generation method provided by the embodiments of this application;
[0017] Figure 2 Schematic diagram of a dynamic driving model provided by the embodiments of this application;
[0018] Figure 3 Schematic diagram of the module based on the hierarchical driving strategy training module provided by the embodiments of this invention;
[0019] Figure 4 Technical framework diagram of the interactive autonomous driving simulation scenario generation method based on multi-agent reinforcement learning provided by the embodiments of this application;
[0020] Figure 5 Another technical framework diagram of the interactive autonomous driving simulation scenario generation method based on multi-agent reinforcement learning provided by the embodiments of this application;
[0021] Figure 6 Schematic diagram of an experimental result provided by the embodiments of this application;
[0022] Figure 7 Schematic diagram of another experimental result provided by the embodiments of this application;
[0023] Figure 8 Schematic structural diagram of a communication device according to an embodiment of this application;
[0024] Figure 9 Schematic structural diagram of a communication device according to another embodiment of this application. Detailed implementation manners
[0025] To make the objectives, technical solutions and advantages of the embodiments of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. Obviously, the described embodiments are some but not all of the embodiments of this invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of this invention fall within the scope of protection of this invention.
[0026] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The terms "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationships will also change accordingly.
[0027] The background technology related to the embodiments of the present application and the existing technical problems are introduced below.
[0028] With the continuous development of autonomous driving technology, testing based on simulation scenarios has become one of the core means to verify and evaluate the functions and performance of autonomous driving systems. At present, simulation testing methods mainly rely on static data sets, preset trajectories or background vehicle driving models with fixed rules to construct test scenarios. That is to say, the background vehicle driving models mainly conduct simulation testing based on static or fixed methods. Exemplarily, at present, heuristic methods, methods based on deep learning, etc. are used to construct test scenarios. Although the current simulation testing methods relying on static data sets, preset trajectories or fixed rules can evaluate the behaviors and performance of autonomous driving systems to a certain extent, due to the lack of dynamic interactivity of background vehicles, test scenarios usually cannot fully simulate the dynamic changes in complex traffic environments, thus limiting the authenticity and effectiveness of test results.
[0029] Exemplarily, using heuristic methods (such as Bayesian optimization and hyperparameter search) can generate some special scenarios to test the reactions of autonomous driving systems in adversarial or dangerous situations. For example, Bayesian optimization is used to generate adversarial scenarios that can induce weaknesses in system performance, and hyperparameter search can be used to identify potential dangerous situations. However, these methods usually require preset parameter ranges, and the generated scenarios are relatively monotonous, lacking sufficient interactivity and complexity. Therefore, the test scenarios they generate are usually short and difficult to effectively reflect the performance of autonomous driving systems in real complex traffic environments.
[0030] Exemplarily, in recent years, imitation learning and deep learning technologies have also been widely applied to the generation of simulation test scenarios. For example, inverse reinforcement learning can extract features from natural driving data and generate trajectories similar to real human driving behaviors. However, although these methods can generate high-fidelity scenarios, they usually rely on a large amount of data for training, and the generated scenarios are often simple reproductions of natural driving trajectories, lacking true dynamic interactivity and complex traffic situations. In addition, deep learning methods such as generative adversarial networks and gated recurrent units are used to generate scenarios, but the limitation of these methods is that their scenario generation ability still depends on predefined and scripted settings, resulting in the behavior of background vehicles being unable to respond to the driving behavior of the vehicle under test in real time, thereby reducing the interactivity.
[0031] In addition, in order to improve the interactivity of scenarios, in recent years, some research has begun to focus on interaction-based test scenario generation, which is used to generate more interactive test scenarios closer to real human driving processes, such as generating anthropomorphic background vehicle behaviors through deep reinforcement learning. However, most of these methods focus on certain specific types of traffic scenarios and lack large-scale interactive background vehicle simulations, and cannot generate dynamic scenarios with multi-element interactions in complex traffic environments.
[0032] Based on the above background technology and existing problems, generally speaking, there is a technical problem of lack of interactivity of background vehicles in the existing technology.
[0033] On the one hand, most of the existing simulation test scenarios rely on static data sets and preset trajectories, and the behaviors of background vehicles are usually fixed, lacking real-time interaction responses to the vehicle under test. This design method cannot effectively simulate the changeable and complex situations in the real traffic environment, resulting in the test results being unable to fully reflect the performance of the autonomous driving system in dynamic traffic scenarios.
[0034] On the other hand, the existing simulation test scenarios lack the complexity of multi-agent interaction. Although the current dynamic interactive test scenarios can simulate some basic behaviors of background vehicles, most methods only focus on single-type scenarios or a small number of background vehicles, and it is difficult to achieve complex interactions between multiple agents. In traditional methods, the behaviors of multiple agents are usually difficult to coordinate or influence each other, and complex traffic situations cannot be simulated.
[0035] In addition, the simulation scenario generation methods based on deep learning (such as inverse reinforcement learning and generative adversarial networks) usually require a large amount of driving data for training, and the computational cost is relatively high. This is an important bottleneck for real-time simulation testing and large-scale scenario generation.
[0036] To solve the above technical problems, this application proposes an interactive autonomous driving simulation test scenario generation method, which includes four modules, namely S1 to S4. Through these four modules, the technical problem of insufficient interactivity in the existing autonomous driving simulation test scenario construction method can be solved.
[0037] Specifically, the current test scenarios generated based on static data sets and preset trajectories cannot achieve dynamic interaction between background vehicles and the vehicle under test, resulting in the constructed test scenarios lacking sufficient interactivity and being unable to truly reflect the performance of the autonomous driving system in complex traffic environments. Therefore, how to generate interactive test scenarios that can simulate more realistic driving environments, especially how to enhance the interaction ability between background vehicles and the vehicle under test, has become the core issue studied in this application.
[0038] To solve this problem, this application proposes an interactive autonomous driving simulation test scenario generation method based on multi-agent reinforcement learning. By designing an interactive driving strategy for background vehicles, it can simulate more complex and realistic dynamic traffic scenarios. Different from the traditional static data set and preset trajectory methods, this application can achieve real interaction between background vehicles and the vehicle under test by dynamically adjusting the behavior of background vehicles, and then construct a traffic test scenario with high interactivity and high complexity.
[0039] The core concept of this application is to generate more realistic and complex autonomous driving simulation test scenarios by combining multi-agent reinforcement learning with an interactive background vehicle driving strategy. Specifically, it includes the following parts:
[0040] 1. Design of the multi-agent reinforcement learning framework: The multi-agent reinforcement learning framework of the present invention simulates the interaction between background vehicles and the vehicle under test by constructing different levels of interactive driving strategy models for background vehicles. This framework is one of the core innovations of this application. It trains background vehicles through the multi-agent reinforcement learning algorithm, enabling them to have dynamic interaction with the vehicle under test. Through this framework, background vehicles can adjust their driving decisions in real time according to environmental changes, the behavior of the vehicle under test, and traffic conditions, simulating a more realistic traffic scenario. During training, through this framework, background vehicles do not just drive along a predetermined trajectory, but make interactive decisions with the vehicle under test through reinforcement learning.
[0041] Among them, the multi-agent reinforcement learning framework can be based on game theory (such as Level-K game theory, Stackelberg game theory, etc.), or it can adopt a deep reinforcement learning framework and use a deep neural network to replace the reinforcement learning model.
[0042] 2. Interactive driving strategies of background vehicles: In traditional testing methods, background vehicles usually move only relying on preset trajectories or fixed rules, resulting in insufficient response to the behavior of the vehicle under test. In this application, the background vehicles are trained through reinforcement learning, endowing them with the ability of "autonomous decision-making", enabling them to adjust their driving strategies and behaviors in real time according to the changes in the simulation environment. The background vehicles in the simulation are not just simple "bystanders", but interact with the vehicle under test through interactive behaviors, simulating more complex traffic situations. Furthermore, each background vehicle will have real-time dynamic interactions with the test vehicle according to different strategies at different interaction levels. By gradually increasing the complexity of the interaction strategy and the diversity of agent behaviors, a more realistic dynamic scenario can be generated in the simulation environment.
[0043] Exemplarily, in the embodiments of this application, this application can utilize the aforementioned game theories (such as the Level-K game theory, Stackelberg game theory, etc.) to design and train interaction strategies at different levels, enabling the background vehicles to gradually enhance the complexity and diversity of their decisions during the interaction with the vehicle under test. For example, the Level-K game theory is essentially a framework based on game strategy reasoning. Under this framework, the driving decisions of background vehicles not only depend on their own states, but also take into account the behaviors and strategies of other agents (such as the vehicle under test and other background vehicles). Through this game-based training method, the background vehicles can be made closer to real driving behaviors in the simulation test.
[0044] 3. Generation of interactive test scenarios: By continuously adjusting the interaction relationship between the background vehicles and the vehicle under test in the test scenario, a challenging and variable interactive traffic test scenario is generated. For example, through interaction strategies at different levels (such as non-interactive, semi-interactive, fully interactive), various traffic scenarios from simple to complex and from low-risk to high-risk can be simulated in the test. This flexible adjustment ability enables the autonomous driving system to be challenged in various dynamic interaction situations and better evaluate its ability to handle complex traffic conditions.
[0045] Furthermore, after generating the interactive test scenario, this application combines it with the existing autonomous driving closed-loop simulation platform for high-fidelity simulation testing. For example, by integrating with existing autonomous driving closed-loop simulation platforms (such as CARLA, SUMO, or VISSIM), the simulation environment can be updated in real time and present the complex interaction process between the background vehicles and the vehicle under test.
[0046] The specific embodiments of this application will be described below with reference to the accompanying drawings. Figure 1 The block diagram of the interactive autonomous driving simulation test scenario generation method provided by the embodiments of this application is shown. Among them, asFigure 1 As shown in the figure, it includes four modules S1 to S4.
[0047] S1: A module for building a vision-based dynamic driving model.
[0048] In the S1 module, the present application can build a dynamic driving model by a method based on a deep neural network, combining image processing and metric data to imitate the driving decision-making process of a human driver. Figure 2 The figure shows a schematic diagram of a dynamic driving model provided by an embodiment of the present application. As Figure 2 shown, the model consists of multiple modules, covering image encoding, metric data processing, time series processing, and final decision output, etc. The whole process can generate vehicle control commands in real time with a small computational overhead to achieve a dynamic and highly interactive autonomous driving scenario.
[0049] As Figure 2 shown, the image and metric related to the background vehicle are respectively input into the image encoder and the metric encoder.
[0050] Exemplarily, the image encoder can be a convolutional neural network for extracting features from the original images collected by in-vehicle sensors. The input of this module includes image data (for example, images collected by in-vehicle cameras). Different levels of image features are gradually extracted through multiple convolutional layers. The features are further compressed through pooling operations, and finally a high-dimensional feature representation is output. This high-dimensional feature representation is used as the input of the subsequent decision-making module to provide visual information support for the driving decision of the background vehicle.
[0051] Exemplarily again, the metric encoder can be responsible for processing the metric data from the vehicle itself and combining it with the image features. The metric information input to this module can include the state information of the vehicle (for example, the speed of the vehicle itself, the distance from the vehicle in front, the relative speed, the current position, etc.), and the current action commands (such as throttle, brake, steering wheel angle). Among them, the action command information can be encoded as a one-hot vector to distinguish different types of control signals. After being processed through multiple hidden layers, the metric information is transformed into a feature representation understandable by a deep learning model. Finally, a comprehensive feature representation containing the current state of the vehicle is output.
[0052] After that, as Figure 2 shown, the feature representations respectively output by the image encoder and the metric encoder will be fused and input into a neural network model, as Figure 2The long short-term memory (LSTM) model shown. In other words, this LSTM module receives the concatenated features from the image encoder and the metric encoder. Exemplarily, this LSTM module is used to process time series data to help the dynamic driving model consider historical state and action information when making decisions. Among them, the decisions in the autonomous driving scenario not only depend on the current observations but also need to refer to the past driving states. The LSTM module can retain and update the past state information through its gating mechanism, so that the model's decision at the current moment can take into account the past driving history. Finally, it outputs the fused features of the historical state information and the current input information.
[0053] Finally, as Figure 2 shown, the fused features of the historical state information and the current input information output by the neural network model are input into the fully connected (FC) layer, which further enables the fully connected layer to extract high-level features and perform non-linear mapping. This fully connected layer can include one or more layers. Exemplarily, the role of the FC layer can be to enhance the model's expressive ability so that it can handle complex decision-making tasks and output the final control commands and action predictions. Finally, the FC layer can output two main parts: the action head (or action information) generates the control actions corresponding to the current driving state; the value head (or value information) predicts the cumulative reward value in this state, which is used to evaluate the long-term effect of the current driving decision.
[0054] In some other embodiments of the present application, Figure 2 the image encoder shown can also be replaced with a lighter network structure such as MobileNe, the metric encoder can also be replaced with a network structure with stronger feature modeling ability such as MLP-Mixer, and LSTM can be replaced with GRU.
[0055] In addition, in some other embodiments of the present application, the parameters of the input image encoder and the metric encoder can also introduce multi-frame image sequences, depth maps, lidar data, or map information, etc., and the parameters output by the FC can also be extended to multi-step trajectory point prediction and the probability distribution of control signals.
[0056] In some other embodiments of the present application, the vision-based dynamic driving model can also be simplified to a certain extent, for example, without including the LSTM model, etc. In this way, although simplifying the vision recognition model will reduce the fidelity of the scenario, as long as it can effectively identify the interaction relationship between the background vehicle and the vehicle under test, it can still achieve the basic purpose of testing the autonomous driving system.
[0057] S2: Based on the hierarchical driving strategy training module.
[0058] Figure 3The module schematic diagram based on the hierarchical driving strategy training module is shown. Exemplarily, as Figure 3 shown, in this application, in order to improve the interactivity in the autonomous driving test scenario, a training framework can be adopted based on game theory. Through the hierarchical training idea, two different levels of training environments, semi-interactive and fully interactive, are constructed. This framework effectively guides the driving behavior of background vehicles, enabling them to conduct multi-agent joint training in the simulation environment, gradually enhancing the interactivity between agents, and finally generating multi-level driving strategies that can be used in interactive autonomous driving test scenarios. In the embodiments of this application, the Level-K game theory, Stackelberg game theory, etc. can be used to train the framework. The embodiments will be introduced below taking the Level-K game theory as an example.
[0059] Among them, the Level-K game theory is a game theory model that describes human decision-making behavior and can explain the thinking levels of humans in complex decision-making environments. This theory assumes that an individual's decision is based on the expectation of the behavior of other participants, and these expectations are usually not completely rational. In the Level-K theory, the thinking levels of decision-makers range from Level-0 to Level-K. In other words, this strategy divides background vehicles into different interaction levels (such as non-interactive, semi-interactive, and fully interactive corresponding to Level-0 to Level-2 respectively), and designs corresponding driving strategies for each level, which can simulate more complex and changeable traffic situations during the simulation process.
[0060] Those skilled in the art should be aware that if the Stackelberg game theory is introduced, the two types of hierarchical strategies, "leader-type driving strategy" and "responsive driving strategy", need to be used as the training objectives. Instead of dividing the interaction levels by "cognitive level", the strategy logic is set according to "role positioning". This application will not elaborate on this.
[0061] First, in the initialization phase, all agents (background vehicles and the vehicle under test) are initialized according to the Level-0 strategy. The Level-0 strategy can adopt reflexive decision-making. For example, the vehicle can make passive avoidance according to the distance interval, and the goal is only to reach the destination without considering the interaction with other vehicles.
[0062] After that, in the semi-interactive training environment (Level-1), the behavior of the vehicle can be alternately controlled by the autopilot mode in the simulation environment (such as CARLA) and the vision-based dynamic driving model. The Autopilot mode can provide basic autonomous driving control, while the dynamic driving model provides more complex behavior decisions at critical decision-making times (such as turning or avoidance). By gradually increasing the control weight of the dynamic driving model within a fixed time step, the interactivity of the background vehicle is gradually enhanced, and finally a semi-interactive driving strategy is generated.
[0063] Then, in the full-interaction training environment (Level-2), all vehicles no longer rely on the autopilot mode, but are completely controlled by a dynamic driving model equipped with a semi-interactive driving strategy. Vehicles will undergo multiple rounds of training through reinforcement learning in this environment to increase the interactivity between vehicles and the complexity of decision-making. The background vehicles in this environment will exhibit more complex driving behaviors and higher interactivity, thus generating a full-interactive driving strategy. Among them, as Figure 3 shown, the training process in the full-interaction training environment is similar to that in the semi-interactive training environment. Through a multi-agent reinforcement learning framework based on Level-K game theory, the interactivity of background vehicles in the autonomous driving simulation environment can be effectively improved, generating two different levels of driving strategies: semi-interactive and full-interactive.
[0064] In some other embodiments of the present application, if the full-interactive background vehicles are omitted, the basic performance of the autonomous driving vehicle can still be tested through non-interactive and semi-interactive background vehicles. That is, Figure 3 it may not include the full-interaction training environment. In this way, although the complexity of the scenario is reduced, effective testing can still be carried out in the basic interactive scenarios.
[0065] This hierarchical training method not only increases the synergy between agents, but also greatly improves the diversity and complexity of the test scenarios, thus providing a more realistic and efficient test environment for the verification of autonomous driving systems.
[0066] S3: Multi-agent reinforcement learning module.
[0067] The present application designs a multi-agent reinforcement learning scheme for the joint training of background vehicles. Compared with the reinforcement learning method for a single vehicle, multi-agent reinforcement learning can better handle the interaction between multiple agents and generate a more realistic and complex driving model. This scheme gradually improves the interactivity of agents through joint training in different training environments, enabling background vehicles to perform more complex driving behaviors in the simulation test scenarios.
[0068] Exemplarily, the present application can adopt a multi-agent reinforcement learning framework based on the proximal policy optimization (PPO) algorithm. As Figure 3 shown, in this framework, the background vehicles (agents) jointly conduct reinforcement learning training. For example, each agent can optimize its policy through the Actor-Critic structure and guide the learning process through the reward function, and finally form a driving strategy that can exhibit intelligent interaction in the simulation environment.
[0069] Exemplarily, four factors such as driving speed, vehicle distance, events, and destination distance can be considered in the reward function:
[0070] 1. Driving speed: According to the average driving speed of vehicles at intersections in real natural driving data, the vehicle speed in the training environment of this application is limited between 0 and 5.6 m / s, and the speed reward R speed is as shown in formula (1):
[0071] R speed = ω1·max(0, min(forward_speed, 5.6)) (1)
[0072] In the formula, forward_speed is the forward speed of the vehicle at the current moment; ω1 is the reward coefficient of this item.
[0073] 2. Vehicle distance: To ensure safety during driving, this application can keep a safety distance of more than 2 meters during driving. The vehicle distance reward R distance is as shown in formula (2):
[0074]
[0075] In the formula, distance represents the distance between the vehicle and the nearest vehicle at the current time
[0076] 3. Event factors: This part mainly considers two types of emergencies, namely collision and lane deviation. Penalties are imposed on such events, and the penalty R event is as shown in formula (3), where Δcollision and Δintersection are the change amounts of collision events and lane deviation events within every two adjacent time steps.
[0077] R event = ω3·Δcollision + ω4·Δintersection (3)
[0078] 4. Destination distance: To ensure that the agent moves steadily towards the destination, the driving behavior towards the target is rewarded, and a penalty for deviating from the target point is added. R destination is as shown in formula (4), where Δdist represents the change amount of the distance from the destination within two adjacent time steps.
[0079]
[0080] Combining the above four main reward items with other fixed bias rewards R bias forms the final reward function R as shown in formula (5):
[0081] R = R speed -Rdistance -R event +R destination +R bias (5)
[0082] The process of the multi-agent reinforcement learning algorithm based on PPO is as follows: First, initialize the driving environment in a simulator (such as CARLA, SUMO, or VISSIM), set multiple background vehicle agents, initialize an online Actor-Critic network and its corresponding target network for each agent, and at the same time create an experience replay pool for each agent to store interaction data. During the training process, the policy is updated iteratively through the main loop. First, reload the initial scene in each episode, and then at each time step, the agent selects an action according to the current state through the Actor network. After applying the action to the simulation environment, obtain the new state and reward, and store the state transition samples in the replay pool. Subsequently, sample a small batch of samples from the experience pool, and optimize the Actor and Critic networks through the PPO algorithm. PPO ensures the stability of training by restricting the range of policy changes, and at the same time combines the value function estimate and the policy gradient to update the network parameters. Every fixed number of episodes, the target network will synchronize the online network parameters through soft update to further enhance the stability of model training. This loop continues until the training converges or reaches the maximum number of episodes, and the algorithm will finally save the trained Actor network as the final interactive driving policy network.
[0083] S4: Interactive test scenario generation module.
[0084] Figure 4 and Figure 5 respectively show two technical framework diagrams of the interactive autonomous driving simulation scenario generation method. As Figure 4 and Figure 5 shown, by using the simulation platform and combining multi-level background vehicle driving strategies obtained based on different reinforcement learning frameworks, complex driving scenarios containing background vehicles with different levels of interactivity can be generated to help test and evaluate the performance of the autonomous driving system.
[0085] Exemplarily, three different levels of interactive driving strategies are obtained through the modules described above: Non-interactive strategy (Level-0): The background vehicle drives autonomously according to preset rules and does not interact with other vehicles at all. Semi-interactive strategy (Level-1): The background vehicle makes basic responses according to the behaviors of surrounding vehicles. Full-interactive strategy (Level-2): The background vehicle can actively perform behavior prediction and interaction, consider the decisions of other vehicles, and conduct cooperative driving.
[0086] Among them, the Level-1 driving strategy significantly increases the dynamic interactivity of the scenario by guiding the background vehicles to actively pursue greater self-driving benefits during driving, making the vehicle under test unable to rely solely on autonomous driving expectations and instead having to cope with more diverse behavioral interactions from the background vehicles. The Level-2 strategy assumes that the surrounding driving vehicles have a certain ability for active interaction. Therefore, vehicles equipped with this strategy tend to interact frequently with the surrounding vehicles to ensure their own safety, which requires the vehicle under test to improve driving efficiency through reasonable interaction. In actual simulation test applications, the interactive driving strategies of the background vehicles in the scenario need to be configured in different proportions according to the test purpose. If focusing on the safety test of the vehicle under test, the background vehicles should be configured with the Level-1 driving strategy, which helps to test the collision avoidance performance of the vehicle under test in diverse dynamic scenarios. If emphasizing the social driving ability of the vehicle under test, the background vehicles should be configured with the Level-2 driving strategy. Generally speaking, both the Level-1 and Level-2 driving strategies can significantly improve the interactivity of the test scenario.
[0087] After that, by setting the city map in the simulation environment, select the scenarios that need to be simulated, or build the required static simulation scenarios by yourself. The starting positions, speeds, target positions, etc. of the host vehicle (the vehicle under test) and the background vehicles will be set according to the test objectives. According to the levels of the background vehicles (Level-0, Level-1, Level-2, Level-3), load the corresponding driving strategies and apply them to the background vehicles. Control factors such as road conditions, weather conditions, and traffic signals to increase the diversity and complexity of the test. For example, as Figure 4 shown, in the reinforcement learning framework based on Level-K, three types of driving strategies can be applied to the background vehicles, and each type is configured in the simulation scenario according to requirements and proportions. Through constructing a static simulation scenario for real-time simulation, an interactive test scenario can be finally obtained.
[0088] Finally, the background vehicles in the simulation test scenario generated by this module can have continuous interactions with the vehicle under test through dynamic and diverse decisions. This interactive test scenario performs better in terms of interactivity, has good usability, and in actual tests, different purpose-oriented diverse test tasks can also be achieved by adjusting the proportions of driving strategies with different interaction levels in the scenario.
[0089] Next, the beneficial effects of the method shown in this application will be introduced in combination with Figure 6 and Figure 7 the experimental results shown. Among them, this application has set a total of five groups of test cases (such as Figure 6 and 7As shown in Groups 1 to 5, there are three types of driving strategies, namely L0 (non-interactive), L1 (semi-interactive), and L2 (fully interactive). Among them, all background vehicles in the test cases of Group 1 adopt a single driving strategy. That is, when all background vehicles in the scenario are L0 / L1 / L2, a traditional non-interactive, semi-interactive, and fully interactive simulation test scenario case is restored. The test scenarios under this case will be used as baseline test cases case1 to case3 for comparing with the interactive simulation test scenarios. The background vehicles in the test cases of Groups 2 to 4 have two different interactive level driving strategies among L0, L1, and L2, namely L0 and L1, L0 and L2, L1 and L2. Each group forms two test cases according to the different proportions of the two different driving strategies (each accounting for 1 / 3 or 2 / 3), namely case4 and case5, case6 and case7, and case8 and case9. The background vehicles in the test cases of Group 5 have three levels of driving strategies, L0, L1, and L2, with each strategy accounting for 1 / 3, forming a total of 1 test case, namely case10.
[0090] As Figure 6 shown, the collision rate of case1 is the lowest, while the arrival rate is the highest. Further analysis of different test cases reveals that when the proportion of L1-type vehicles in the scenario increases, the collision rate and arrival rate reach local maxima and local minima within the group respectively (marked by red circles in the figure). In the test cases of Group 3 (case6 and case7) that do not include the L1-type driving strategy, the collision rate is significantly lower than that of other groups, while the arrival rate is higher than that of other groups. As the proportion of the L1-type driving strategy increases, the collision rate of the test cases gradually rises, while the arrival rate continues to decline.
[0091] The above results indicate that the L1-type driving strategy significantly increases the dynamic interactivity of the scenario by guiding background vehicles to actively pursue greater self-driving benefits during driving, making the vehicle under test unable to rely solely on autonomous driving expectations but need to cope with more diverse behavioral interactions of background vehicles.
[0092] As Figure 7 shown, the time to reach the destination of the baseline test case case1 is the shortest. As the proportion of background vehicles with the L2-type driving strategy in the scenario increases, the time to reach the destination of the test cases increases and reaches local maxima within each group (marked by red circles in the figure), indicating that the L2-type driving strategy has a positive effect on enhancing the dynamic interactivity of the scenario. Facing background vehicles with the L2-type driving strategy, the vehicle under test needs to achieve diverse trajectory planning to avoid collisions, resulting in an increase in its intersection passing time.
[0093] Generally speaking, in view of the problems of insufficient background vehicle interactivity, lack of complexity in test scenarios, and high computational costs in the prior art, the present application proposes an interactive test scenario generation method based on multi-agent reinforcement learning. This method addresses the deficiencies of the prior art in the following aspects:
[0094] 1. Improving the interactivity of the scenario: By adopting a multi-agent reinforcement learning framework, the present application designs driving strategies with different interaction levels for background vehicles, enabling background vehicles to make diverse driving decisions according to real-time traffic conditions, thereby simulating more complex and realistic traffic scenarios. This method can enhance the dynamic interaction between background vehicles and the vehicle under test in the scenario, addressing the problem of insufficient interactivity in existing methods.
[0095] 2. Application of multi-agent reinforcement learning: The present application implements a multi-agent reinforcement learning framework guided by the Level-K game theory. This multi-level reinforcement learning framework can gradually improve the driving strategies of background vehicles according to different interaction levels, obtaining driving strategies with multiple interaction levels, thereby simulating more variable traffic situations and addressing the problem of insufficient complexity in generating test scenarios in the prior art.
[0096] 3. Low data dependence and scalability: The present application uses the reinforcement learning method, which has lower data dependence compared to the deep learning method and can generate high-fidelity dynamic test scenarios without relying on a large amount of actual driving data. In addition, compared with traditional methods, the reinforcement learning method can more flexibly adjust the driving strategies of background vehicles to adapt to different test requirements, enhancing the scalability of simulation tests.
[0097] Figure 8 It is a schematic block diagram of the communication device 800 according to an embodiment of the present application. The processing module 810 is used to implement the method as Figures 1 to 5 shown.
[0098] It should be understood that the communication device 800 is only an example. The device according to an embodiment of the present application may further include other modules or units, or include modules similar to the functions of the respective modules in Figure 8 , or not necessarily include all the modules in Figure 8 .
[0099] Figure 9 It is a schematic structural diagram of the communication device 900 according to an embodiment of the present application. It should be understood that Figure 9 the shown communication device 900 is only an example, and the communication device 900 according to an embodiment of the present application may further include other modules or units, or include modules similar to the functions of the respective modules in Figure 9 .
[0100] The communication device 900 may include one or more processors 910, one or more memories 920, a receiver 930, and a transmitter 940. The receiver 930 and the transmitter 940 may be integrated together and referred to as a transceiver. The memory 920 is used to store program code executed by the processor 910. Among them, the memory 920 may be integrated in the processor 910, or the processor 910 is coupled to one or more memories 920 for retrieving instructions in the memory 920.
[0101] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0102] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0103] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0104] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.
[0105] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0106] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0107] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0108] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0109] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0110] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0111] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0112] As described above, the above are only the specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for generating an interactive autonomous driving simulation test scenario based on multi-agent, characterized in that, Including: Based on image data and metric data, construct a dynamic driving model for simulating the driving decision-making process of a vehicle as a background vehicle. The image data includes images collected by the vehicle, and the metric data includes the state information of the vehicle and the current action commands of the vehicle. The dynamic driving model generates action information and value information through an image encoder, a metric encoder, and at least one neural network model. The action information is used to indicate control actions corresponding to the driving state of the vehicle, and the value information is used to evaluate the long-term effect of the current driving decision of the vehicle; Based on the hierarchical training idea and the dynamic driving model, determine and train driving strategies at multiple interaction levels. Among them, the vehicle as the background vehicle is trained in the training environments at the multiple interaction levels and is controlled by a fixed regularity model or the dynamic driving model to obtain the driving strategies at the multiple interaction levels; Construct a simulation test environment including the vehicle under test and the vehicle as the background vehicle; According to the dynamic driving model and the driving strategies at the multiple levels, based on the simulation test environment, configure the driving strategies at the multiple interaction levels for the vehicle as the background vehicle, and, based on the multi-agent reinforcement learning algorithm, jointly train the vehicle as the background vehicle and the vehicle under test to generate an autonomous driving simulation test scenario with interaction between the background vehicle and the vehicle under test.
2. The method according to claim 1, wherein The image encoder is a convolutional neural network and is used to extract image features from the image data to output a high-dimensional feature representation, and the metric encoder is used to generate a feature representation of the current state of the vehicle based on the metric data; The at least one neural network model includes a long short-term memory network LSTM for processing the time series dependence in driving decisions and a fully connected layer for outputting vehicle control commands and action predictions based on the fused features.
3. The method according to claim 1, characterized in that, The state information of the vehicle includes one or more of the vehicle's own speed, distance from the vehicle in front, relative speed, and current position, and the current action commands of the vehicle include one or more of the throttle, brake, and steering wheel angle.
4. The method according to claim 1, wherein The driving strategies at multiple interaction levels are generated based on game theory and a reinforcement learning framework.
5. The method according to claim 4, wherein The driving strategies at multiple interaction levels are generated based on the Level-K game theory, where the multiple interaction levels include: Non-interactive where the vehicle as the background vehicle drives based on fixed rules; semi-interactive where the vehicle as the background vehicle gradually introduces the dynamic driving model through an alternating control method; fully interactive where the vehicle as the background vehicle is completely controlled by the dynamic driving model.
6. The method according to claim 1, characterized in that, The construction of the simulation test environment including the vehicle under test and the vehicle as the background vehicle includes: Configure the interaction level ratio of the background vehicle according to the test target; Set the static scene parameters in the simulation test environment, and the static scene parameters include road conditions, weather, and traffic signals.
7. The method according to claim 1, characterized in that, Based on the multi-agent reinforcement learning algorithm, the joint training of the vehicle as the background vehicle includes: Initialize the Actor-Critic network and the target network for each background vehicle; Store interaction data through an experience replay pool and sample mini-batch samples to optimize network parameters; Guide policy optimization through a reward function.
8. The method according to claim 7, wherein The reward function includes driving speed, vehicle distance, events, and destination distance.
9. An interactive autonomous driving simulation test field system based on multi-agent, characterized in that, It includes: A dynamic driving model construction module for generating a dynamic driving model that simulates the driving decision-making process of a vehicle as a background vehicle; A hierarchical driving policy training module for determining and training driving policies at different interaction levels; A multi-agent reinforcement learning module for training the vehicle as a background vehicle in training environments at multiple interaction levels, and the vehicle as a background vehicle is controlled by a fixed regularity model or the dynamic driving model; An interactive test scenario generation module for configuring and generating interactive autonomous driving simulation test scenarios.
10. A computer-readable storage medium, characterized in that, Stored with a computer program, and when the computer program is executed by a processor, it implements the method described in any one of claims 1 to 8.
Citation Information
Cited By
Natural and adversarial AI collaborative automatic driving test scene generation method
CN121956936A
Automatic driving simulation agent method
CN121960199A