An autonomous driving testing method based on multi-coalition cluster confrontation
By adopting multi-alliance cluster adversarial method in autonomous driving tests, and using reinforcement learning and alliance games to dynamically generate high-adversarial test scenarios, the problem of difficulty in generating high-risk boundary scenarios in traditional testing methods is solved, and the testing efficiency is improved.
Patent Information
- Application Number
- CN202410992688.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-23
AI Technical Summary
Traditional autonomous driving simulation testing methods are difficult to generate high-risk boundary scenarios, and the test scenarios are low in compatibility with the test vehicles, resulting in low testing efficiency.
The autonomous driving test method based on multi-alliance cluster confrontation is adopted, and the background vehicle cluster is divided through reinforcement learning algorithms, and the background vehicle alliance is highly confrontational behavior decision-making is achieved through alliance game methods, and a test scenario with high confrontationality with test vehicles is dynamically generated.
It improves the diversity and testing difficulty of autonomous driving test scenarios, can find dangerous boundary scenarios of autonomous driving faster, and significantly improves simulation testing efficiency.
Smart Images

Figure CN118936910B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving virtual simulation testing, and in particular to an autonomous driving testing method based on multi-alliance cluster confrontation. Background Art
[0002] In recent years, autonomous driving technology has become a hot research direction in the automotive and transportation fields. Compared with human-driven vehicles, autonomous driving systems can significantly improve traffic safety, reduce accident rates, and greatly improve driving comfort and economy. However, in the future, strict autonomous driving system testing will be an indispensable key link to ensure the safety and stability of autonomous driving systems. Autonomous driving test methods include road testing, field testing, simulation testing, etc. Among them, autonomous driving simulation testing can simulate richer and more complex scenarios, conduct virtualization and automation testing, have high scenario coverage, high safety and low cost, and is currently the most effective autonomous driving testing method.
[0003] However, as the level of autonomous driving continues to improve, long-tail scenarios and boundary scenarios have become a major bottleneck restricting autonomous driving simulation testing. The traditional method of generating autonomous driving test scenarios based on parameter combinations is mainly achieved through digital mapping of trajectory data. It does not have the ability to dynamically interact with the test vehicle. The test scenario has a low degree of fit with the test vehicle, and it is easy to generate a large number of meaningless redundant scenarios. It is difficult to generate high-risk boundary scenarios, and the efficiency of key test scenario generation is low. The autonomous driving test method based on environmental confrontation fully considers the dynamic interaction relationship between the background vehicle and the test vehicle, and can generate highly confrontational test scenarios with the test vehicle in real time by controlling the behavior of the background vehicle. The test scenarios generated by this method have a high degree of fit and pertinence with the test vehicle, making it easier to find high-risk boundary scenarios for the test vehicle, effectively improving the efficiency of autonomous driving testing. Summary of the invention
[0004] The present invention aims to overcome the shortcomings of the prior art and discloses an autonomous driving test method based on multi-alliance cluster confrontation.
[0005] The present invention can be implemented by the following technical solutions:
[0006] An automatic driving test method based on multi-coalition cluster confrontation is characterized in that, firstly, adaptive division of background vehicle confrontation clusters is realized through reinforcement learning algorithm, background vehicles are divided into groups with different confrontation intensity levels, and test vehicles are made to interact with background vehicle groups with different confrontation intensity levels, so as to more comprehensively test the automatic driving capability of the tested vehicle, effectively improve the diversity of scenes and the test difficulty, and facilitate finding dangerous boundary scenes; at the same time, each background vehicle group is regarded as an alliance, and the vehicles within the alliance have a cooperative relationship, and the interactive high-adversarial behavior decision of the background vehicle alliance is realized through the alliance game method; finally, trajectory planning is carried out for the background vehicle cluster according to the behavior decision result, so that the background vehicle can travel according to the planned trajectory.
[0007] An autonomous driving test method based on multi-coalition cluster confrontation, characterized by comprising:
[0008] S1: Initialization of the autonomous driving test environment.
[0009] S2: Background vehicle adversarial clustering decision.
[0010] S3: Background vehicle adversarial behavior decision.
[0011] S4: Background vehicle trajectory planning.
[0012] S5: Execute the above steps S2, S3, and S4 repeatedly until the cluster confrontation test task is completed.
[0013] The S1 autonomous driving test environment initialization includes: selecting the required test map on the virtual simulation test platform, setting environmental information such as the number of lanes and lane width; determining the number of background vehicles and their generation locations, test tasks, selecting the start and end points of the test in the test map, and generating a global reference path for the vehicle under test.
[0014] The S2 background car adversarial clustering decision:
[0015] A reinforcement learning algorithm is used to realize the decision-making process of adversarial clustering in background vehicles, and the background vehicle clustering process is modeled as a Markov process M(S,A,P,R,γ), where S is the state space of the background vehicle cluster, A is the action space of the background vehicle cluster, P is the state transition probability, R is the immediate reward obtained after the background vehicle cluster performs an action, and γ is the attenuation factor; all background vehicles are regarded as reinforcement learning algorithm agents, and the environment initialization information obtained in step S1 is used as the environment information input of the reinforcement learning algorithm; the optimal environment vehicle clustering result, that is, the action space of the background vehicle cluster, is the output of the reinforcement learning algorithm and provided to S3.
[0016] The S2 background vehicle adversarial clustering decision further includes:
[0017] The construction process of the reinforcement learning model is as follows:
[0018] The state space is
[0019] S={i∈0,1,…,n|s i}
[0020] s i ={x i ,y i ,v x,i ,v y,i ,a x,i ,a y,i}
[0021] Where n is the number of background vehicles in the test environment, s i It is the driving status information of the background vehicle and the test vehicle. i ,y i ,v x,i ,v y,i ,a x,i ,a y,i are the longitudinal position, lateral position, longitudinal velocity, lateral velocity, longitudinal acceleration and lateral acceleration of the background vehicle and the test vehicle respectively;
[0022] Among them, the action space is composed of different background car cluster division schemes. The action space is
[0023] A={A 1 ×A 2}
[0024] A 1 ={a 1 ,a 2 ,…,a p}
[0025] A 2 = {q 1 ,...,q v}
[0026] Where A is the total action space of reinforcement learning, A 1 ={a 1 ,a 2 ,...,a p} are all cluster division schemes under different adversarial strengths of n vehicles, A 2 = {q 1 ,…,q v} is the confrontation intensity of each alliance.
[0027] Among them, the reward function R includes the adversarial strength reward r 1 , Acceleration Reward 2 , collision penalty r 3and driving range penalty r 4 ;
[0028] R=w 1 r 1 +w 2 r 2 +w 3 r 3 +w 4 r 4
[0029]
[0030] Among them, x 0 ,y 0 is the horizontal and vertical position of the test vehicle, a i is the acceleration of the i-th background car, a max is the maximum acceleration of the background vehicle, d max is the maximum distance between the background car and the test car, d i is the distance between the i-th background car and the test car, w 1 ,w 2 ,w 3 ,w 4 is the weight of each reward;
[0031] The training process of the reinforcement learning model is as follows, implemented using the DQN algorithm:
[0032] First, a Q target network with the same structure as the Q network is established, and random parameters are used to initialize the Q network and the Q target network, and the maximum number of training rounds is determined;
[0033] In each training epoch:
[0034] According to the current network Q π (s,a) selects action a with a greedy strategy t , execute actions, update reports and status information (s i ,a i ,r i ,s t+1 ), where π refers to a specific strategy;
[0035] Will (s i ,a i ,r i ,s t+1 ) into the playback pool D. If there is enough data in D, sample N data from D and input them into the target Q network to calculate the learning target Minimize the target loss and use it to update the current Q network weights;
[0036] Update the Q target network and repeat the above process until the maximum training round;
[0037] When the reinforcement learning model is trained to convergence, the action it outputs is the optimal result of the environment vehicle cluster division.
[0038] The S3 background car adversarial behavior decision:
[0039] According to the background vehicle group division results, the background vehicle cluster is divided into alliances with different confrontation strengths; the coordinated behavior decision of the background vehicle cluster is realized through the alliance game method, and a cooperative relationship is established between the alliances to jointly realize the confrontation process of the test vehicle, thereby improving the confrontation strength of the scene. Preferably, the alliance game belongs to a mixed strategy dynamic game.
[0040] The S3 background car adversarial behavior decision, specifically, the action space of the alliance is
[0041] A coal ={i∈1,…,m|a m,i ,}
[0042] a m,i = {b 1 ,b 2 ,b 3 ,b 4 ,b 5 ,b 6 ,b 7}
[0043] Among them, A coal is the action space of the alliance, m is the number of vehicles in the alliance, a m,i is the action space of the i-th vehicle in the alliance, b 1 ,b 2 ,b 3 ,b 4 ,b 5 ,b 6 ,b 7 These are the actions that can be performed by vehicles in the alliance, including left lane change, right lane change, lane keeping, large acceleration, large deceleration, small acceleration, and small deceleration;
[0044] After establishing the action space, it is necessary to predict the state of the vehicles in the alliance after selecting different actions; predict the state of the background vehicles at the next moment after performing the corresponding action. The prediction process is as follows:
[0045] For the i-th background car in the m-th alliance:
[0046] Known current vehicle status information: x(k), y(k), v x (k),v y (k) are longitudinal position, lateral position, lateral velocity and longitudinal velocity respectively, and the prediction time domain is T;
[0047] The predicted vehicle state at the next moment is:
[0048]
[0049] v x (k+1)=v x (k)+a x T
[0050] v y (k+1)=v y (k)+a y T
[0051] Among them, when the vehicle action is to select different actions, a x and a y Take different values;
[0052] After completing the action prediction, a profit function is established to calculate the profit obtained by the alliance when selecting different actions; the profit function e includes the adversarial profit e 1 、Speed Benefit 2 and target tracking benefits 3 , the calculation process is as follows:
[0053]
[0054] e=w 1 E 1 +w 2 E 2 +w 3 E 3
[0055] Among them, x targ ,t targ The estimated target position after the vehicle selects different actions, w 1 ,w 2 ,w 3 is the weight coefficient of each benefit, and the antagonistic benefit e 1 The weight coefficient w 1 Related to the strength of confrontation, for the i-th alliance, w 1,i =w×q i , w is a constant;
[0056] The probability of vehicles choosing different strategies is used as the optimization variable to establish the alliance game mixed strategy optimization problem, whose objective function and constraints are:
[0057]
[0058] Among them, n coal is the number of alliances, n act is the number of actions that the alliance can choose, uij is the probability of the alliance choosing the action; the constraint condition is that the sum of the probabilities of all actions chosen by the alliance is 1; solve the above optimization problem to obtain the mixed strategy decision results of each alliance, sample the actions according to the strategy distribution, and the sampling result is the behavior decision result of the background car.
[0059] The S4 plans the background vehicle trajectory:
[0060] First, the Frenet curve coordinate system is established with the road centerline as a reference. The subsequent trajectory planning steps are completed in the Frenet coordinate system.
[0061] Secondly, the position range x of the end point of trajectory planning is determined according to the behavior decision results of the background vehicle lim =[x min ,x max ],y lim =[y min ,y max ] and the time range of reaching the destination t lim =[t min ,t max ]; uniformly sample the calculated position range and time range according to a certain position interval and time interval, combine the position information and time information obtained by sampling, and obtain the state information of the trajectory planning end point;
[0062] Use the quintic polynomial method to fit the trajectory from the current starting point to different end states to form a trajectory cluster;
[0063] Establish a trajectory evaluation index to evaluate the trajectories in the trajectory cluster, and select the optimal trajectory as the trajectory planning result output; the trajectory evaluation index is as follows:
[0064] E t1 =v x (t f )-v x (t 0 )
[0065] E t2 =||(x targ ,y targ )-(x(t f ),y(t f ))||
[0066]
[0067] E t =w t1 E t1 +w t1 E t2 +w t3 Et3
[0068] Among them, E t1 is the speed gain, E t2 is the target return, E t3 is the collision gain, v x is the longitudinal velocity of the trajectory, t 0 and t f are the start and end time of the trajectory, respectively, x targ ,y targ is the expected end position of the trajectory, w t1 ,w t2 ,w t3 is the weight of each indicator;
[0069] The optimal trajectory is selected through trajectory evaluation indicators as the reference trajectory of the background vehicle, and the planning result is input into the lower-level trajectory tracking controller so that the background vehicle can travel according to the reference trajectory.
[0070] The S5 loops through the above steps S2, S3, and S4 until the cluster confrontation test task is completed:
[0071] When both the tested vehicle and the background vehicle complete the decision-making planning process, the status information of the tested vehicle and the background vehicle is updated according to the planning results, and the corresponding environmental information is also updated. The termination condition of the test task is determined based on the updated environmental information: if the current test time has reached the predetermined test time or the tested vehicle collides with the background vehicle, the test task ends, otherwise the above steps S2, S3, and S4 are executed repeatedly until the cluster confrontation test task is completed.
[0072] This paper proposes an autonomous driving test method based on multi-coalition cluster confrontation, aiming to improve the efficiency and accuracy of autonomous driving simulation testing. This method dynamically generates test scenarios that are highly confrontational with the test vehicle through reinforcement learning and alliance game, which can find the dangerous boundary scenarios of autonomous driving more quickly and improve the efficiency of simulation testing.
[0073] The advantages of the present invention are as follows:
[0074] 1. The present invention realizes adaptive division of background vehicle clusters through reinforcement learning algorithm, realizes continuous adjustment of strategies through interaction with the environment, has strong environmental adaptability, can dynamically generate high-fit risk scenarios for different tested objects, and improves the generalization of the test method.
[0075] 2. The present invention models the cooperative relationship between background vehicle clusters through alliance game, improves the overall scene confrontation strength through the collaborative cooperation of background vehicle clusters, can generate more complex scenes, and improve test efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1The figure is a flow chart of the testing process of the present invention.
[0077] Figure 2 This is a schematic diagram of initializing a simulation scene according to an embodiment of the present invention.
[0078] Figure 3 Schematic diagram of the multi-alliance cluster confrontation process according to an embodiment of the present invention. DETAILED DESCRIPTION
[0079] The technical solution of the present invention belongs to an autonomous driving virtual simulation test method, which can be applied to scenes such as multi-lane highways, ramp merging, and unprotected left turns. The technical solution of the present invention is applied to an autonomous driving virtual simulation test platform (hereinafter referred to as the test platform):
[0080] The test platform is the upper host;
[0081] The test platform can collect the status information of the tested autonomous vehicle in real time and transmit the information to the environment vehicle cluster;
[0082] The method of the present invention is run on a test platform, and the obtained planning results are input into a lower-layer trajectory tracking controller, which controls the real-time operation of the environmental vehicle cluster.
[0083] The present invention proposes an autonomous driving test method based on multi-alliance cluster confrontation, which runs on a test platform: first, the adaptive division of the background vehicle confrontation cluster is realized through a reinforcement learning algorithm, and the background vehicles are divided into groups with different confrontation intensity levels. The test vehicle is allowed to interact with the background vehicle groups with different confrontation intensity levels, which can more comprehensively test the autonomous driving ability of the tested vehicle, effectively improve the diversity of scenes and the difficulty of testing, and facilitate the finding of dangerous boundary scenes. At the same time, each background vehicle group is regarded as an alliance, and the vehicles within the alliance have a cooperative relationship. The interactive high-adversarial behavior decision of the background vehicle alliance is realized through the alliance game method. Finally, the trajectory of the background vehicle cluster is planned according to the behavior decision results, so that the background vehicle can travel according to the planned trajectory.
[0084] The technical solution of the present invention is further introduced below in conjunction with the accompanying drawings and embodiments.
[0085] Example
[0086] In this embodiment,
[0087] Definition of the vehicle under test: A high-level autonomous driving vehicle that needs to undergo simulation testing, equipped with high-performance sensors and capable of autonomous perception, decision-making planning, and control;
[0088] Definition of background vehicle: other vehicles generated in the simulation test platform that have an interactive relationship with the vehicle under test.
[0089] like Figure 1As shown, an autonomous driving test method based on multi-coalition cluster confrontation includes the following steps:
[0090] S1: Initialization of the autonomous driving test environment
[0091] The method of the present invention is implemented based on an autonomous driving simulation platform, which can provide a variety of autonomous driving virtual simulation scenarios and vehicle models. By connecting the planning control algorithm to the autonomous driving simulation platform, the movement of the vehicle in the virtual scenario can be controlled. The autonomous driving simulation platform is an autonomous driving virtual simulation test platform, which can provide test scenarios, vehicle models, sensor models, etc., to achieve autonomous driving closed-loop simulation.
[0092] Select the required test map on the virtual simulation test platform and set the number of lanes, lane width and other environmental information. Determine the number of background vehicles and their generation locations, edit the test task, select the start and end points of the test on the map, and generate a global reference path for the vehicle under test. The schematic diagram of the initial simulation scenario is shown in the figure.
[0093] S2: Background vehicle adversarial clustering decision:
[0094] A reinforcement learning algorithm is used to realize the decision-making process of adversarial clustering in background vehicles, and the background vehicle clustering process is modeled as a Markov process M(S,A,P,R,γ), where S is the state space of the background vehicle cluster, A is the action space of the background vehicle cluster, P is the state transition probability, R is the immediate reward obtained after the background vehicle cluster performs an action, and γ is the attenuation factor; all background vehicles are regarded as intelligent agents of the reinforcement learning algorithm, and the environment initialization information obtained in step S1 is used as the environment information input of the reinforcement learning algorithm; the optimal environment vehicle clustering result, that is, the action space of the background vehicle cluster, is the output of the reinforcement learning algorithm and provided to S3.
[0095] Details are as follows.
[0096] The construction process of the reinforcement learning model is as follows:
[0097] The state space is
[0098] S={i∈0,1,…,n|s i}
[0099] s i ={x i ,y i ,v x,i ,v y,i ,a x,i ,a y,i}
[0100] Where n is the number of background vehicles in the test environment, s i It is the driving status information of the background vehicle and the test vehicle.i ,y i ,v x,i ,v y,i ,a x,i ,a y,i are the longitudinal position, lateral position, longitudinal velocity, lateral velocity, longitudinal acceleration and lateral acceleration of the background vehicle and the test vehicle respectively.
[0101] Among them, the action space is composed of different background car cluster division schemes. The action space is
[0102] A={A 1 ×A 2}
[0103] A 1 ={a 1 ,a 2 ,…,a p}
[0104] A 2 = {q 1 ,...,q v}
[0105] Where A is the total action space of reinforcement learning, A 1 ={a 1 ,a 2 ,…,a p} are all cluster division schemes under different adversarial strengths of n vehicles, A 2 = {q 1 ,…,q v} is the confrontation intensity of each alliance.
[0106] Among them, the reward function R includes the adversarial strength reward r 1 , Acceleration Reward 2 , collision penalty r 3 and driving range penalty r 4 wait.
[0107] R=w 1 r 1 +w 2 r 2 +w 3 r 3 +w 4 r 4
[0108]
[0109] Among them, x 0 ,y 0 is the horizontal and vertical position of the test vehicle, a i is the acceleration of the i-th background car, a maxis the maximum acceleration of the background vehicle, d max is the maximum distance between the background car and the test car, d i is the distance between the i-th background car and the test car, w 1 ,w 2 ,w 3 ,w 4 The weight of each reward.
[0110] After the reinforcement learning model is built, it needs to be pre-trained.
[0111] The model training process is as follows, and the embodiment is implemented using the DQN algorithm:
[0112] First, a Q target network with the same structure as the Q network is established, and random parameters are used to initialize the Q network and the Q target network, and the maximum training rounds are determined.
[0113] In each training epoch:
[0114] According to the current network Q π (s,a) selects action a with a greedy strategy t , execute actions, update reports and status information (s i ,a i ,r i ,s t+1 ), where π refers to a specific strategy.
[0115] Will (s i ,a i ,r i ,s t+1 ) into the playback pool D. If there is enough data in D, sample N data from D and input them into the target Q network to calculate the learning target Minimize the target loss and use it to update the current Q network weights;
[0116] Update the Q target network and repeat the above process until the maximum number of training rounds.
[0117] When the reinforcement learning model is trained to convergence, the action it outputs is the optimal result of the environment vehicle cluster division.
[0118] S3: Background car adversarial behavior decision:
[0119] According to the background vehicle group division results, the background vehicle cluster is divided into alliances with different confrontation strengths. The collaborative behavior decision of the background vehicle cluster is realized through the alliance game method, and a cooperative relationship is established between the alliances to jointly realize the confrontation process of the test vehicle, thereby enhancing the confrontation strength of the scene. The background vehicle cluster confrontation process is shown in the figure. Preferably, the alliance game belongs to a mixed strategy dynamic game. During the game process, each alliance makes decisions in a certain order, and the latter alliance knows the decision results of all previous alliances when making decisions. In addition, the decision result of each alliance is not a fixed strategy, but a probability distribution of different strategies, which can improve the diversity of strategies and better cope with complex and changeable dynamic game environments.
[0120] Preferably, the action space of the alliance is
[0121] A coal ={i∈1,…,m|a m,i ,}
[0122] a m,i = {b 1 ,b 2 ,b 3 ,b 4 ,b 5 ,b 6 ,b 7}
[0123] Among them, A coal is the action space of the alliance, m is the number of vehicles in the alliance, a m,i is the action space of the i-th vehicle in the alliance, b 1 ,b 2 ,b 3 ,b 4 ,b 5 ,b 6 ,b 7 These are the actions that can be performed by vehicles in the alliance, including changing lanes to the left, changing lanes to the right, lane keeping, large acceleration, large deceleration, small acceleration, and small deceleration.
[0124] After establishing the action space, it is necessary to predict the state of the vehicles in the alliance after selecting different actions. In order to improve the real-time performance of the algorithm, in this embodiment, only the state of the background vehicles at the next moment after executing the corresponding action is predicted. The prediction process is as follows:
[0125] For the i-th background car in the m-th alliance:
[0126] Known current vehicle status information: x(k), y(k), v x (k),v y(k) are the longitudinal position, lateral position, lateral velocity and longitudinal velocity respectively, and the prediction time domain is T.
[0127] The predicted vehicle state at the next moment is:
[0128]
[0129] v x (k+1)=v x (k)+a x T
[0130] v y (k+1)=v y (k)+a y T
[0131] Among them, when the vehicle action is to select different actions, a x and a y Take different values.
[0132] After completing the action prediction, it is necessary to establish a profit function to calculate the profit obtained by the alliance when choosing different actions.
[0133] Preferably, the benefit function E includes the adversarial benefit e 1 、Speed Benefit 2 and target tracking benefits 3 , the calculation process is as follows:
[0134]
[0135] E=w 1 E 1 +w 2 E 2 +w 3 E 3
[0136] Among them, x targ ,y targ The estimated target position after the vehicle selects different actions, w 1 ,w 2 ,w 3 is the weight coefficient of each benefit, and the antagonistic benefit e 1 The weight coefficient w 1 Related to the strength of confrontation, for the i-th alliance, w 1,i =w×q i , w is a constant.
[0137] The probability of vehicles choosing different strategies is used as the optimization variable to establish the alliance game mixed strategy optimization problem, whose objective function and constraints are:
[0138]
[0139] Among them, n coal is the number of alliances, n act is the number of actions that the alliance can choose, u ij The probability of the alliance choosing the action is 1. The constraint is that the sum of the probabilities of all actions chosen by the alliance is 1. Solve the above optimization problem to obtain the mixed strategy decision results of each alliance, sample the actions according to the strategy distribution, and the sampling result is the behavior decision result of the background car.
[0140] S4: Background vehicle trajectory planning
[0141] The trajectory of the background vehicle is planned according to the behavior decision result of the background vehicle. In this embodiment, the Lattice method is used for path planning.
[0142] First, the Frenet curve coordinate system is established with the road centerline as a reference, and the subsequent trajectory planning steps are completed in the Frenet coordinate system.
[0143] Secondly, the position range x of the end point of trajectory planning is determined according to the behavior decision results of the background vehicle lim =[x min ,x max ],y lim =[y min ,y max ] and the time range of reaching the destination t lim =[t min ,t max ]. Uniform sampling is performed within the calculated position range and time range according to a certain position interval and time interval, and the position information and time information obtained by sampling are combined to obtain the state information of the trajectory planning end point.
[0144] The trajectories from the current starting point to different end states are fitted using the quintic polynomial method to form trajectory clusters.
[0145] A dynamics verification module is established to screen the endpoint trajectories of the trajectory cluster and delete the trajectories that do not meet the vehicle dynamics requirements.
[0146] Establish trajectory evaluation indicators to evaluate the trajectories in the trajectory cluster and select the optimal trajectory as the trajectory planning result output. The trajectory evaluation indicators are as follows:
[0147] E t1 =v x (t f )-v x (t 0 )
[0148] E t2 =||(x targ ,ytarg )-(x(t f ),y(t f ))||
[0149]
[0150] E t =w t1 E t1 +w t1 E t2 +w t3 E t3
[0151] Among them, E t1 is the speed gain, E t22 is the target return, E t3 is the collision gain, v x is the longitudinal velocity of the trajectory, t 0 and t f are the start and end time of the trajectory, respectively, x targ ,y targ is the expected end position of the trajectory, w t1 ,w t2 ,w t3 is the weight of each indicator.
[0152] The optimal trajectory is selected as the reference trajectory of the background vehicle through the trajectory evaluation index, and the planning result is input into the lower-level trajectory tracking controller so that the background vehicle can travel according to the reference trajectory. The trajectory tracking controller can be implemented by the proportional-integral-differential control method, which will not be described in detail here.
[0153] S5: Execute the above steps S2, S3, and S4 repeatedly until the cluster confrontation test task is completed.
[0154] When both the tested vehicle and the background vehicle complete the decision-making planning process, the status information of the tested vehicle and the background vehicle is updated according to the planning results, and the corresponding environmental information is also updated. The termination condition of the test task is determined based on the updated environmental information: if the current test time has reached the predetermined test time or the tested vehicle collides with the background vehicle, the test task ends, otherwise the above steps S2, S3, and S4 are executed repeatedly until the cluster confrontation test task is completed.
Claims
1. An autonomous driving test method based on multi-coalition cluster confrontation, characterized in that: Firstly, the adaptive division of background vehicle adversarial clusters is realized through reinforcement learning algorithm, and background vehicles are divided into groups with different adversarial strength levels, so that the test vehicle interacts with background vehicle groups with different adversarial strength levels; at the same time, each background vehicle group is regarded as an alliance, and the vehicles within the alliance have a cooperative relationship. The interactive high-adversarial behavior decision of the background vehicle alliance is realized through the alliance game method; finally, the trajectory of the background vehicle cluster is planned according to the behavior decision results, so that the background vehicle can drive according to the planned trajectory; S1: Initialization of the autonomous driving test environment; S2: Background vehicle adversarial clustering decision; S3: Background vehicle adversarial behavior decision; S4: background vehicle trajectory planning; S5: loop through the above steps S2, S3, and S4 until the cluster confrontation test task is completed; The S2 background vehicle adversarial clustering decision: a reinforcement learning algorithm is used to implement the background vehicle adversarial clustering decision process, and the background vehicle clustering process is modeled as a Markov process ,in, is the state space of the background car cluster, is the action space of the background car cluster, is the state transition probability, It is the immediate reward obtained after the background vehicle group performs an action. is the attenuation factor; all background vehicles are regarded as the intelligent agents of the reinforcement learning algorithm, and the environment initialization information obtained in step S1 is used as the environment information input of the reinforcement learning algorithm; the optimal environment vehicle cluster division result, that is, the action space of the background vehicle cluster, is the output of the reinforcement learning algorithm and provided to S3; The S2 background vehicle adversarial clustering decision further includes: The construction process of the reinforcement learning model is as follows: The state space is in, is the number of background vehicles in the test environment, It is the driving status information of the background vehicle and the test vehicle. are the longitudinal position, lateral position, longitudinal velocity, lateral velocity, longitudinal acceleration and lateral acceleration of the background vehicle and the test vehicle respectively; Among them, the action space is composed of different background car cluster division schemes. The action space is in is the total action space of reinforcement learning, for All cluster division schemes under different adversarial strengths of vehicles, The intensity of the confrontation for each alliance; Among them, the reward function Includes combat strength bonus , Acceleration Rewards , Collision Penalty and driving range penalty ; in, is the horizontal and vertical position of the test vehicle, For the The acceleration of the background car, is the maximum acceleration of the background vehicle, is the maximum distance between the background car and the test car, For the The distance between the background car and the test car, is the weight of each reward; The training process of the reinforcement learning model is as follows, implemented using the DQN algorithm: First, a Q target network with the same structure as the Q network is established, and random parameters are used to initialize the Q network and the Q target network, and the maximum number of training rounds is determined; In each training epoch: According to the current network Selecting actions with a greedy strategy , execute actions, update reports and status information ;in, refers to a specific strategy; Will Put into the playback pool ,like There is enough data from Medium Sampling Data input target Q network calculation learning target , minimize the target loss and use it to update the current Q network weights; Update the Q target network and repeat the above process until the maximum training round; When the reinforcement learning model is trained to convergence, the action it outputs is the optimal environment vehicle cluster division result.
2. The method for testing autonomous driving based on multi-coalition cluster confrontation according to claim 1, characterized in that: The S1 autonomous driving test environment initialization includes: selecting the required test map on the virtual simulation test platform, setting environmental information such as the number of lanes and lane widths; determining the number of background vehicles and their generation locations, test tasks, selecting the start and end points of the test in the test map, and generating a global reference path for the tested vehicle.
3. The autonomous driving test method based on multi-coalition cluster confrontation according to claim 1, characterized in that: The S3 background car adversarial behavior decision: According to the background vehicle group division results, the background vehicle cluster is divided into alliances with different confrontation strengths; the collaborative behavior decision-making of the background vehicle cluster is realized through the alliance game method, and cooperative relationships are established between the alliances to jointly realize the confrontation process of the test vehicle, thereby improving the confrontation strength of the scene.
4. The method for testing autonomous driving based on multi-coalition cluster confrontation according to claim 3, characterized in that: The alliance game is a mixed strategy dynamic game.
5. The method for testing autonomous driving based on multi-coalition cluster confrontation according to claim 3, characterized in that: The S3 background car adversarial behavior decision, specifically, the action space of the alliance is in, is the action space of the alliance, is the number of vehicles in the alliance, The first in the alliance The action space of the vehicle, These are the actions that can be performed by vehicles in the alliance, including left lane change, right lane change, lane keeping, large acceleration, large deceleration, small acceleration, and small deceleration; After establishing the action space, it is necessary to predict the state of the vehicles in the alliance after selecting different actions; predict the state of the background vehicles at the next moment after performing the corresponding action. The prediction process is as follows: For The first in the alliance Background vehicles: Known current vehicle status information: They are longitudinal position, lateral position, lateral velocity and longitudinal velocity respectively, and the prediction time domain is ; The predicted vehicle state at the next moment is: Among them, when the vehicle action is to select different actions, and Take different values; After completing the action prediction, a profit function is established to calculate the profit obtained by the alliance when choosing different actions; the profit function Including adversarial benefits , speed benefits and target tracking benefits , the calculation process is as follows: in, The estimated target location after selecting different actions for the vehicle, is the weight coefficient of each benefit, the antagonistic benefit The weight coefficient Related to the intensity of confrontation, alliance, , is a constant; The probability of vehicles choosing different strategies is used as the optimization variable to establish the alliance game mixed strategy optimization problem, whose objective function and constraints are: in, is the number of alliances, is the number of actions that the alliance can choose, is the probability of the alliance choosing a certain action; the constraint condition is that the sum of the probabilities of all actions chosen by the alliance is 1; solve the above optimization problem to obtain the mixed strategy decision results of each alliance, sample the actions according to the strategy distribution, and the sampling result is the behavior decision result of the background car.
6. The autonomous driving test method based on multi-coalition cluster confrontation according to claim 1, characterized in that: The S4 plans the background vehicle trajectory: First, the Frenet curve coordinate system is established with the road centerline as a reference. The subsequent trajectory planning steps are completed in the Frenet coordinate system. Secondly, the location range of the end point of trajectory planning is determined based on the behavior decision results of the background vehicle. , and the time frame for reaching the destination ; Perform uniform sampling within the calculated position range and time range according to a certain position interval and time interval, combine the sampled position information and time information, and obtain the status information of the trajectory planning end point; Use the quintic polynomial method to fit the trajectory from the current starting point to different end states to form a trajectory cluster; Establish a trajectory evaluation index to evaluate the trajectories in the trajectory cluster, and select the optimal trajectory as the trajectory planning result output; the trajectory evaluation index is as follows: in, For speed gain, For target income, For collision benefit, is the longitudinal velocity of the trajectory, and are the start and end time of the trajectory, is the expected end position of the trajectory, is the weight of each indicator; The optimal trajectory is selected through trajectory evaluation indicators as the reference trajectory of the background vehicle, and the planning result is input into the lower-level trajectory tracking controller so that the background vehicle can travel according to the reference trajectory.
7. The method for testing autonomous driving based on multi-coalition cluster confrontation according to claim 1, characterized in that: The S5 loops through the above steps S2, S3, and S4 until the cluster confrontation test task is completed: When both the tested vehicle and the background vehicle have completed the decision-making and planning process, the status information of the tested vehicle and the background vehicle are updated according to the planning results, and the corresponding environmental information is also updated; the termination condition of the test task is determined based on the updated environmental information: if the current test time has reached the predetermined test time or the tested vehicle collides with the background vehicle, the test task ends; otherwise, the above steps S2, S3, and S4 are executed repeatedly until the cluster confrontation test task is completed.
Citation Information
Patent Citations
Registration test evaluation method for automatic driving automobile in lane changing scene
CN117892631A