A multi-driving style high-risk automatic driving cut-in scene test method and system
By generating high-risk autonomous driving entry scenarios with multiple driving styles through the Cutin-TimeGAN network and Stackelberg game framework, the problem of not being able to achieve closed-loop interaction and controllable optimization of trajectory clusters in existing technologies is solved, which improves the diversity and adversarial nature of the test and dynamically adjusts the generated trajectory to discover risk interaction processes that are more in line with the test vehicle.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-20
AI Technical Summary
Existing autonomous driving testing methods cannot achieve closed-loop interaction with the strategy of the vehicle under test, cannot adaptively generate the most dangerous trajectory for different driving styles, and lack controllable optimization at the trajectory cluster level, resulting in insufficient test coverage.
By collecting raw driving trajectory data, the Cutin-TimeGAN network model is used to generate risk entry scenario trajectory clusters with consistent driving styles. A Stackelberg game framework is constructed for closed-loop interaction testing, and the optimal interaction trajectory is generated. Combined with physical feasibility constraints and cost functions, high-risk scenario testing with multiple driving styles is achieved.
It achieves closed-loop interaction with the strategy of the vehicle under test, and the generated entry scenarios are more in line with real behavior, improving the diversity and adversarial nature of the test, taking into account the authenticity of different driving styles and high-risk adversarial nature, and dynamically adjusting the generated trajectory to discover risk interaction processes that are more in line with the vehicle under test.
Smart Images

Figure CN121387751B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic driving car testing, in particular to a multi-driving style high-risk automatic driving cut-in scene testing method and system. BACKGROUND
[0002] When an automatic driving car faces a risk cut-in scene, it is often in a critical interaction state of strong coupling and high uncertainty, which is easy to trigger a collision or a near-miss accident. The scene-based testing method is an important testing method before the automatic driving car is put into use. The existing scene-based testing method mainly includes: a testing method based on natural driving data playback, which directly uses natural driving data for playback, although it is real, but lacks controllability and high-risk conditions; a testing method based on parameterized script predefinition, which generates scenes by setting rules and parameters, although it is controllable, but lacks real driving style and diversity. At the same time, many high-risk scene generation methods often cannot simultaneously consider real driving style and adversarial enhancement, resulting in insufficient test coverage, especially in critical interaction scenes under different driving styles. In order to solve the above technical problems, Chinese patent application CN120278193A provides an automatic driving risk lane change testing scene generation method, which first generates a risk lane change trajectory of a background vehicle BV using an improved Traj-TimeGAN; then constructs a critical safety distance model, under the premise that the AV takes a given braking behavior, the initial state of the AV is inversely calculated through an analytical formula, so that the two vehicles are in a critical state of just not colliding at a certain time C; then each generated BV trajectory and the AV initial state obtained by inverse calculation are directly used as a critical lane change test case, and then the lane change test scene is generalized; although the testing accuracy of high-risk critical lane change is realized, it still has the following shortcomings: 1) only open-loop critical scenes are generated, there is no closed-loop interaction test with the tested vehicle strategy, and the generated scene forms are basically the same for different tested vehicles, and the most dangerous trajectory cannot be adaptively generated for the specific control strategy of the tested vehicle; 2) each generated trajectory is regarded as an equivalent candidate scene, ignoring the balance between different driving style types of trajectory clusters, and lacking trajectory cluster level controllable optimization.
[0003] Therefore, a high-style cut-in scene testing method is provided, which can realize closed-loop interaction with the tested vehicle strategy, adaptively generate the most dangerous trajectory for the specific control strategy of the tested vehicle, and simultaneously consider the balance of different driving style trajectory clusters and support trajectory cluster level controllable optimization. SUMMARY
[0004] The purpose of the present application is to overcome the defects of the prior art and provide a multi-driving style high-risk automatic driving cut-in scene test method and system, which learns the cut-in trajectory distribution of multi-driving style from natural driving data, generates candidate trajectories with consistent and diversified styles, and uses them as style references to build a high-risk, controllable and repeatable cut-in interaction test scene in dynamic game, thereby comprehensively testing the safety of a vehicle under test (VUT) in a risk cut-in scene.
[0005] The purpose of the present application can be achieved by the following technical solutions:
[0006] According to a first aspect of the present application, a multi-driving style high-risk automatic driving cut-in scene test method is provided, comprising:
[0007] Collecting original driving trajectory data and constructing physical feasibility constraints, after driving style clustering of the original driving trajectory data, using a Cutin-TimeGAN network model combined with the physical feasibility constraints to generate a risk cut-in scene trajectory cluster with consistent driving style; the driving style includes conservative, ordinary and aggressive;
[0008] Constructing a cost function, calculating the cost function value for each risk cut-in scene trajectory cluster, and selecting a target cut-in reference trajectory as a target reference state vector based on the cost function value;
[0009] Constructing a Stackelberg game framework, constructing an utility function according to the state vectors of the VUT and the test opponent vehicle in the framework, and combining the target reference state vector;
[0010] Calculating the predicted collision time of the VUT and the test opponent vehicle, and constructing a risk confrontation utility function based on the predicted collision time;
[0011] Based on the utility function and the risk confrontation utility function, a risk cut-in interaction Stackelberg game optimization problem is constructed and solved to obtain the optimal interaction trajectory of the test opponent vehicle, and a closed-loop high-risk automatic driving cut-in scene test is performed based on the optimal interaction trajectory.
[0012] As a preferred technical solution, the physical feasibility constraints include:
[0013] For each vehicle, its state always moves forward, i.e. the vehicle's longitudinal displacement at the previous time is less than the vehicle's longitudinal displacement at the next time;
[0014] For the vehicle after the scene cut-in ends, the vehicle's lateral displacement point at the end time is located within the left and right boundaries of the lane;
[0015] For each vehicle, the lateral displacement difference between adjacent time frames does not exceed a difference threshold.
[0016] As a preferred technical solution, the loss function of the Cutin-TimeGAN network model comprises an embedder reconstruction loss, a discriminator loss and a generator adversarial loss.
[0017] The embedder reconstruction loss is: , represents a computational mathematical expectation operation, represents input driving trajectory data, represents a driving trajectory reconstructed by an embedder;
[0018] The discriminator loss is: , represents a discrimination result output by a discriminator based on a latent vector H, represents a cross-entropy loss calculation operation, represents a discrimination result output by a discriminator based on a generated latent representation , represents a weight coefficient, represents a discrimination result output by a discriminator based on a representation generated by random noise ;
[0019] The generator adversarial loss is: , represents an adversarial loss function, and ; represents a supervision loss, and , represents a real latent representation at t+1 time, represents a prediction result at t+1 time; represents a statistical matching loss, and , represents a mean calculation operation, represents trajectory data generated by a Cutin-TimeGAN network model; represents a standard deviation calculation operation.
[0020] As a preferred technical solution, for the i-th trajectory, the corresponding cost function is:
[0021] ,
[0022] wherein, , and all represent style weight coefficients; represents a total planning step length; represents a unit time length; denotes the longitudinal velocity of the vehicle at time t; denotes the jthplanning step at time t; denotes the reference vehicle speed; denotes the longitudinal acceleration of the vehicle at time t; denotes the lateral acceleration of the vehicle at time t;
[0023] As a preferred technical solution, in the Stackelberg game framework, the VUT is set as the follower F, and the test opponent vehicle is set as the leader L, and the state vectors of the two are set as:
[0024]
[0025]
[0026]
[0027] wherein, denotes the state vector of the follower F at time t; denotes the state vector of the leader L at time t; and denote the longitudinal positions of the follower and the leader at time t, respectively; and denote the longitudinal velocities of the follower and the leader at time t, respectively; and denote the lateral positions of the follower and the leader at time t, respectively; and denote the lateral velocities of the follower and the leader at time t, respectively;
[0028] The state weight coefficient matrix is constructed as: wherein, denotes the weight proportion coefficient of different states;
[0029] The control input weight matrix is constructed as: wherein, denotes the weight proportion coefficient of different control inputs; and the control input vector is set as: , denotes the longitudinal control input of the vehicle at time t, denotes the lateral control input of the vehicle at time t;
[0030] The state transition of the follower and the leader follows:
[0031]
[0032] ,
[0033] wherein, represents the control input vector of the follower at the kth moment; represents the state transition matrix; represents the input gain matrix; represents the control input vector of the leader at the kth moment.
[0034] As a preferred technical solution, the utility function is:
[0035] ,
[0036] ,
[0037] wherein, represents the utility function of the follower, i.e. VUT; represents the state vector of the follower at the k+1th moment; represents the target reference state vector of the follower; represents the state weight coefficient matrix; represents the control input vector of the follower at the kth moment; represents the control input weight matrix; represents the total step length of planning; represents the utility function of the leader, i.e. the test confrontation vehicle; represents the state vector of the leader at the k+1th moment; represents the target reference state vector of the leader; represents the control input vector of the leader at the kth moment.
[0038] As a preferred technical solution, the risk confrontation utility function is:
[0039] ,
[0040] wherein, represents the total step length of planning; represents the length of the trajectory traveled by the follower; represents the speed of the follower; represents the length of the trajectory traveled by the leader; represents the speed of the leader.
[0041] As a preferred technical solution, the risk cut-in interaction Stackelberg game optimization problem is:
[0042] ,
[0043] ,
[0044] ,
[0045] wherein, denotes the leader's control input vector; denotes the risk driving regulation coefficient; denotes the compliance natural driving regulation coefficient; denotes the leader's utility function, i.e. the test opponent vehicle; denotes the risk opponent utility function; denotes the follower's control input vector at the kth moment; denotes the state transition matrix; denotes the input gain matrix; denotes the leader's control input vector at the kth moment; denotes the leader's longitudinal position at the kth moment; denotes the follower's longitudinal position at the kth moment; denotes the preset safety distance; denotes the i-th vehicle's control input vector; denotes the control input vector lower bound; denotes the control input vector upper bound; denotes the i-th vehicle's speed; denotes the maximum vehicle speed; denotes the follower's state vector at the kth moment; denotes the leader's state vector at the kth moment.
[0046] As a preferred technical solution, the method for solving the risk cut-in interactive Stackelberg game optimization problem is:
[0047] The risk cut-in interactive Stackelberg game optimization problem is converted into a leader optimal control problem and a follower optimal control problem, i.e. wherein, denotes the leader's state vector; denotes the leader's control input vector; denotes the follower's optimal control input vector; denotes the follower's control input vector; denotes the follower's state vector;
[0048] According to the convex optimization property of the follower optimal control problem, the risk cut-in interactive Stackelberg game optimization problem is converted into a single-layer convex optimization problem by introducing the KKT condition, which is:
[0049] ,
[0050] ,
[0051] ,
[0052] ,
[0053] wherein, and represent the Lagrange coefficients; represents the risky driving regulation coefficient; represents the compliance natural driving regulation coefficient; represents the leader, i.e., the test opponent vehicle, utility function; represents the risky opponent utility function; represents the follower's k-th time control input vector; represents the state transition matrix; represents the input gain matrix; represents the leader's k-th time control input vector; represents the leader's k-th time longitudinal position; represents the follower's k-th time longitudinal position; represents the preset safety distance; represents the i-th vehicle's control input vector; represents the control input vector lower bound; represents the control input vector upper bound; represents the i-th vehicle's speed; represents the maximum vehicle speed; represents the Lagrange function; represents the follower's k-th time state vector; represents the leader's k-th time state vector;
[0054] The single-layer convex optimization problem is solved by using CasADi and Ipopt solvers.
[0055] According to the second aspect of the present application, a multi-driving style high-risk automatic driving cut-in scene test system is provided for implementing the above method.
[0056] Compared with the prior art, the present application has the following beneficial effects:
[0057] 1) In view of the problem in the prior art that the closed-loop interaction test of the measured vehicle strategy cannot be realized, after obtaining the risk cut-in trajectory cluster with multiple driving styles based on TimeGAN, the present application optimizes a trajectory with the minimum current cost function value as a reference trajectory, and incorporates the reference trajectory into the construction of the game optimization problem, so that the scene generation process becomes a closed-loop game process of the VUT and the opponent vehicle; in the closed-loop game process, the risky driving regulation coefficient And compliance natural driving adjustment coefficient The above two parameters are dynamically adjusted for each different driving style, so that the generated cut-in trajectory is automatically adjusted according to the control strategy of the tested vehicle. Compared with the single-direction critical scene in the prior art, the application can find a risk interaction process that is more consistent with the real behavior of the tested vehicle, significantly improves the test diversity and antagonism, and at the same time, takes into account the real driving style and high risk antagonism.
[0058] 2) The application does not need to collect any trajectory manually, and different groups of cut-in trajectories are automatically generated by the TimeGAN algorithm to form test scenes, and dynamic interaction testing is realized. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 The flowchart of the method of the application;
[0060] Figure 2 The driving trajectory graph and speed change graph in the risk cut-in scene test process in the embodiment of the application. DETAILED DESCRIPTION
[0061] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the application.
[0062] In order to solve the problems in the prior art, the application provides a multi-driving style high-risk automatic driving cut-in scene test method, and the flowchart thereof is shown in Figure 1 The method comprises the following steps in detail:
[0063] S1, collect original driving trajectory data and construct physical feasibility constraints, after driving style clustering of the original driving trajectory data, generate risk cut-in scene trajectory clusters with consistent driving style by using a Cutin-TimeGAN network model combined with physical feasibility constraints.
[0064] S11, original driving trajectory data collection and processing.
[0065] In this embodiment, the collected original driving trajectory data is derived from the trajectory data of the risk cut-in scene of the automatic driving vehicle on the expressway extracted from the public data set.
[0066] S111. Load a multi-time period vehicle trajectory dataset, where the vehicle trajectory data includes timestamps, vehicle IDs, and vehicle driving status information, including location, speed, heading, and lane markings; filter the driving status information of each vehicle based on the vehicle ID, and identify the entry event by the lane markings of the vehicle itself and the driving status information of other vehicles in the target lane, and obtain complete entry trajectory data.
[0067] S112. Filter outliers in the cutting trajectory data and perform Z-score normalization to obtain the final complete... n Vehicle entry trajectory set , of which i vehicle trajectory , T Given the duration of vehicle cut-in, the vehicle cut-in trajectory at time t includes the following state information: the vehicle's longitudinal position. Horizontal position Longitudinal velocity lateral velocity and heading angle , can be represented as: .
[0068] S12, Driving Style Clustering.
[0069] To identify the driving styles of different drivers during the cut-in process, a feature vector is constructed using the cut-in duration T, where the feature vector for the i-th vehicle is: .
[0070] The density-based clustering algorithm DBSCAN is used to cluster the feature vectors of the scene data. In this embodiment, the core idea of DBSCAN is to define the feature vectors. of - Neighborhood:
[0071] ,
[0072] in, Let the eigenvector of the j-th vehicle be denoted as ; if a point is within its eigenvector... - The number of samples in the neighborhood is not less than the minimum number of samples. These are called core points. Core points gradually expand through "density reachability" and "density connectivity," forming a cluster set. Based on trajectory feature data, driving styles are categorized into conservative, average, and aggressive within this cluster set.
[0073] S13. Construct physical feasibility constraints.
[0074] To ensure that the generated trajectory meets practical requirements, the following constraints are established in this invention:
[0075] i) for each vehicle, its state always moves forward, that is, the longitudinal displacement of the vehicle at the previous time is less than that at the next time, and for the ith vehicle, it can be expressed as: , denotes the longitudinal displacement of the ith vehicle at time t.
[0076] ii) for the vehicle after the end of the scene cut-in, the lateral displacement point of the vehicle at the end time is located within the left and right boundaries of the lane, and for the ith vehicle, it can be expressed as: , denotes the left boundary constraint of the lane, denotes the lateral displacement point of the ith vehicle at time t, denotes the right boundary constraint of the lane.
[0077] iii) for each vehicle, the lateral displacement difference between adjacent time frames does not exceed the difference threshold, and for the ith vehicle, it can be expressed as: .
[0078] S14, risk cut-in scene trajectory cluster generation of multiple driving styles.
[0079] On the basis of a traditional time series generative adversarial network (TimeGAN), a cut-in time series generative adversarial network (Cutin-TimeGAN) is proposed for a high-speed cut-in scene. The specific improvements include: introducing traffic behavior feature constraints to ensure the physical feasibility of the generated trajectories; introducing prior styles to ensure that the generated trajectories remain consistent under different driving styles. In the Cutin-TimeGAN network provided in the present application, there are an embedder E , a generator G , a restorer R , a supervisor S and a discriminator D . In detail, there are:
[0080] The embedder E inputs the cut-in trajectory sequence data after driving style clustering, and encodes the trajectory sequence data into a low-dimensional latent vector through a multi-layer recurrent neural network, retaining the time sequence dependence feature.
[0081] The restorer R inputs the latent vector H , and decodes and reconstructs the trajectory through a symmetric network structure, ensuring the effectiveness of the encoding.
[0082] The generator GInput random noise sequence Z and driving style tags Output the fake potential vector Simulates the temporal trajectory of real-world risk entry scenarios.
[0083] Supervisor S Input the true latent vector H With generating latent vectors Output time prediction results .
[0084] Discriminator D The generator is optimized by classifying and distinguishing between real and fake latent vectors.
[0085] When training the Cutin-TimeGAN network, three types of collaborative loss functions are constructed: embedder reconstruction loss, discriminator loss, and generator adversarial loss, to ensure the real distribution and time sequence rationality of the generated driving trajectories.
[0086] In order to make the input via embedder E get The reconstructed trajectory is then obtained through the restorer R. The reconstruction loss is minimized in this process; the reconstruction loss for building the embedder is: , This indicates the computational expectation operation. This represents the input driving trajectory data. This represents the driving trajectory reconstructed by the embedder.
[0087] The goal of a discriminator is to distinguish true latent representations. Generate latent representations and only random noise The generated representation The discriminator loss can be: , This represents the discrimination result output by the discriminator based on the latent vector H. This indicates the cross-entropy loss calculation operation. The discriminator is based on generating latent representations. The output of the discrimination result, Indicates the weighting coefficient (in this embodiment) ), The discriminator is based on random noise. The generated representation The output is the discrimination result.
[0088] The goal of the generator is to make the discriminator think... , All are real, so the generator can be constructed to generate an adversarial loss:
[0089] ,
[0090] wherein, represents an adversarial loss function, which is used to measure the difference between x and the real label, since the goal of the generator is to 'trick' the discriminator, so in this loss term, the label is taken, that is, it is hoped that the discriminator will think that the generated sample is a real sample, and ;
[0091] represents a supervised loss, which aims to constrain the temporal consistency, and requires the generated latent representation to be able to predict the real , , represents the real latent representation at t+1 time, represents the prediction result at t+1 time;
[0092] represents a statistical matching loss, in order to ensure the consistency of the overall statistical distribution, the matching of the first moment (mean) and the second moment (standard deviation) is introduced, and , represents a mean calculation operation, represents the trajectory data generated by the Cutin-TimeGAN network model; represents a standard deviation calculation operation.
[0093] After training the network model using the above loss function, according to the actual requirements, determine the network type, hidden layer dimension, number of layers, training iteration number, batch size and other key parameters, load the driving trajectory data, and use the trained Cutin-TimeGAN network to generate risk cut-in scene trajectory clusters. By inputting the trajectory data of different driving styles after clustering, it is ensured that the driving style of the network model is consistent in each generation of risk cut-in scene trajectory clusters.
[0094] For the above generated risk cut-in scene trajectory cluster, filter the cut-in trajectory that meets the above physical feasibility constraint as an effective trajectory segment, and obtain the final risk cut-in trajectory cluster that meets the requirements.
[0095] S2, construct a cost function, calculate the cost function value for each risk cut-in scene trajectory cluster, and select a target cut-in reference trajectory as a target reference state vector based on the cost function value.
[0096] S21, cost function construction.
[0097] Taking the ith trajectory as an example, a corresponding cost function can be constructed as:
[0098]
[0099] all represent style weight coefficients; represents the total step length of planning; represents the unit time length; represents t j the vehicle longitudinal speed at the moment, and represents the moment of the jth planning step; represents the reference vehicle speed; represents t j the vehicle longitudinal acceleration at the moment; represents t j the vehicle lateral acceleration at the moment.
[0100] According to the cost function constructed above, the corresponding function value is calculated, and under the premise of meeting the safety, feasibility and style constraints, the trajectory with the minimum cost function value is selected as the target cut-in reference trajectory.
[0101] S22, target reference state vector construction.
[0102] According to the target cut-in reference trajectory data, the state vector at the target moment is constructed as:
[0103]
[0104]
[0105] and respectively represent the target reference longitudinal position of the VUT and the test opponent vehicle; and respectively represent the target reference longitudinal speed of the VUT and the test opponent vehicle; and respectively represent the target reference lateral position of the VUT and the test opponent vehicle; and respectively represent the target reference lateral speed of the VUT and the test opponent vehicle.
[0106] S3, construct a Stackelberg game framework, and construct an utility function according to the state vectors of the VUT and the test opponent vehicle in the framework and the target reference state vector.
[0107] In the Stackelberg game framework, we have:
[0108] Let VUT as the follower F and the test opponent vehicle as the leader L, set the state vector form of both as:
[0109]
[0110]
[0111] wherein, represents the state vector of the follower F at the kth moment; represents the state vector of the leader L at the kth moment; and respectively represent the longitudinal position of the follower and the leader at the kth moment; and respectively represent the longitudinal speed of the follower and the leader at the kth moment; and respectively represent the lateral position of the follower and the leader at the kth moment; and respectively represent the lateral speed of the follower and the leader at the kth moment; the initial state can be set as:
[0112]
[0113]
[0114] The state transition of the follower and the leader follows:
[0115]
[0116]
[0117] wherein, represents the control input vector of the follower at the kth moment; represents the state transition matrix; represents the input gain matrix; represents the control input vector of the leader at the kth moment; then in the prediction interval , based on the initial state , the future state trajectory can be obtained through the above state transition as:
[0118]
[0119] The state weight coefficient matrix is constructed as: wherein, represents the weight proportion coefficient of different states.
[0120] The control input weight matrix is constructed as: wherein, and represent the weight proportion coefficients of different control inputs; and the control input vector is set as: , represents the longitudinal control input of the vehicle at time k, represents the lateral control input of the vehicle at time k.
[0121] On the basis of the above architecture, for the time interval , the utility functions of the follower and the leader are respectively designed as:
[0122] ,
[0123] ,
[0124] wherein, represents the utility function of the follower, i.e., the VUT; represents the state vector of the follower at time k+1; represents the target reference state vector of the follower; represents the state weight coefficient matrix; represents the control input vector of the follower at time k; represents the control input weight matrix; represents the total planning step length; represents the utility function of the leader, i.e., the test countermeasure vehicle; represents the state vector of the leader at time k+1; represents the target reference state vector of the leader; represents the control input vector of the leader at time k.
[0125] S4, the predicted collision time of the VUT and the test countermeasure vehicle is calculated, and a risk countermeasure utility function is constructed based on the predicted collision time.
[0126] In addition, the leader needs to actively counterattack the VUT in addition to meeting the driving compliance and driving smoothness, and therefore needs to additionally design a risk countermeasure utility function, which is:
[0127] ,
[0128] wherein, represents the total planning step length; represents the length of the trajectory traveled by the follower; represents the speed of the follower; represents the length of the trajectory traveled by the leader; represents the speed of the leader.
[0129] S5, construct a risk-cutting interactive Stackelberg game optimization problem based on the utility function and the risk-antagonistic utility function, and solve it to obtain an optimal interactive trajectory of the test-antagonistic vehicle, and perform closed-loop high-risk automatic driving cut-in scene testing based on the optimal interactive trajectory.
[0130] S51, construct a risk-cutting interactive Stackelberg game optimization problem.
[0131] On the basis of the functions constructed in the foregoing steps, in order to ensure the interaction safety of the follower and the leader, the longitudinal collision constraint needs to be met: In combination with the utility function and the risk-antagonistic utility function, the optimization problem can be obtained as:
[0132] ,
[0133] ,
[0134] ,
[0135] wherein, represents the control input vector of the leader; represents the risk driving adjustment coefficient; represents the compliance natural driving adjustment coefficient; represents the utility function of the leader, i.e., the test-antagonistic vehicle; represents the risk-antagonistic utility function; represents the control input vector of the follower at the kth moment; represents the state transition matrix; represents the input gain matrix; represents the control input vector of the leader at the kth moment; represents the longitudinal position of the leader at the kth moment; represents the longitudinal position of the follower at the kth moment; represents a preset safety distance; represents the control input vector of the ith vehicle; represents the lower bound of the control input vector; represents the upper bound of the control input vector; represents the speed of the ith vehicle; represents the maximum vehicle speed; represents the state vector of the follower at the kth moment; represents the state vector of the leader at the kth moment.
[0136] S52, solve the game optimization problem.
[0137] In the Stackelberg game framework, the optimization problem of the leader and the follower is a typical bi-level optimal control problem, the follower makes a response action according to the strategy of the leader, and therefore the upper optimization is the leader optimal control and the lower optimization is the follower optimal control, that is, the risk-cut interactive Stackelberg game optimization problem can be transformed into the leader optimal control problem and the follower optimal control problem, which is represented as:
[0138] ,
[0139] wherein, denotes the leader state vector; denotes the leader control input vector; denotes the follower optimal control input vector; denotes the follower control input vector; denotes the follower state vector.
[0140] In the foregoing modeling, the utility function of the follower is quadratic, the constraints are composed of the system dynamics equation and the collision constraint, and both are in the affine form, which is a convex optimization problem, then according to the convex optimization property of the above-mentioned, the KKT condition is introduced to transform the risk-cut interactive Stackelberg game optimization problem into a single-layer convex optimization problem, which is:
[0141] ,
[0142] ,
[0143] ,
[0144] ,
[0145] wherein, and both denote the Lagrange coefficient; denotes the risk driving regulation coefficient; denotes the compliance natural driving regulation coefficient; denotes the utility function of the leader, that is, the test confrontation vehicle; denotes the risk confrontation utility function; denotes the control input vector of the follower at the kth moment; denotes the state transition matrix; denotes the input gain matrix; denotes the control input vector of the leader at the kth moment; denotes the longitudinal position of the leader at the kth moment; denotes the longitudinal position of the follower at the kth moment; denotes the preset safety distance; represents a control input vector of the i-th vehicle; represents a lower bound of the control input vector; represents an upper bound of the control input vector; represents a speed of the i-th vehicle; represents a maximum vehicle speed; represents a Lagrange function; represents a state vector of the follower at the k-th time; represents a state vector of the leader at the k-th time.
[0146] Finally, a single-layer convex optimization problem is solved by using CasADi and Ipopt solvers to obtain an optimal trajectory of the leader interaction strategy.
[0147] The obtained optimal trajectory of the leader interaction strategy is sent to a simulation environment (such as CARLA) or a real vehicle test platform in real time for closed-loop measurement and setting, and statistical indexes (TTC, speed change curve) are output.
[0148] After one round of S1-S5 is executed, whether to continue testing is selected according to requirements, if it is required to continue, driving trajectory data generated after the current round of testing is executed is collected, and S1-S5 is executed again.
[0149] In order to verify the feasibility of the method, the method provided in the application is used to test a high-risk automatic driving cut-in scene, and a result graph as shown in Figure 2 can be drawn according to a simulation result. Figure 2 It can be seen that in the relative distance-time step diagram, the curve shows that the distance between the two vehicles gradually decreases during the test, and the collision risk gradually increases; in the speed-time step diagram, the curve shows that the test vehicle can adjust the trajectory in real time according to the speed change of the tested vehicle, forming dynamic interaction, which proves that the method provided in the application is feasible.
[0150] In addition, the application provides a multi-driving style high-risk automatic driving cut-in scene test system and an electronic device, which are used to implement the above method. The electronic device of the application includes a central processing unit (CPU), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded into a random access memory (RAM) from a storage unit. In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0151] A number of components in the device are connected to the I / O interface, including: input units, such as a keyboard, a mouse, etc.; output units, such as various types of displays, speakers, etc.; storage units, such as a magnetic disk, an optical disk, etc.; and communication units, such as a network card, a modem, a wireless communication transceiver, etc. The communication units allow the device to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0152] The processing unit performs various methods and processes described above, such as the methods S1-S5. For example, in some embodiments, the methods S1-S5 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the methods S1-S5 described above can be performed. Alternatively, in other embodiments, the CPU can be configured to perform the methods S1-S5 by any other suitable means, such as by means of firmware.
[0153] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.
[0154] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flow charts and / or block diagrams to be implemented. The program code can execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0155] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage medium can include, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage medium would include one or more lines of electrical wire, portable computer diskette, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0156] The above description is only specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A testing method for high-risk autonomous driving entry scenarios with multiple driving styles, characterized in that, include: Raw driving trajectory data is collected and physical feasibility constraints are constructed. After clustering the raw driving trajectory data by driving style, the Cutin-TimeGAN network model is used in conjunction with the physical feasibility constraints to generate risk entry scenario trajectory clusters with consistent driving styles. The driving styles include conservative, normal, and aggressive. The Cutin-TimeGAN network model includes an embedder. E Generator G Recovery device R Supervisor S and discriminator D There are: Embedders E Input trajectory sequence data after driving style clustering, and encode the trajectory sequence data into a low-dimensional latent vector through a multi-layer recurrent neural network. Recovery device R Input latent vector H Trajectory is reconstructed by decoding latent vectors using a symmetric network structure. ; generator G Input random noise sequence Z and driving style tags Output the fake potential vector Supervisor S Input the true latent vector H With generating latent vectors Output time prediction results Discriminator D Classify and distinguish between real and fake latent vectors to drive generator optimization; Construct a cost function, calculate the cost function value for each risk entry scenario trajectory cluster, and select the target entry reference trajectory as the target reference state vector based on the cost function value; Construct a Stackelberg game framework, and based on the state vectors of the VUT and the test adversary vehicle in the framework, combine them with the target reference state vector to construct a utility function; the utility function is as follows: , , in, This represents the utility function of the follower, i.e., the VUT; This represents the state vector of the follower at time k+1. This represents the target reference state vector of the follower; Represents the state weight coefficient matrix; This represents the control input vector of the follower at time k. This represents the control input weight matrix; Indicates the total planning step length; This represents the utility function of the leader, i.e., the test vehicle. This represents the leader's state vector at time k+1. This represents the leader's target reference state vector; This represents the leader's control input vector at time k. Calculate the estimated collision time of the VUT and the test vehicle, and construct a risk mitigation utility function based on the estimated collision time; the risk mitigation utility function is: , in, Indicates the total planning step length; Indicates the length of the path traveled by the follower; Indicates the speed of the follower; Indicates the length of the leader's travel route; Indicates the leader's speed; Based on the aforementioned utility function and risk adversarial utility function, a risk-based entry interaction Stackelberg game optimization problem is constructed and solved to obtain the optimal interaction trajectory of the test adversarial vehicle. Closed-loop high-risk autonomous driving entry scenario testing is then conducted based on this optimal interaction trajectory. The aforementioned risk-based entry interaction Stackelberg game optimization problem is as follows: , , , in, This represents the leader's control input vector; This indicates the risk driving adjustment factor; Indicates the compliant natural driving adjustment coefficient; This represents the utility function of the leader, i.e., the test vehicle. This represents the risk mitigation utility function; This represents the control input vector of the follower at time k. Represents the state transition matrix; Represents the input gain matrix; This represents the leader's control input vector at time k. This represents the leader's vertical position at time k. This represents the vertical position of the follower at time k. Indicates the preset safety distance; This represents the control input vector for the i-th vehicle; Indicates the lower bound of the control input vector; This indicates the upper bound of the control input vector; Indicates the speed of the i-th vehicle; Indicates the maximum vehicle speed; This represents the state vector of the follower at time k. This represents the leader's state vector at time k.
2. The testing method for high-risk autonomous driving entry scenarios with multiple driving styles according to claim 1, characterized in that, The aforementioned physical feasibility constraints include: For each vehicle, its state is always forward movement, that is, the longitudinal displacement of the vehicle at the previous moment is less than the longitudinal displacement of the vehicle at the next moment. For vehicles that have finished entering the scene, the lateral displacement point of the vehicle at the end time is located within the left and right boundaries of the lane. For each vehicle, the difference in lateral displacement between adjacent time frames does not exceed the difference threshold.
3. The testing method for high-risk autonomous driving entry scenarios with multiple driving styles according to claim 1, characterized in that, The loss function of the Cutin-TimeGAN network model includes the embedder reconstruction loss, the discriminator loss, and the generator adversarial loss. The embedder reconstruction loss is as follows: , This indicates the computational expectation operation. This represents the input driving trajectory data. This represents the driving trajectory reconstructed by the embedder; The discriminator loss is: , This represents the discrimination result output by the discriminator based on the latent vector H. This indicates the cross-entropy loss calculation operation. The discriminator is based on generating latent representations. The output of the discrimination result, Indicates the weighting coefficient. The discriminator is based on random noise. The generated representation The output of the discrimination result; The generator adversarial loss is: , Describing the adversarial loss function, we have ; Indicates monitoring losses, there are , This represents the true potential representation at time t+1. This represents the prediction result at time t+1; Represents the statistical matching loss, with , This indicates the operation of calculating the mean. This represents trajectory data generated by the Cutin-TimeGAN network model; This indicates the standard deviation calculation operation.
4. The testing method for high-risk autonomous driving entry scenarios with multiple driving styles according to claim 1, characterized in that, For the i-th trajectory, the corresponding cost function is: , in, , and All represent style weight coefficients; Indicates the total planning step length; Indicates the unit of time; express The longitudinal speed of the vehicle at any given time, and This represents the time of the j-th planning step; Indicates reference speed; express t j The longitudinal acceleration of the vehicle at any given moment; express The vehicle's lateral acceleration at all times.
5. The testing method for high-risk autonomous driving entry scenarios with multiple driving styles according to claim 1, characterized in that, In the Stackelberg game framework described above, we have: Let the VUT be the follower F, and the test adversary vehicle be the leader L. Let their state vectors be defined as follows: , , in, This represents the state vector of follower F at time k. This represents the state vector of leader L at time k. and These represent the vertical positions of the follower and the leader at time k, respectively. and Let represent the longitudinal velocities of the follower and the leader at time k, respectively. and These represent the lateral positions of the follower and the leader at time k, respectively. and Let represent the lateral velocities of the follower and leader at time k, respectively. The state weight coefficient matrix is constructed as follows: ,in, The weighting coefficients representing different states; The control input weight matrix is constructed as follows: ,in, This represents the weighting coefficients for different control inputs; and the control input vector is defined as: , This represents the longitudinal control input of the vehicle at time k. This represents the lateral control input of the vehicle at time k; The state transitions between followers and leaders are governed by: , , in, This represents the control input vector of the follower at time k. Represents the state transition matrix; Represents the input gain matrix; This represents the leader's control input vector at time k.
6. The testing method for high-risk autonomous driving entry scenarios with multiple driving styles according to claim 1, characterized in that, The method for solving the aforementioned risk-entry interaction Stackelberg game optimization problem is as follows: The aforementioned risk-based Stackelberg game optimization problem is transformed into a leader-optimal control problem and a follower-optimal control problem, namely: ,in, Represents the leader's state vector; This represents the leader's control input vector; This represents the follower's optimal control input vector; This represents the follower control input vector; Represents the follower state vector; Based on its convex optimization property, the follower optimal control problem is transformed into a single-layer convex optimization problem by introducing KKT conditions, which are: , , , , in, and Both represent the Lagrange coefficients; This indicates the risk driving adjustment factor; Indicates the compliant natural driving adjustment coefficient; This represents the utility function of the leader, i.e., the test vehicle. This represents the risk mitigation utility function; This represents the control input vector of the follower at time k. Represents the state transition matrix; Represents the input gain matrix; This represents the leader's control input vector at time k. This represents the leader's vertical position at time k. This represents the vertical position of the follower at time k. Indicates the preset safety distance; This represents the control input vector for the i-th vehicle; Indicates the lower bound of the control input vector; This indicates the upper bound of the control input vector; Indicates the speed of the i-th vehicle; Indicates the maximum vehicle speed; Represent the Lagrange function; This represents the state vector of the follower at time k. This represents the leader's state vector at time k. The CasaADi and Ipopt solvers were used to solve the single-layer convex optimization problem.
7. A testing system for high-risk autonomous driving scenarios with multiple driving styles, characterized in that, The system is used to implement the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Expressway shunting area forced lane change decision-making method oriented to intelligent network connection environment
CN120412313A
Self-driving automobile test method based on cooperative hunting confrontation of multiple traffic participants
CN120430075A