Background vehicle anthropomorphic trajectory generation method for automatic driving test scene
By constructing a dynamic Bayesian game framework and a multi-objective utility function, combined with a random perturbation mechanism to optimize the background vehicle trajectory generation, the problems of multi-objective rigidity and insufficient anthropomorphism in the existing technology of background vehicle trajectory generation are solved, and higher anthropomorphism and reliability are achieved.
Patent Information
- Application Number
- CN202510715282.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing background vehicle trajectory generation methods in autonomous driving test scenarios have a rigid multi-objective trade-off mechanism and are unable to flexibly respond to complex and changing traffic environments. The personalized modeling of driver behavior styles is rough, making it difficult to accurately simulate the diversity of human driving. There is a lack of anthropomorphic evaluation methods, resulting in a disconnect between simulation results and real driving scenarios.
A dynamic Bayesian game framework is constructed using a driving simulator. Through driving style clustering and multi-objective utility functions, anthropomorphic trajectories are generated in real time. Combined with a random perturbation mechanism, the behavior of background vehicles is optimized to achieve autonomous decision-making and anthropomorphic interaction of background vehicles in complex scenarios.
It achieves improved diversity and authenticity of background vehicle behavior, and can autonomously adjust the weight distribution of goals such as safety, efficiency, and urgency according to real-time scene characteristics, demonstrating human-like decision-making hesitation and strategy switching characteristics, and improving the anthropomorphism and reliability of the test scene.
Smart Images

Figure CN120671349A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving testing technology, and in particular to a method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios. Background Art
[0002] In recent years, with the continuous advancement of autonomous driving technology, the number of autonomous vehicles on the road has increased significantly. This has been accompanied by a number of autonomous vehicle accidents, which have exposed the shortcomings of current autonomous driving systems in handling complex traffic scenarios and highlighted the need for strengthening related testing. Current testing methods primarily include virtual simulation testing, real-world road testing, and combined virtual-real testing. While real-world road testing allows for direct evaluation of autonomous vehicles in a real traffic environment, it suffers from long testing cycles and high costs. To mitigate the shortcomings of real-world road testing, virtual simulation testing and combined virtual-real testing have gradually become the primary means of autonomous driving testing. Both testing methods are conducted in simulated test scenarios, significantly reducing testing costs and time, improving testing efficiency, and enabling the tested vehicle to complete numerous complex scenarios in a short period of time. Consequently, they have become widely used.
[0003] Currently, in simulation test scenarios, the vehicle under test needs to interact with background vehicles in the scenario to evaluate its driving performance in dynamic scenarios such as lane changing, following a vehicle, and overtaking. Therefore, the behavior of background vehicles has an important impact on the performance evaluation of the autonomous driving system. For example, the closer the behavior of the background vehicle is to the intention of the real driver, the higher the authenticity of the simulation test scenario and the more accurate the test results. However, in existing simulation test scenarios, data-driven methods such as neural networks and reinforcement learning are mainly used to explore the driver's driving habits and style, so as to achieve background vehicle control in the test scenario that is closer to human driving behavior. In existing autonomous driving test scenarios, the generation of background vehicle trajectories has at least the following defects:
[0004] 1) The multi-objective trade-off mechanism is rigid and cannot flexibly respond to complex and changing traffic environments. For example, when faced with complex and changing traffic scenarios, it is unable to accurately balance multiple objectives such as safety, efficiency, urgency, and energy consumption in real time based on driving style.
[0005] 2) The personalized modeling of driver behavior styles is crude and difficult to accurately simulate the diversity of human driving. If fixed weights or simple adjustment strategies are used, it will be difficult to adapt to the dynamic changes in the attention paid to various targets by different driving styles, making the background vehicle decision-making lack rationality and adaptability. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios, so as to solve the technical problems mentioned in the prior art.
[0007] A method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios includes the following steps:
[0008] S1. Extracting driving characteristics from several sets of interactive scenarios according to a predetermined interactive logic and conducting simulation training to construct a driving simulator. The driving simulator uses the driving characteristics from each set of interactive scenarios as driving behavior data for several background vehicles and configures a predetermined driving style for each of them. A corresponding driving strategy is then set for each driving style to construct a dynamic Bayesian game framework.
[0009] S2. Real-time acquisition of the driving style of the background vehicles in any set of interactive scenarios, and iterative training of the multi-objective utility function.
[0010] S3. Calculating a benefit value corresponding to the driving style of the background vehicle using the multi-objective utility function, and selecting a corresponding driving strategy from a dynamic Bayesian game framework as the optimal response action of the background vehicle based on the calculated benefit value;
[0011] S4. Generate a trajectory control instruction according to the selected driving strategy, and perform anthropomorphic optimization on the trajectory control instruction using a random perturbation mechanism to obtain an anthropomorphic trajectory of the background vehicle.
[0012] Optionally, in S1, the interaction scenario includes at least one of lane change game, intersection merging, and following a vehicle in congestion.
[0013] Optionally, in S1, the driving characteristics include one or more of the vehicle's acceleration distribution, lane change frequency, following distance, steering wheel angle, and pedal travel in any detection cycle.
[0014] Optionally, in S1, the method for setting the driving style includes:
[0015] Analyzing the driving behavior data of each background vehicle using a clustering algorithm, dividing the driving behavior data of the background vehicles into corresponding clusters based on their characteristic similarities through several iterative calculations, and taking each cluster as a driving style;
[0016] The competitiveness level c of the corresponding driving style is quantified by the center value of the cluster:
[0017] c=w1·F norm +w2·D norm +w3·A norm ;
[0018]
[0019] Where c∈(0,1); F is the lane change frequency, F min is the minimum lane change frequency, F max is the maximum lane change frequency, F norm is the normalized lane change frequency, w1 is the characteristic weight of the lane change frequency; D is the average following distance, D min is the minimum average following distance, D max is the maximum average following distance, D norm is the normalized average following distance, w2 is the feature weight of the average following distance; A is the acceleration variance, A min is the minimum acceleration variance, A max is the maximum acceleration variance, A norm is the normalized acceleration variance, w3 is the feature weight of the acceleration variance;
[0020] The background vehicles are classified into corresponding driving style categories based on the competitiveness level c of the driving style.
[0021] Optionally, when calculating the competitiveness level c of the driving style, based on the differences in the degree of attention paid to the driving goal by drivers with different driving styles, the weight parameter corresponding to each driving feature is dynamically adjusted using an elastic network regression model, wherein the elastic network regression model is set to:
[0022]
[0023] Among them, X i =[v avg , a rms , f LC , d follow ] T is the standardized vehicle driving feature vector, v avg is the average speed, a rms is the root mean square acceleration; f LC is the lane change frequency, d follow is the following distance; β ko is the base value of the target weight in the elastic net regression model; β ki is the elastic network regression coefficient; ε k is the residual term, Residual ε k The mean is 0 and the variance is Normal distribution; p is the number of standardized vehicle driving feature vectors;
[0024] When the elastic network regression model regresses to obtain weight combinations of different driving styles, a mapping relationship between the driving styles of background vehicles and the weight combinations is established as follows:
[0025]
[0026] Where N is the number of training samples, λ is the overall regularization degree, α∈[0,1], β T is the elastic net regression coefficient vector; β1 is the L1 norm, and β2 is the square of the L2 norm.
[0027] Optionally, the driving style is set to aggressive, conservative and balanced.
[0028] Optionally, in S1, the driving strategy of the dynamic Bayesian game framework is set as:
[0029] Scramble strategy: Accelerate and / or change lanes in the current lane;
[0030] Yield strategy: slow down in the current lane or maintain a safe distance from the adjacent background vehicle;
[0031] Heuristic strategy: maintain the current driving state and observe the driving states of the remaining background vehicles in the adjacent driving lane or the same driving lane.
[0032] Optionally, in S2, the multi-objective utility function u is set to:
[0033] u=ω te f te +ω sf f sf +ω ug f ug +ω ef f ef ;
[0034]
[0035] Among them, ω te is the efficiency weight, ω sf is the safety weight, ω ug is the urgency weight, ω ef are energy consumption weights, which are used to simulate the trade-off mechanism of human drivers in complex scenarios; f te is the efficiency cost, exp(·) is the exponential function operation with the natural constant e as the base, v sv is the current speed of the background vehicle, v min is the minimum speed allowed in the current interaction scenario, v max is the maximum speed allowed by the current interaction scene, ε is the smoothing factor; f sf is the safety cost, which is used to reflect the collision risk of background vehicles, TTC is the collision time, τ is the time constant, and λ d is the weight coefficient, d min is the minimum changing distance; f ugis the urgency cost, which is used to reflect the urgency of the background vehicle approaching the target point, X int is the coordinate of the lane change starting point along the road direction, X term Y is the coordinate of the lane change destination along the road direction, int Y is the coordinate of the lane change starting point perpendicular to the road direction, term is the coordinate of the lane change endpoint perpendicular to the road direction; f ef is the energy cost, which is used to quantify the energy cost of the control action, v term is the desired speed, v cur is the current speed, a max is the maximum acceleration of the background vehicle.
[0036] Optionally, in S3, the method of selecting a corresponding driving strategy from a dynamic Bayesian game framework specifically includes:
[0037] Calculating the utility values of the scramble strategy and the concession strategy respectively according to the multi-objective utility function, obtaining a scramble strategy utility value and a concession strategy utility value, and comparing the magnitude relationship between the scramble strategy utility value and the concession strategy utility value;
[0038] If the utility value of the scramble strategy is greater than the utility value of the concession strategy, then the scramble strategy is selected;
[0039] If the utility value of the scramble strategy is less than the utility value of the concession strategy, then the scramble strategy is selected;
[0040] If the utility value of the contention strategy is equal to the utility value of the concession strategy, the trial strategy is selected.
[0041] Optionally, in S4, a trajectory control instruction is generated according to the current driving state of the background vehicle and the selected driving strategy, and the control information of the trajectory control instruction includes the target speed v, acceleration a and steering angle δ of the background vehicle; and the control parameter k is used as the control parameter k. v 、k a 、k δ As influencing factors, the acceleration a and the steering angle δ of the background vehicle are adjusted in real time;
[0042] a=c·k a ·a max ·sign(k v ·v target -v)+(1-c)·(v target -v)+η·rand;
[0043] δ=c·k δ ·δ max ·sign(e y )+(1-c)·e y +ηy ·rand;
[0044] Among them, k a is the acceleration influence factor, a max is the absolute value of the maximum acceleration, sign(·) is the operation sign function, k v is the target speed influence factor, v target is the target speed, v is the current speed, k δ is the steering angle influencing factor, δ max is the maximum steering angle, e y is the lateral distance deviation, η and η y All are random perturbation amplitudes, η=0.2a max , η y =0.1δ max , rand∈[-1, 1] is an evenly distributed random number.
[0045] The beneficial effects that the present invention can produce include:
[0046] 1. This method dynamically generates multi-dimensional driving behaviors: By integrating driving style clustering with a dynamic multi-objective weighting adaptation mechanism, it overcomes the technical bottleneck of traditional methods that rely on a single background vehicle behavior pattern. It can autonomously adjust the weighting of objectives such as safety, efficiency, and urgency based on real-time scene characteristics, generating a continuous spectrum of driving styles ranging from conservative to aggressive, significantly improving the diversity and realism of background vehicle behaviors in test scenarios.
[0047] 2. This method enables precise modeling of complex interaction logic: Based on a dynamic Bayesian game framework, it achieves a deep coupling of vehicle interaction strategies and competitive level parameters. Background vehicles can autonomously choose to compete, yield, or explore strategies based on the state of their surroundings. They exhibit human-like decision-making hesitation and strategy switching in scenarios such as intersection merging and lane maneuvering, addressing the mechanical nature of interaction behavior in existing technologies.
[0048] 3. This method effectively improves the verifiability of anthropomorphism: Through the Turing test, an objective and quantitative anthropomorphism evaluation system was established. Double-blind human-computer interaction tests verified the naturalness of behavioral patterns, providing a more reliable benchmark for evaluating the performance of autonomous driving systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flowchart of the method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to the present invention;
[0050] Figure 2 This is a flowchart of the driving feature extraction and personalized behavior modeling of the present invention;
[0051] Figure 3 This is a flow chart of the Bayesian game strategy generation and decision optimization of the present invention. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] At present, with the rapid development of autonomous driving technology, the testing phase is the key to ensuring its safety and reliability. However, in existing autonomous driving test scenarios, there are significant defects in the generation of background vehicle trajectories, such as the rough personalized modeling of driver behavior style, which makes it difficult to accurately simulate the diversity of human driving; the rigid multi-objective trade-off mechanism, which cannot flexibly respond to complex and changing traffic environments; and the lack of anthropomorphic evaluation methods, which leads to a disconnect between simulation results and real driving scenarios. To solve the above defects, refer to Figure 1-Figure 3 As shown, the present invention proposes a method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios to achieve dynamic generation and verification of anthropomorphic interactive behaviors of background vehicles, which includes the following steps:
[0054] Step 1: According to the set interaction logic, driving characteristics from several sets of interactive scenarios are extracted and simulated to build a driving simulator. These interactive scenarios are rich and diverse, including but not limited to lane change game scenarios, in which vehicles must find the right time to change lanes in the traffic flow, involving speed and spacing games with surrounding vehicles; intersection merging scenarios, in which vehicles must safely and efficiently merge into the main road traffic flow in complex intersection traffic flows; and congested following scenarios, in which vehicles maintain an appropriate following distance and speed in low-speed, dense traffic to avoid collisions. The driving simulator uses the driving characteristics from each set of interactive scenarios as driving behavior data for several background vehicles and assigns a set driving style to each of them. A corresponding driving strategy is then assigned to each driving style to build a dynamic Bayesian game framework.
[0055] In the above, driving characteristics include one or more of the vehicle's acceleration distribution (reflecting the frequency and intensity of vehicle acceleration / deceleration), lane change frequency (reflecting the vehicle's activeness in changing lanes), following distance (directly related to driving safety), steering wheel angle (reflecting the vehicle's steering operation) and pedal travel (covering key information such as the operating range of the accelerator pedal and brake pedal) within any detection cycle.
[0056] In the above, the method for setting a driving style includes: using a clustering algorithm to analyze the driving behavior data of each background vehicle, and through several iterative calculations, dividing the driving behavior data of the background vehicles into corresponding clusters based on their characteristic similarities, and using each cluster as a driving style; the driving style can be set as aggressive, conservative, and balanced, wherein the competitiveness level c of the corresponding driving style is quantified by the center value of the cluster;
[0057] c=w1·F norm +w2·D norm +w3·A norm ;
[0058]
[0059]
[0060] Where c∈(0,1); F is the lane change frequency, F min is the minimum lane change frequency, F max is the maximum lane change frequency, F norm is the normalized lane change frequency, w1 is the characteristic weight of the lane change frequency; D is the average following distance, D min is the minimum average following distance, D max is the maximum average following distance, D norm is the normalized average following distance, w2 is the feature weight of the average following distance; A is the acceleration variance, A min is the minimum acceleration variance, A max is the maximum acceleration variance, A norm is the normalized acceleration variance, and w3 is the characteristic weight of the acceleration variance. Background vehicles are categorized into corresponding driving style categories based on the competitiveness level c of the driving style. For example, drivers with an aggressive driving style tend to frequently change lanes, accelerate and decelerate quickly, and pursue driving efficiency. Consequently, they exhibit high lane change frequency, short following distances, and high acceleration variance. Drivers with a conservative driving style prioritize driving safety and exhibit low lane change frequency, a large following distance, and gentle acceleration. A balanced driving style seeks a balance between efficiency and safety.
[0061] Furthermore, when calculating the competitiveness level c of driving styles, the weight parameters corresponding to each driving characteristic are dynamically adjusted based on the differences in the degree of attention paid by drivers with different driving styles to driving goals (such as safety, efficiency, urgency, and energy consumption). The elastic network regression model is set as follows:
[0062]
[0063] Among them, Xi =[v avg , a rms , f LC , d follow ] T is the standardized vehicle driving feature vector, v avg is the average speed, a rms is the root mean square acceleration, which is a characteristic quantity that measures the amplitude of acceleration fluctuation; f LC is the lane change frequency, d follow is the following distance; β ko is the base value of the target weight in the elastic net regression model, used to adjust the weight baseline; β ki is the elastic network regression coefficient; ε k is the residual term, Residual ε k The mean is 0 and the variance is Normal distribution; p is the number of standardized vehicle driving feature vectors;
[0064] When the elastic network regression model regresses the weighted combinations of different driving styles, the mapping relationship between the driving styles of background vehicles and the weighted combinations is established as follows:
[0065]
[0066] Where N is the number of training samples, λ is the overall regularization degree (value is 0.1), α∈[0,1], β T is the elastic network regression coefficient vector, obtained by optimizing the objective function and used to establish the mapping relationship between driving style characteristics and multi-objective weights; β1 is the L1 norm, and β2 is the square of the L2 norm; this facilitates the subsequent rapid matching of driving styles and corresponding weight parameters based on the driving behavior data of background vehicles.
[0067] In the above, according to the driving style set by the background vehicle, a corresponding driving strategy is configured for each category in the dynamic Bayesian game framework. Specifically, the driving strategy of the dynamic Bayesian game framework is set as:
[0068] Aggressive Scramble Strategy: Accelerate and / or change lanes in the current lane. When a background vehicle detects a favorable driving opportunity, such as a large gap in the lane ahead or slower speeds of surrounding vehicles that affect its own driving, it may employ a scramble strategy, rapidly accelerating or changing lanes to secure a more favorable position and improve driving efficiency.
[0069] Yield Strategy (Conservative): Slow down in the current lane or maintain a safe distance from adjacent background vehicles. When encountering potentially dangerous situations, such as a sudden deceleration of the vehicle ahead or a forced lane change attempt by a vehicle in an adjacent lane, the background vehicle will adopt a yield strategy, reducing speed and increasing following distance to ensure safe driving.
[0070] Heuristic Strategy (Balanced): Maintains the current driving state while observing the driving status of other background vehicles in adjacent lanes or the same lane. In uncertain traffic situations, such as when the road ahead is unclear or the behavior of surrounding vehicles is unpredictable, background vehicles will first adopt a heuristic strategy, maintaining their current state and gathering more information before making a decision.
[0071] Step 2: Obtain the driving styles of background vehicles in any set of interactive scenarios in real time, and use this as a basis for iterative training of the multi-objective utility function. The multi-objective utility function aims to simulate the trade-offs made by human drivers in complex scenarios by comprehensively considering multiple key dimensions such as efficiency, safety, urgency, and energy consumption, making the decisions of background vehicles more closely resemble real human driving behavior. Specifically, the multi-objective utility function u is set as:
[0072] u=ω te f te +ω sf f sf +ω ug f ug +ω ef f ef ;
[0073]
[0074] Among them, ω te is the efficiency weight, ω sf is the safety weight, ω ug is the urgency weight, ω ef are energy consumption weights, which are used to simulate the trade-off mechanism of human drivers in complex scenarios; f te is the efficiency cost, exp(·) is the exponential function operation with the natural constant e (about 2.71828) as the base, v sv is the current speed of the background vehicle, v min is the minimum speed allowed in the current interaction scenario, v max is the maximum speed allowed by the current interaction scene, ε is the smoothing factor; f sf is the safety cost, which is used to reflect the collision risk of background vehicles. TTC is the collision time, which is the time required from the current moment to the collision between the vehicle and the obstacle while maintaining the current speed and acceleration. τ is the time constant, and λ is the collision time. dis the weight coefficient. The larger its value is, the more significant the impact on the safety cost is, indicating that the system has stricter safety constraints on the minimum distance. min is the minimum changing distance; f ug is the urgency cost, which is used to reflect the urgency of the background vehicle approaching the target point, X int is the coordinate of the lane change starting point along the road direction, X term Y is the coordinate of the lane change destination along the road direction, int Y is the coordinate of the lane change starting point perpendicular to the road direction, term is the coordinate of the lane change endpoint perpendicular to the road direction; f ef is the energy cost, which is used to quantify the energy cost of the control action, v term is the desired speed, v cur is the current speed, a max is the maximum acceleration of the background vehicle. By continuously optimizing and adjusting these parameters during iterative training, we ensure that the multi-objective utility function can adapt to different driving styles, enabling background vehicles to make reasonable decisions in various scenarios and enhancing the flexibility and adaptability of the strategy.
[0075] Step 3: Use the multi-objective utility function to calculate the benefit value corresponding to the driving style of the background vehicle, and select the corresponding driving strategy from the dynamic Bayesian game framework as the optimal response action of the background vehicle based on the calculated benefit value. Specifically, it includes: calculating the utility value of the competition strategy and the concession strategy respectively according to the multi-objective utility function, and obtaining the competition strategy utility value u(A F ) and the utility value of the concession strategy u(A Y ), and compare the utility value of the competition strategy u(A F ) and the utility value of the concession strategy u(A Y ) between the size relationship; Among them, if the competition strategy utility value u(A F ) is greater than the utility value u(A Y ), it means that in the current situation, the use of the contention strategy can enable the background vehicles to obtain higher comprehensive benefits, such as reaching the destination faster and using road resources more efficiently, so the contention strategy is selected; if the contention strategy utility value u(A F ) is less than the utility value u(A Y ), it means that the concession strategy is more in line with the current scenario requirements, can better ensure driving safety and avoid potential dangers, and the concession strategy is selected at this time; if the competition strategy utility value u(A F ) is equal to the utility value of the concession strategy u(A Y ), it means that the comprehensive benefits of the two strategies are comparable under the current circumstances, and it is difficult to directly judge the pros and cons. In this case, a trial strategy is chosen. By maintaining the current driving state and further observing the surrounding traffic conditions, a more sufficient basis for subsequent decision-making is provided.
[0076] Step 4: Generate trajectory control instructions based on the selected driving strategy, and use the random perturbation mechanism to perform anthropomorphic optimization on the trajectory control instructions to obtain the anthropomorphic trajectory of the background vehicle. Specifically, based on the current driving state of the background vehicle and the selected driving strategy, the trajectory control instructions are generated. The control information of the trajectory control instructions includes the target speed v, acceleration a, and steering angle δ of the background vehicle. In order to make the generated trajectory more anthropomorphic, a random perturbation mechanism is introduced to control the parameter k v 、k a 、k δ The acceleration a and steering angle δ of the background vehicle are adjusted in real time as influencing factors; the trajectory control instructions are optimized in an anthropomorphic manner to obtain the final anthropomorphic trajectory.
[0077] a=c·k a ·a max ·sign(k v ·v target -v)+(1-c)·(v target -v)+η·rand;
[0078] δ=c·k δ ·δ max ·sign(e y )+(1-c)·e y +η y ·rand;
[0079] Among them, k a is the acceleration influence factor, a max is the absolute value of the maximum acceleration; sign(·) is the sign function, which returns 1 when a positive number is input, 0 when a 0 is input, and -1 when a negative number is input; k v is the target speed influence factor, v target is the target speed, v is the current speed, k δ is the steering angle influencing factor, δ max is the maximum steering angle, e y is the lateral distance deviation, η and η y All are random perturbation amplitudes, η=0.2a max , η y =0.1δ max , rand∈[-1, 1] is an evenly distributed random number used to simulate the uncertainty of the driver's throttle and steering wheel control.
[0080] In this embodiment, different driving strategies correspond to different control parameter settings, as follows:
[0081] (1) Control parameter k of the scramble strategy v 、k a、k δ Set to:
[0082] k v =1+0.5c;k a =1+0.6c;k δ =1+0.6c;
[0083] (2) Control parameter k of the concession strategy v 、k a 、k δ Set to:
[0084] k v =1-0.3c;k a =1-0.4c;k δ =1-0.4c;
[0085] (3) Control parameter k of the trial strategy v 、k a 、k δ Both are set to 1.
[0086] Step 5. In order to scientifically and accurately verify the degree of anthropomorphism of the background vehicle driving behavior data, the present invention uses the classic Turing test method to verify the anthropomorphism level of the anthropomorphic trajectory of the background vehicle. Specifically, testers with driving experience are recruited to interact with the background vehicle in real time in typical interactive scenarios (such as lane change games, intersection merging, congested car following, etc.). It should be noted that the testers need to rely on their own driving experience and intuition to judge whether the background vehicle is under human control and evaluate its driving style (aggressive / conservative / balanced). Based on the test results, the differences between the driving behavior data of the background vehicle and the real human driving behavior are deeply analyzed, and the driving simulator is iteratively optimized. By adjusting the algorithm parameters, such as the weight parameters in the elastic network regression model, the weight coefficients in the multi-objective utility function, etc., and optimizing the strategy selection mechanism, the driving behavior of the background vehicle in complex traffic scenarios is made more realistic and natural, and the gap between the simulation results and the real scene is gradually narrowed. It is ensured that the driving behavior and trajectory of the background vehicle in the simulated test scene are highly close to the real situation, providing more reliable and effective data support for autonomous driving testing.
[0087] Example 1
[0088] The present invention uses Carla autonomous driving simulation software to generate and test background vehicle trajectories.
[0089] Before the simulation environment is started, a high-precision driving simulator is used to collect driving behavior data of human drivers. Specifically, 50 experienced drivers are recruited to participate in the simulated driving experiment. Among them, each driver needs to complete 10 repeated tests in typical scenarios such as lane change games, intersection merging, and congested following. The simulator records key feature data such as speed sequence, acceleration distribution, lane change frequency, and following distance at a sampling frequency of 100Hz. The core features of the driving behavior data are calculated: acceleration variance A, average following distance D, and lane change frequency F, and the above three core features are normalized to eliminate dimensional differences and unify the direction to positive (that is, the larger the value, the stronger the competitiveness); then the improved K-means clustering algorithm (number of clusters = 3, 100 iterations) is used to classify the driving style, and finally the driving behavior data is divided into three types of driving styles:
[0090] Radical (32%): F norm ≥0.7, D norm ≥0.7, A norm ≥0.65, c≥0.68;
[0091] Conservative (41%): F norm ≤0.3, D norm ≤0.3, A norm ≤0.35, c≤0.32;
[0092] Balanced type (27%): The characteristic values of acceleration variance A, average following distance D, and lane change frequency F are distributed in the middle, 0.32<c<0.68.
[0093] Based on the clustering results, the elastic network regression model is used to dynamically generate weight parameters for the multi-objective utility functions corresponding to the three types of driving styles. The specific weight ranges are:
[0094] Aggressive type (c ≥ 0.68): efficiency weight ω te =0.68, safety weight ω sf =0.12, urgency weight ω ug =0.09, energy consumption weight ω ef =0.11;
[0095] Conservative (c≤0.32): efficiency weight ω te =0.18, safety weight ω sf =0.72, urgency weight ω ug =0.06, energy consumption weight ω ef =0.04;
[0096] Balanced type (0.32<c<0.68): efficiency weight ω te =0.25, safety weight ωsf =0.25, urgency weight ω ug =0.25, energy consumption weight ω ef =0.25.
[0097] When the simulation environment is started, an initial c value is assigned to each background vehicle, and then a baseline value is set according to the vehicle type and dynamically corrected in combination with the initial position. Specifically, the baseline value c for a car is base =0.5, truck baseline value c base =0.3, sports car c base =0.7; vehicles c in the first 20% of the traffic flow init =c base +0.1rand, vehicles in the bottom 20% c init =c base -0.1rand, the remaining vehicles c init =c base +0.05rand. Where rand∈[-0.05, 0.05] is a uniformly distributed random number. For example, a sports car in the first 20% of the traffic flow has an initial competitive value c init =0.7+0.1×0.3=0.73, which is an aggressive type; a truck 20% behind the traffic flow has an initial competitive value c init =0.3-0.1×0.2=0.28, which is conservative.
[0098] During driving, background vehicles dynamically adjust the competitive level c value based on real-time driving data. For example, in a lane change scenario, if the vehicle successfully seizes the target lane c new =c current +0.05×(1-c current ), forced to give up changing lanes c new =c current -0.03×c current The background vehicle can match the driving style and the corresponding multi-objective utility function weight parameters according to the current competitiveness c. During the driving process, the background vehicle makes decisions through the dynamic Bayesian game framework. For example, when the background vehicle A (aggressive, c = 0.82) and the autonomous driving test vehicle B are running parallel on the highway, when the background vehicle A detects that the autonomous driving test vehicle B intends to change lanes, the utility value of each strategy is calculated to obtain: the competition strategy u (A F )=0.76, concession strategy u(A Y )=0.34, then the background vehicle A chooses the scramble strategy and accelerates to v target =k v ·v current =1.41v current , and ultimately successfully prevented the autonomous driving test vehicle B from changing lanes.
[0099] Anthropomorphic Validation Phase: 30 professional drivers (with 5 years of driving experience or more) were recruited. Each driver interacted with a total of 100 background vehicles (32 aggressive, 41 conservative, and 27 balanced) in real-time across three typical scenarios: changing lanes on a highway, merging at an intersection without a signal, and following a vehicle in traffic. Each test round lasted three minutes. Test results showed that the average proportion of background vehicles misidentified as human drivers (behavior confusion) reached 79.9%, with aggressive vehicles having the highest confusion (83.5%), conservative vehicles at 76.8%, and balanced vehicles at 79.3%. These data demonstrate that the anthropomorphic trajectories of background vehicles generated by the algorithm closely resemble real-world driving habits, validating the effectiveness and reliability of the method.
[0100] Example 2
[0101] The difference between this embodiment and embodiment 1 is that the background vehicle trajectory is generated and tested and verified using Unity autonomous driving simulation software. At the same time, according to different test requirements, when the background vehicle accesses V2X (Vehicle-to-Everything) communication, its c-value change trend can be broadcast to the vehicle under test in real time. This function is of great significance in actual testing. For example, in a complex traffic intersection scenario, the driving intention of the background vehicle may change rapidly with the changes in the surrounding environment. By sending the c-value change trend to the vehicle under test, the responsiveness of the system under test to the changes in the intention of the background vehicle can be tested, and the decision-making accuracy and timeliness of the autonomous driving system in the face of dynamic traffic environments can be evaluated, meeting the diverse needs of autonomous driving system performance evaluation in different test scenarios.
Claims
1. A method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios, characterized by: The steps include: S1. Extracting driving characteristics from several sets of interactive scenarios according to a predetermined interactive logic and conducting simulation training to construct a driving simulator. The driving simulator uses the driving characteristics from each set of interactive scenarios as driving behavior data for several background vehicles and configures a predetermined driving style for each of them. A corresponding driving strategy is then set for each driving style to construct a dynamic Bayesian game framework. S2. Real-time acquisition of the driving style of the background vehicles in any set of interactive scenarios, and iterative training of the multi-objective utility function. S3. Calculating a benefit value corresponding to the driving style of the background vehicle using the multi-objective utility function, and selecting a corresponding driving strategy from a dynamic Bayesian game framework as the optimal response action of the background vehicle based on the calculated benefit value; S4. Generate a trajectory control instruction according to the selected driving strategy, and perform anthropomorphic optimization on the trajectory control instruction using a random perturbation mechanism to obtain an anthropomorphic trajectory of the background vehicle.
2. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: In S1, the interaction scenario includes at least one of lane change game, intersection merging, and congestion following.
3. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: In S1, the driving characteristics include one or more of the vehicle's acceleration distribution, lane change frequency, following distance, steering wheel angle, and pedal travel within any detection period.
4. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: In S1, the method for setting the driving style includes: Analyzing the driving behavior data of each background vehicle using a clustering algorithm, dividing the driving behavior data of the background vehicles into corresponding clusters based on their characteristic similarities through several iterative calculations, and taking each cluster as a driving style; The competitiveness level c of the corresponding driving style is quantified by the center value of the cluster: c=w1·F norm +w2·D norm +w3·A norm ; Where c∈(0,1); F is the lane change frequency, F min is the minimum lane change frequency, F max is the maximum lane change frequency, F norm is the normalized lane change frequency, w1 is the characteristic weight of the lane change frequency; D is the average following distance, D min is the minimum average following distance, D max is the maximum average following distance, D norm is the normalized average following distance, w2 is the feature weight of the average following distance; A is the acceleration variance, A min is the minimum acceleration variance, A max is the maximum acceleration variance, A norm is the normalized acceleration variance, w3 is the feature weight of the acceleration variance; The background vehicles are classified into corresponding driving style categories based on the competitiveness level c of the driving style.
5. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 4, characterized in that: When calculating the competitiveness level c of driving style, the weight parameters corresponding to each driving characteristic are dynamically adjusted using the elastic network regression model, based on the fact that drivers with different driving styles have different levels of attention to the driving goal. The elastic network regression model is set as: Among them, X i =[v avg , a rms , f LC , d follow ] T is the standardized vehicle driving feature vector, v avg is the average speed, a rms is the root mean square acceleration; f LC is the lane change frequency, d follow is the following distance; β ko is the base value of the target weight in the elastic net regression model; β ki is the elastic network regression coefficient; ε k is the residual term, Residual ε k The mean is 0 and the variance is Normal distribution; p is the number of standardized vehicle driving feature vectors; When the elastic network regression model regresses to obtain weight combinations of different driving styles, a mapping relationship between the driving styles of background vehicles and the weight combinations is established as follows: Where N is the number of training samples, λ is the overall regularization degree, α∈[0,1], β T is the elastic net regression coefficient vector; β1 is the L1 norm, and β2 is the square of the L2 norm.
6. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: The driving style is set to aggressive, conservative, and balanced.
7. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: In S1, the driving strategy of the dynamic Bayesian game framework is set as: Scramble strategy: Accelerate and / or change lanes in the current lane; Yield strategy: slow down in the current lane or maintain a safe distance from the adjacent background vehicle; Heuristic strategy: maintain the current driving state and observe the driving states of the remaining background vehicles in the adjacent driving lane or the same driving lane.
8. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: In S2, the multi-objective utility function u is set as: u=ω te f te +oh sf f sf +oh ug f ug +oh ef f ef ; Among them, ω te is the efficiency weight, ω sf is the safety weight, ω ug is the urgency weight, ω ef are energy consumption weights, which are used to simulate the trade-off mechanism of human drivers in complex scenarios; f te is the efficiency cost, exp(·) is the exponential function operation with the natural constant e as the base, v sv is the current speed of the background vehicle, v min is the minimum speed allowed in the current interaction scenario, v max is the maximum speed allowed by the current interaction scene, ε is the smoothing factor; f sf is the safety cost, which is used to reflect the collision risk of background vehicles, TTC is the collision time, τ is the time constant, and λ d is the weight coefficient, d min is the minimum changing distance; f ug is the urgency cost, which is used to reflect the urgency of the background vehicle approaching the target point, X int is the coordinate of the lane change starting point along the road direction, X term Y is the coordinate of the lane change destination along the road direction, int Y is the coordinate of the lane change starting point perpendicular to the road direction, term is the coordinate of the lane change endpoint perpendicular to the road direction; f ef is the energy cost, which is used to quantify the energy cost of the control action, v term is the desired speed, v cur is the current speed, a max is the maximum acceleration of the background vehicle.
9. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: In S3, the method of selecting a corresponding driving strategy from the dynamic Bayesian game framework specifically includes: Calculating the utility values of the scramble strategy and the concession strategy respectively according to the multi-objective utility function to obtain a scramble strategy utility value and a concession strategy utility value, and comparing the magnitude relationship between the scramble strategy utility value and the concession strategy utility value; If the utility value of the scramble strategy is greater than the utility value of the concession strategy, then the scramble strategy is selected; If the utility value of the scramble strategy is less than the utility value of the concession strategy, then the scramble strategy is selected; If the utility value of the contention strategy is equal to the utility value of the concession strategy, the trial strategy is selected.
10. The method for generating anthropomorphic trajectories of background vehicles for autonomous driving test scenarios according to claim 1, characterized in that: In S4, a trajectory control instruction is generated according to the current driving state of the background vehicle and the selected driving strategy. The control information of the trajectory control instruction includes the target speed v, acceleration a and steering angle δ of the background vehicle; and the control parameter k is used as the control parameter. v 、k a 、k δ As influencing factors, the acceleration a and the steering angle δ of the background vehicle are adjusted in real time; a=c·k a ·and max ·sign(k v ·in target -v)+(1-c)·(v target -v)+η·rand; δ=c·k δ ·d max ·sign(e y )+(1-c)·e y +n y ·rand; Among them, k a is the acceleration influence factor, a max is the absolute value of the maximum acceleration, sign(·) is the operation sign function, k v is the target speed influence factor, v target is the target speed, v is the current speed, k δ is the steering angle influencing factor, δ max is the maximum steering angle, e y is the lateral distance deviation, η and η y All are random perturbation amplitudes, η=0.2a max , η y =0.1δ max , rand∈[-1, 1] is an evenly distributed random number.
Citation Information
Patent Citations
Lane changing decision-making method and system for autonomous vehicle based on rolling game
CN110297494A
Automatic driving vehicle lane changing behavior vehicle road collaborative decision algorithm based on Bayesian game
CN115056798A
Intelligent vehicle intersection straight-going speed decision-making method fusing driving style game
CN115352449A
Automatic driving simulation test scene generation method introducing driver style
CN115952734A
Personalized driver-based anthropomorphic lane changing trajectory optimization method
CN116534055A
Cited By
Method and system for generating man-machine cooperation conflict test case
CN120994569A