Multidimensional natural driving scene generation and simulation test method based on reinforcement learning

Through the method based on reinforcement learning, a fuzzy driving style rules are constructed and an intelligent body strategy network is trained, which solves the problem of insufficient natural driving scenarios and behavior generation in existing simulation tests, and realizes diversified driving behavior pattern generation, which improves simulation test efficiency and the ability of autonomous driving system research and development.

CN120012589AActive Publication Date: 2025-05-16XIDIAN UNIV

Patent Information

Application Number
CN202510117815.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

When generating natural driving scenarios and behaviors, existing simulation testing methods have problems such as insufficient coverage of real-world complex scenarios and limited modeling of driving behavior characteristics, which affect the effectiveness of simulation testing.

Method used

Using a method based on reinforcement learning, a multi-dimensional natural driving scenario is generated by constructing fuzzy driving style rules, combining dynamic characteristics, rule compliance, psychological characteristics and collaboration methods, and using reinforcement learning algorithms to train the intelligent body strategy network.

Benefits of technology

It has achieved the generation of diverse behavioral patterns covering from cautious to radical driving, and obtained more complete and accurate behavioral feature representations, which has improved the effectiveness of simulation testing and the R&D and verification capabilities of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012589A_ABST
    Figure CN120012589A_ABST
Patent Text Reader

Abstract

The invention relates to a reinforcement learning-based multi-dimensional natural driving scene generation and simulation test method, which comprises the following steps of: quantizing driving behavior characteristics by combining a fuzzy logic classification model on the basis of dynamic characteristics, rule compliance, psychological characteristics and a cooperation mode, and constructing to obtain a driving style fuzzy rule; constructing a driving environment and configuring quantitative characteristics of different driving styles according to the driving style fuzzy rule; by utilizing a reinforcement learning algorithm, based on the quantitative characteristics of the driving environment and different driving styles, with a driving style fuzzy rule as a target, calculating rewards of different driving scenes according to a reward function so as to train and update the intelligent agent strategy network, and obtaining a trained intelligent agent strategy network; the trained agent strategy network is used for generating a multi-dimensional natural driving scene and performing a simulation test. According to the method, diversified behavior modes covering cautious driving to aggressive driving can be generated, and more complete and accurate behavior feature representation can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and specifically relates to a multi-dimensional natural driving scene generation and simulation test method based on reinforcement learning. Background Art

[0002] In recent years, with the rapid development of intelligent driving technology, real vehicle testing plays a key role in verifying the performance and safety of autonomous driving systems. However, real vehicle testing faces major challenges such as high cost, high risk, and difficulty in fully covering complex scenarios. To address these issues, simulation testing, as an efficient and flexible alternative, is receiving widespread attention from the research and industry communities.

[0003] Simulation testing can reproduce real driving scenarios through a virtual environment, and has significant advantages such as low cost, strong controllability, and high repeatability. For example, in complex traffic scenarios, the perception system and decision-making algorithm of autonomous driving vehicles can be tested. Through simulation tools such as Carla and LGSVL Simulator, extreme conditions such as rainy and snowy weather, low visibility, and high-density traffic can be simulated, which are difficult to achieve in actual testing. In addition, Waymo's research also shows that billions of miles of autonomous driving data can be generated every day through simulation testing, and the cost of collecting this data is much lower than that of real vehicle testing, and it is easier to find boundary problems in long-tail scenarios.

[0004] However, simulation testing is not perfect, and its limitations are mainly reflected in the deviation between the simulation model and the real world. Since the algorithms of simulation tools are based on idealized assumptions, they may not fully reflect the dynamic complexity of real driving, such as the random behavior of traffic participants or the uncertainty of the environment. In addition, the confidence of the simulation system itself depends on the accuracy of the underlying model and the data coverage. Existing simulation test methods and scenario generation technologies face challenges in practical applications: first, insufficient coverage of complex driving scenarios in the real world, affecting the authenticity and universality of the simulation; second, the modeling of driving behavior characteristics is limited, resulting in a lack of diversity and naturalness in the generated behavior. These problems restrict the effectiveness of simulation testing and hinder the development of autonomous driving systems. Summary of the invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a multi-dimensional natural driving scene generation and simulation test method based on reinforcement learning. The technical problem to be solved by the present invention is achieved by the following technical solutions:

[0006] The embodiment of the present invention provides a multi-dimensional natural driving scene generation and simulation test method based on reinforcement learning, comprising the steps of:

[0007] S1. Based on dynamic characteristics, rule compliance, psychological characteristics and cooperation mode, the driving behavior characteristics are quantified by combining the fuzzy logic classification model to construct the driving style fuzzy rules;

[0008] S2, constructing a driving environment and configuring quantitative features of different driving styles according to the driving style fuzzy rules;

[0009] S3. Using a reinforcement learning algorithm, based on the quantitative characteristics of the driving environment and different driving styles, taking the driving style fuzzy rules as the target, calculating the rewards of different driving scenarios according to the reward function to train and update the intelligent agent policy network, and obtaining a trained intelligent agent policy network; the trained intelligent agent policy network is used to generate multi-dimensional natural driving scenarios and perform simulation tests.

[0010] In one embodiment of the present invention, step S1 comprises:

[0011] S11, obtaining a vehicle kinematic state index, and using a fuzzy subset to define fuzzy values ​​of different vehicle kinematic state indexes to obtain a vehicle dynamic state fuzzy index;

[0012] S12, obtaining a traffic compliance state index, and using a fuzzy subset to define fuzzy values ​​of different traffic compliance state indexes to obtain a traffic compliance state fuzzy index;

[0013] S13. Calculate the risk aversion coefficient, risk decision threshold and risk conversion coefficient and weight them to construct a comprehensive risk preference index;

[0014] S14. Calculate the collaboration initiative coefficient and information sharing degree and weight them to construct a comprehensive index of collaboration mode;

[0015] S15. Construct a driving style fuzzy rule according to the vehicle dynamics state fuzzy index, the traffic compliance state fuzzy index, the risk preference comprehensive index and the cooperation mode comprehensive index.

[0016] In one embodiment of the present invention, the vehicle kinematic state index includes speed, acceleration, steering angle and yaw rate; wherein the fuzzy value of the speed includes low speed, medium speed and high speed, the fuzzy value of the acceleration includes low acceleration, medium acceleration and high acceleration, the fuzzy value of the steering angle includes low steering angle and high steering angle, and the fuzzy value of the yaw rate includes low yaw rate and high yaw rate;

[0017] The traffic compliance status indicators include speeding status, following distance safety status and lane keeping standard status. The fuzzy value of the speeding status includes speeding and not speeding. The fuzzy value of the following distance safety status includes following distance safety and following distance unsafe. The lane keeping standard status includes good lane keeping, average lane keeping and poor lane keeping.

[0018] In one embodiment of the present invention, the calculation formula of the risk aversion coefficient is:

[0019]

[0020] Among them, a x is the longitudinal acceleration, a y is the lateral acceleration, δ is the steering angle, v is the speed; α1 and α2 are weight coefficients;

[0021] The calculation formula of the risk decision threshold is:

[0022]

[0023] Among them, DHW, THW, TTC are following distance, headway time, and collision time, respectively, and β1, β2, and β3 are weight coefficients;

[0024] The calculation formula of the risk conversion coefficient is:

[0025]

[0026] Among them, Δa x ,Δa y ,Δv,Δσ, They are the longitudinal acceleration change, lateral acceleration change, speed change, steering angle change and yaw rate change respectively;

[0027] The comprehensive risk preference indicator is:

[0028] R=w1×R a +w2×R d +w3×R RTC

[0029] Among them, w1, w2, w3 are weights.

[0030] In one embodiment of the present invention, the calculation formula of the collaboration initiative coefficient is:

[0031]

[0032] Among them, y,d c are the lateral position and the lateral distance to the adjacent vehicle, respectively, γ1 and γ2 are weight coefficients;

[0033] The calculation formula of the information sharing degree is:

[0034] I s =ω1log(v)+ω2ρ

[0035] Among them, v is the speed, ρ is the traffic flow density, ω1, ω2 are weight coefficients;

[0036] The comprehensive indicators of the collaboration mode are:

[0037] CI=μ1×C i +μ2×I s

[0038] Among them, μ1, μ2 are weight coefficients.

[0039] In one embodiment of the present invention, the driving style fuzzy rules include: very defensive, high priority defensive, low priority defensive, high priority normal movement, low priority normal movement, very sporty, high priority aggressive and low priority aggressive.

[0040] In one embodiment of the present invention, step S2 comprises:

[0041] A driving environment is constructed in a simulation platform and quantitative features of different driving styles are configured according to the driving style fuzzy rules, wherein the driving environment includes complex road conditions, dynamic traffic flow and multi-type vehicle models.

[0042] In one embodiment of the present invention, step S3 includes:

[0043] Acquire the real-time state information, historical state trajectory and environmental information of the vehicle, and use a Transformer encoder to encode the real-time state information, the historical state trajectory and the environmental information to obtain the vehicle state at the current moment;

[0044] Based on the vehicle state at the current moment, a driving behavior action is generated using an Actor network, the driving behavior action is executed, the vehicle state at the next moment is obtained, and a reward for the current driving scenario is calculated using a multi-objective reward function;

[0045] Based on the time difference method, the Critic network is used to evaluate the vehicle state and the driving behavior at the current moment; wherein the Actor network and the Critic network form the agent strategy network;

[0046] Using the MAPPO reinforcement learning algorithm to update the Actor network and the Critic network;

[0047] The weights in the multi-objective reward function are adjusted to optimize the agent policy network.

[0048] In one embodiment of the present invention, the real-time state information includes: vehicle speed, acceleration, yaw rate, steering angle, speeding state, following distance state and lane keeping state; the historical state trajectory includes the state of the previous time step; the environmental information includes road type and traffic flow;

[0049] The driving behavior action includes continuous output and discrete output. The continuous output includes acceleration and steering angle, and the discrete output includes whether to change lanes, slow down or accelerate.

[0050] In one embodiment of the present invention, the multi-objective reward function is:

[0051] R=p1·R1+p2·R2+p3·R3+p4·R4+p5·R5

[0052] Among them, R1 is the comprehensive dynamics reward, R2 is the comprehensive traffic compliance reward, R3 is the risk preference reward, R4 is the collaborative mode reward, R5 is the driving style matching reward, and p1, p2, p3, p4, and p5 are weight coefficients;

[0053] R1=w speed ·R speed +w acc ·R acc +w yaw ·R Ryaw +w steering ·R steering

[0054]

[0055] Among them, R speed is the speed reward, R acc is the acceleration reward, R Ryaw is the yaw rate reward, R steering is the steering angle reward, w speed 、w acc 、w yaw 、w steering is the weight coefficient;

[0056] R2=w overspeed ·R overspeed +w distance ·R distance +w lane ·R lane

[0057]

[0058] Among them, Roverspeed For speeding reward, R distance is the following distance reward, R lane lane keeping reward; w overspeed 、w distance 、w lane is the weight coefficient;

[0059]

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] The method of the present invention introduces risk preference and naturalness indicators of collaborative mode, and constructs fuzzy rules of driving style based on dynamic characteristics, rule compliance, psychological characteristics and collaborative mode. Then, the reward function in the reinforcement learning framework is used to train and update the agent strategy network to achieve dynamic generation of the driving behavior model. This method can generate a variety of behavioral patterns covering from cautious to aggressive driving, obtain a more complete and accurate representation of behavioral characteristics, and thus generate a variety of driving behaviors, providing scientific and efficient technical support for the research and development, verification and future technical standardization, safety assessment and large-scale application of autonomous driving systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 A flowchart of a method for generating and simulating a multi-dimensional natural driving scene based on reinforcement learning provided by an embodiment of the present invention;

[0063] Figure 2 A schematic diagram of a driving style modeling framework with multi-dimensional features provided by an embodiment of the present invention;

[0064] Figure 3 The driving style scenario generation and simulation process of reinforcement learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0066] Embodiment 1

[0067] See also Figure 1 , Figure 1 A flowchart of a method for generating and simulating a multi-dimensional natural driving scene based on reinforcement learning provided by an embodiment of the present invention. The method comprises the following steps:

[0068] S1. Based on dynamic characteristics, rule compliance, psychological characteristics and cooperation mode, the driving behavior characteristics are quantified by combining the fuzzy logic classification model, and the driving style fuzzy rules are constructed.

[0069] See also Figure 2 , Figure 2 A schematic diagram of a multi-dimensional MDRC (Multidimensional Metrics for Dynamic Characteristics, Rule Compliance, Risk Appetite and Collaboration) driving style modeling framework provided in an embodiment of the present invention. Figure 2 The framework covers four key dimensions: dynamic characteristics, rule compliance, psychological characteristics, and collaborative methods. Traditional research has mainly focused on dynamic characteristics and rule compliance, such as speed, acceleration, and following distance, but these methods fail to fully reflect the driver's internal decision-making mechanism and interactive behavior in complex traffic scenarios. Therefore, this embodiment adds two dimensions: psychological characteristics and collaborative methods, which are used to reveal the driver's risk preference and decision-making tendency, as well as his willingness to interact with the environment, so as to improve the accurate characterization of driving behavior and the ability to adapt to complex scenarios. Dynamic characteristics describe the vehicle's motion state, rule compliance evaluates the driver's compliance with traffic rules, psychological characteristics reveal the driver's tendency in risk decision-making, and collaborative methods measure the driver's interactive behavior with the environment. By constructing a multidimensional feature space, the framework can fully capture the complexity of driving behavior and refine the behavioral analysis in different driving scenarios. To achieve classification, this embodiment introduces a fuzzy logic method to convert multidimensional features into a classification basis for driving styles. Fuzzy logic is good at handling uncertainties in complex systems. Compared with hard boundary classification methods, it has higher flexibility and robustness, can accurately describe the diversity and transition of driving behaviors, and effectively reduce classification errors caused by fuzzy data division. The framework further supports mapping classification results to simulation platforms, verifying and optimizing driving behavior models, and provides an important reference for the development of intelligent driving technology.

[0070] Step S1 specifically includes:

[0071] S11. Obtain a traffic compliance status indicator, and use a fuzzy subset to define fuzzy values ​​of different traffic compliance status indicators to obtain a traffic compliance status fuzzy indicator.

[0072] Specifically, the vehicle kinematic characteristics are captured through key dynamic indicators such as speed, acceleration, steering angle, yaw rate, lateral position and following distance, and fuzzy subsets are used to deal with the complexity and nonlinearity of variables to obtain the fuzzy indicators of traffic compliance state, as shown in Table 1.

[0073] Table 1 Fuzzy indicators of vehicle dynamics state

[0074] index speed Acceleration Yaw rate Steering angle Fuzzy Value Low / Medium / High Speed Low / Medium / High Acceleration Low / high yaw rate Low / high steering angle

[0075] It can be seen from Table 1 that the vehicle kinematic state indicators include speed, acceleration, steering angle and yaw rate; wherein the fuzzy value of the speed includes low speed, medium speed and high speed, the fuzzy value of the acceleration includes low acceleration, medium acceleration and high acceleration, the fuzzy value of the steering angle includes low steering angle and high steering angle, and the fuzzy value of the yaw rate includes low yaw rate and high yaw rate.

[0076] Furthermore, the speed threshold and the interval range are shown in Table 2, the acceleration threshold and the interval range are shown in Table 3, the yaw rate threshold and the interval range are shown in Table 4, and the steering angle threshold and the interval range are shown in Table 5, wherein the acceleration, steering angle and yaw rate are all absolute values.

[0077] Table 2 Speed ​​threshold and range

[0078] Road Type Low speed (km / h) Medium speed (km / h) High speed (km / h) highway ≤60 60-100 >100 Urban roads ≤30 30-50 >50 Ordinary roads ≤40 40-70 >70

[0079] Table 3 Acceleration threshold and range

[0080] Classification basis <![CDATA[Low acceleration (m / s 2 )]]> <![CDATA[Medium acceleration (m / s 2 )]]> <![CDATA[High acceleration (m / s 2 )]]> Comfort |a|≤1.0 1.0≤|a|<2.5 |a|≥2.5 Rate of change |a|≤0.5 0.5≤|a|<1.0 |a|≥1.0

[0081] Table 4 Yaw rate threshold and range

[0082] Low yaw rate (rad / s) High yaw rate (rad / s) ≤0.2 >0.2

[0083] Table 5 Steering angle threshold and range

[0084] Low steering angle (°) High steering angle (°) ≤10° >10°

[0085] S12. Obtain a traffic compliance status indicator, and use a fuzzy subset to define fuzzy values ​​of different traffic compliance status indicators to obtain a traffic compliance status fuzzy indicator.

[0086] Specifically, we focus on the core indicators of traffic compliance status, such as whether the speed is excessive, whether the following distance is safe, and whether the lane keeping is standardized, and divide them into fuzzy subsets. Through fuzzy logic, we can accurately evaluate the compliance status and enhance the adaptability and robustness of the intelligent transportation system in complex environments. The fuzzy indicators of traffic compliance status are shown in Table 6.

[0087] Table 6 Traffic compliance status fuzzy indicators

[0088] index Is the speed exceeding the limit? Is the following distance safe? Is lane keeping standard? Fuzzy Value Conform / Not Conform Conform / Not Conform Good / Fair / Poor

[0089] As can be seen from Table 2, the traffic compliance status indicators include speeding status, following distance safety status and lane keeping standard status. The fuzzy value of the speeding status includes speeding and not speeding. The fuzzy value of the following distance safety status includes following distance safety and following distance unsafe. The lane keeping standard status includes good lane keeping, average lane keeping and poor lane keeping.

[0090] Furthermore, when judging the speed overspeed state, each frame records the holding state once, and the mode of the speed overspeed situation in the total number of frames is taken as the final state. When the speed is greater than the speed limit by 10%, it is the speeding standard, and the threshold range of the speed overspeed state is shown in Table 7.

[0091] Table 7 Threshold range of speed exceeding

[0092] Road Type Speeding (km / h) highway ≥132 Urban roads (No center line) ≥33; (Single lane) ≥55 Ordinary roads (No center line) ≥44; (Single lane) ≥77

[0093] When judging the safety state of the following distance, the holding state is recorded once per frame, and the mode of the following distance within the total number of frames is taken as the final state. The threshold range of whether the following distance is safe is shown in Table 8.

[0094] Table 8 Threshold range for whether the following distance is safe

[0095] Following distance ≥ 100m: Compliant Following distance <50m: Not compliant

[0096] When judging the lane keeping state, the lane keeping state is recorded once per frame, and the mode of the lane keeping states within the total number of frames is taken as the final state. The threshold range for whether the lane keeping is standard is shown in Table 9.

[0097] Table 9 Threshold range for lane keeping compliance

[0098] Lane keeping status Normative evaluation Deviation from lane center < 10% better 10%-50% deviation from lane center generally Deviation from lane center > 50% bad

[0099] S13. Calculate the risk aversion coefficient, risk decision threshold and risk conversion coefficient and weight them to construct a comprehensive risk preference index.

[0100] In order to reveal the driver's internal decision-making mechanism and risk attitude in more depth, the driver's risk-taking tendency and coping ability are quantified. On this basis, the risk aversion coefficient, decision threshold and risk conversion coefficient are calculated and weighted to construct a comprehensive risk preference index to improve the accuracy of driving style classification. In dynamic driving behavior, the collaborative mode dimension is introduced to consider the interaction between the driver and the external environment, especially the intelligent transportation system. A comprehensive index is constructed according to specific weights based on the collaborative initiative coefficient and the degree of information sharing to evaluate the collaborative efficiency of the driver in the intelligent transportation system. This improves the adaptability and overall efficiency of driving behavior modeling in the intelligent transportation system, especially in multi-vehicle interactions and complex traffic scenarios.

[0101] Specifically, the Risk Avoidance Coefficient (RAC) is used to measure the driver's sensitivity to risk and reflect his risk tolerance. The formula for calculating the risk avoidance coefficient is:

[0102]

[0103] Among them, a x is the longitudinal acceleration, a y is the lateral acceleration, δ is the steering angle, v is the speed; α1 and α2 are weight coefficients, which need to be determined by fitting experimental data.

[0104] The Risk Decision Threshold (RDT) defines the critical point at which the driver takes risk-averse behavior, reflecting his decision-making acumen. The formula for calculating the Risk Decision Threshold is:

[0105]

[0106] Among them, DHW, THW, and TTC are respectively the following distance (reflecting the spatial distance between the vehicle and the front vehicle), the headway time (describing the time required to reach the position of the front vehicle at the current speed), and the collision time (indicating the time when a collision may occur in the current state), and β1, β2, and β3 are all weight coefficients.

[0107] The Risk Transition Coefficient (RTC) reflects the driver's adjustment rate of risk preference in different situations. The formula for calculating the Risk Transition Coefficient is:

[0108]

[0109] Among them, Δa x ,Δa y ,Δv,Δσ, They are the longitudinal acceleration change, lateral acceleration change, speed change, steering angle change and yaw rate change respectively.

[0110] The comprehensive risk appetite indicator is:

[0111] R=w1×R a +w2×R d +w3×R RTC

[0112] Among them, w1, w2, w3 are weights.

[0113] S14. Calculate the collaboration initiative coefficient and information sharing degree and weight them to construct a comprehensive index of collaboration mode.

[0114] Specifically, the Collaboration Proactivity Coefficient (CPC) measures the driver's willingness to actively collaborate in traffic situations. The calculation formula for the Collaboration Proactivity Coefficient is:

[0115]

[0116] Among them, y,d c They are the lateral position (indicating the degree of deviation of the vehicle in the lane) and the lateral distance to the adjacent vehicle, and γ1 and γ2 are weight coefficients.

[0117] Information Sharing Degree (ISD) evaluates the driver's communication ability in the intelligent transportation system.

[0118] The calculation formula for the degree of information sharing is:

[0119] I s =ω1log(v)+ω2ρ

[0120] Among them, v is the speed, ρ is the traffic flow density, ω1, ω2 are weight coefficients;

[0121] The comprehensive indicators of the collaboration mode are:

[0122] CI=μ1×C i +μ2×I s

[0123] Among them, μ1, μ2 are weight coefficients.

[0124] Furthermore, the comprehensive index of risk preference and the comprehensive index of collaboration mode are values ​​of 0-1, and their threshold ranges are shown in Table 10.

[0125] Table 10 Threshold ranges of comprehensive risk preference indicators and collaboration methods

[0126]

[0127] S15. Construct a driving style fuzzy rule according to the vehicle dynamics state fuzzy index, the traffic compliance state fuzzy index, the risk preference comprehensive index and the cooperation mode comprehensive index.

[0128] After completing the fuzzy processing of the indicators and setting the rules, we construct a series of fuzzy rules for the given input parameters based on expert experience. First, operate on the actual data set by extracting relevant data (such as speed, acceleration, following distance, etc.). Calculate the fuzzy membership and generate the driving style classification results according to the preset fuzzy rules. For each data, the system automatically determines its driving style category based on the rule mapping table, as shown in Table 11, and statistically classifies the results to verify the overall distribution rationality and classification accuracy of the data set. In Table 11, the risk preference and collaboration mode values ​​will eventually be normalized to values ​​between 0 and 1. Low means below 0.3, and extremely low means below 0.15; high means above 0.8, and extremely high means above 0.9; lane keeping specifications are more than 80% of the deviations from the lane center ≥50% during the entire simulation test cycle, which is frequent and large deviations.

[0129] Table 11 Driving style fuzzy rules

[0130]

[0131]

[0132] As shown in Table 10, the driving style fuzzy rules include: very defensive, high priority defensive, low priority defensive, high priority normal movement, low priority normal movement, very sporty, high priority aggressive, and low priority aggressive.

[0133] Furthermore, the classification results are mapped to the simulation platform (such as CARLA, SUMO). Specifically, the dynamic parameters of the simulated vehicle are set according to the classification results of different driving styles. The simulation platform generates the behavior pattern of the corresponding driving style by inputting these dynamic parameters. During the simulation process, the classification model is verified and optimized using actual driving behavior data, and the dynamic parameters are adjusted to make the simulation results more consistent with the actual driving behavior characteristics, ultimately achieving closed-loop optimization of the classification model and simulation system, and providing support for the development of intelligent transportation systems. Then, these parameters are used to generate different styles of driving behavior in the simulation system. Finally, a system framework that can dynamically adjust the driving style is constructed to achieve a smooth transition from defensive to aggressive, providing support for subsequent testing and optimization.

[0134] Through the above steps, this embodiment achieves efficient classification and simulation mapping of multi-dimensional driving styles on the basis of comprehensively constructing a driving behavior feature system, while effectively reducing the complexity of data processing and the risk of insufficient model adaptability. The constructed fuzzy logic classification and dynamic simulation framework can not only accurately reflect the behavioral characteristics of different driving styles, but also has the advantages of maintaining high robustness and scalability in complex traffic scenarios. This system provides a solid technical foundation and reliable data support for further exploration of driving scene generation and simulation based on reinforcement learning, creating more possibilities for the optimized design of intelligent driving systems.

[0135] S2. Constructing a driving environment and configuring quantitative features of different driving styles according to the driving style fuzzy rules.

[0136] For example, a diverse driving environment is constructed in the SUMO simulation platform, including complex road conditions (such as multiple lanes, highways, intersections), dynamic traffic flow (supporting changes in different traffic density) and multi-type vehicle models (such as cars, buses and trucks). Then, by configuring SUMO XML files (such as route.xml, network.xml), the quantitative characteristics of different driving styles are clarified (such as the rapid acceleration and braking behavior of aggressive driving, and the low-speed and uniform driving characteristics of conservative driving). At the same time, multi-dimensional data is collected, including vehicle status (speed, acceleration), driving actions (acceleration, steering angle) and environmental characteristics (traffic flow density) to provide support for the initial input of the model.

[0137] S3. Using a reinforcement learning algorithm, based on the quantitative characteristics of the driving environment and different driving styles, taking the driving style fuzzy rules as the target, calculating the rewards of different driving scenarios according to the reward function to train and update the intelligent agent policy network, and obtaining a trained intelligent agent policy network; the trained intelligent agent policy network is used to generate multi-dimensional natural driving scenarios and perform simulation tests.

[0138] In response to the challenges in real-time traffic management and adaptability to complex environments, this embodiment introduces the Transformer model into the optimization of multi-agent communication mechanisms. Compared with the inefficiency of traditional sequence processing methods (such as RNN or LSTM) when processing long sequences, Transformer can process the entire sequence at the same time, greatly improving computational efficiency. With the help of the self-attention mechanism, the agent can accurately filter out the information most relevant to the current decision. In order to generate scenarios that conform to multi-dimensional driving styles, this embodiment constructs the input and output structure of reinforcement learning, defines the state and action space, and evaluates the quality of driving behavior in combination with the reward function. Subsequently, the initial scenario is generated using agent strategy training, and the reward weights are dynamically optimized to ensure the diversity of the generated scenarios and driving style characteristics. Finally, a closed-loop feedback mechanism is established to analyze and optimize the failed scenarios, gradually improve the adaptability and comprehensiveness of scenario generation, and provide support for intelligent driving system testing. The overall process is as follows: Figure 3 As shown, Figure 3 The driving style scenario generation and simulation process of reinforcement learning provided by an embodiment of the present invention.

[0139] Specifically, the input and output structure of reinforcement learning is constructed, including the state space, action space, and reward function. First, the state space is used to describe the real-time status of the vehicle and the environment, such as environmental variables such as vehicle speed, acceleration, and distance to the vehicle in front. Then, the action space defines the control input of the driving behavior, such as speed, acceleration, steering wheel steering angle, etc. Next, the reward function is designed to evaluate the driving behavior and scene quality by combining short-term rewards and long-term rewards. Short-term rewards are used to provide immediate feedback on the compliance and safety of driving behavior, such as obstacle avoidance behavior, rule compliance, etc.; long-term rewards evaluate the overall rationality and complexity of the scene, such as whether the driving behavior generated by multi-dimensional features meets the expected goals. Finally, the input-output structure and reward function are integrated into the reinforcement learning framework to lay the foundation for scene generation.

[0140] Specifically, the input of the state space includes real-time state information, historical state trajectory and environmental information. Real-time state information covers vehicle dynamics state (speed, acceleration, yaw rate, steering angle), traffic compliance state (whether speeding, following distance, lane keeping state), which reflects the vehicle's current movement and traffic rules compliance. The historical state trajectory is the state of the previous time step as an additional feature input, which provides the agent with information in the time dimension, which helps to analyze the state change trend and make more reasonable decisions. Environmental information includes road type, traffic flow, etc. Environmental factors have an important impact on vehicle behavior decisions. This information expands the state space and enables the agent to respond differently according to different environments. The role of the state space is to integrate the above-mentioned information and provide the agent with a complete description of the current environmental state so that the agent can choose appropriate actions from the action space. Therefore, the state space has no direct output.

[0141] Specifically, the input of the action space is the decision made by the agent based on the information provided by the state space. In the decision-making process, various dimensions of the state (such as vehicle dynamics state, traffic compliance state, historical trajectory, environmental information, etc.) are considered, and executable actions are determined through a certain strategy network or decision logic. These state information is the input source of the action space. The output of the action space includes continuous output and discrete output; continuous output includes acceleration and steering angle, these continuous control variables are used to fine-tune the vehicle's motion state; discrete output includes whether to change lanes, slow down or accelerate, etc., in the form of discrete decision instructions, so that the vehicle can make clear action choices in different traffic scenarios.

[0142] The logical relationship between state space and action space is:

[0143] State-driven action selection: The state space contains the agent's observation information of the environment, reflecting the current state of the environment. The agent decides what action to take based on the information in the state space, that is, the state is the basis for action selection. For example, in an autonomous driving scenario, the vehicle's current speed, distance from the vehicle in front, lane position and other state information will prompt the agent to select appropriate actions from the action space (such as acceleration, deceleration, steering and other action sets). Action affects state transition: After the agent selects and executes an action from the action space, this action will act on the environment and change the state of the environment. For example, in a robot walking task, after the robot executes the action of "stepping forward", its position, posture and other state information in the environment will change accordingly, thus entering a new state. The new state will become the basis for the next action selection, and so on.

[0144] Furthermore, random actions are used to generate initial scenarios and train the reinforcement learning agent policy network. First, the agent generates different driving scenarios through randomly initialized strategies, such as simulating the probability of pedestrians appearing on urban roads, vehicle spacing, and traffic flow density. Then, the score of each scenario is calculated based on the reward function to evaluate whether the scenario conforms to the driving style classification and behavior logic. Next, the agent policy network is updated using a deep reinforcement learning algorithm to gradually generate high-quality scenarios, such as generating aggressive driving style overtaking scenarios on highways or defensive style avoidance behavior scenarios at urban intersections. Finally, through multiple rounds of training, the scenario generation strategy converges to the optimal solution to ensure the diversity and test coverage of the generated driving scenarios.

[0145] In this embodiment, a fusion network is designed in combination with the Transformer framework, that is, the Transformer encoder is used to extract the environment and vehicle state features, the Actor network generates a driving strategy based on the features, and the Critic network evaluates the value of the action. The Actor network and the Critic network form the agent strategy network. Furthermore, this embodiment uses the MAPPO reinforcement learning algorithm to update the Actor network and the Critic network.

[0146] The process of training and updating the agent policy network includes the following steps:

[0147] S1. Network parameter initialization.

[0148] Specifically, the learning rate is initially set to 1e -4 , the initial discount factor γ = 0.99, and the time difference error (TD error) is used to optimize the Critic network.

[0149] S2, feature extraction and training process of environmental state sequence. Specifically including:

[0150] S21, obtaining the real-time status information, historical status trajectory and environmental information of the vehicle, and encoding the real-time status information, the historical status trajectory and the environmental information using a Transformer encoder to obtain the vehicle status at the current moment.

[0151] Specifically, the real-time state information includes: vehicle speed, acceleration, yaw rate, steering angle, speeding state, following distance state and lane keeping state; the historical state trajectory includes the state of the previous time step; the environmental information includes road type and traffic flow. Exemplarily, the vehicle speed, acceleration, steering angle and path deviation are obtained from the SUMO simulation and encoded as time series data. At the same time, the Transformer encoder converts the input state feature vector into a high-dimensional representation.

[0152] S22. Based on the vehicle state at the current moment, generate a driving behavior action using an Actor network, execute the driving behavior action, obtain the vehicle state at the next moment, and calculate the reward of the current driving scenario using a multi-objective reward function.

[0153] Specifically, the driving behavior action includes continuous output and discrete output, the continuous output includes acceleration and steering angle, and the discrete output includes whether to change lanes, decelerate or accelerate. Exemplarily, based on the output characteristics of the encoder, a driving behavior action (such as acceleration, deceleration, lane change) is generated, and the action is fed back to the SUMO simulation for execution, the simulation environment returns a new state, and a multi-objective reward function is used to calculate the reward of the current driving scene.

[0154] The multi-objective reward function is:

[0155] R=p1·R1+p2·R2+p3·R3+p4·R4+p5·R5

[0156] Among them, R1 is the comprehensive dynamics reward, R2 is the comprehensive traffic compliance reward, R3 is the risk preference reward, R4 is the collaborative mode reward, R5 is the driving style matching reward, and p1, p2, p3, p4, and p5 are weight coefficients;

[0157] R1=w speed ·R speed +w acc ·R acc +w yaw ·R Ryaw +w steering ·R steering

[0158]

[0159]

[0160] Among them, R speed is the speed reward, R acc is the acceleration reward, R Ryaw is the yaw rate reward, R steering is the steering angle reward, w speed 、w acc 、w yaw 、w steering is the weight coefficient;

[0161] R2=w overspeed ·R overspeed +w distance ·R distance +w lane ·R lane

[0162]

[0163] Among them, R overspeed For speeding reward, R distance is the following distance reward, R lane lane keeping reward; w overspeed 、w distance 、w lane is the weight coefficient;

[0164]

[0165] It can be understood that when the risk preference and cooperation mode within the current driving style are in the range of low, medium and high, the reward is increased by 0.01, otherwise it is increased by -0.02.

[0166] S23. Based on the time difference method, the Critic network is used to evaluate the vehicle state and the driving behavior at the current moment; wherein the Actor network and the Critic network form the intelligent agent strategy network.

[0167] For example, the Q value or V value is evaluated for the current state and action, and the long-term benefit is predicted in combination with the next state characteristics. The TD error is then used as the loss function of the Critic network to update the network weights so that its evaluation gradually approaches the real environment feedback value.

[0168] The loss function based on TD error is:

[0169]

[0170] Among them, R t is the reward at time t, γ is the discount factor, V(S t ; θ) is the state S t The value of is estimated, and θ is the parameter of the Critic network.

[0171] S25. Use the MAPPO reinforcement learning algorithm to update the Actor network and the Critic network, and stop training when a preset number of iterations is reached.

[0172] S25. Adjust the weights in the multi-objective reward function to optimize the agent strategy network.

[0173] Specifically, for the first time, you can set p1, p2, p3, p4, and p5 to 0.2, and w speed 、w acc 、w yaw 、w steering Both are 0.25, w overspeed 、w distance 、wlane The weight coefficients are all 0.33. According to the training results of S21 to S25, different weight coefficients are adjusted to repeat the training until the agent strategy network converges to the optimal solution that meets the fuzzy rules of driving style, and a trained agent strategy network is obtained. The trained agent strategy network can be used to generate multi-dimensional natural driving scenes and perform simulation tests.

[0174] Furthermore, after obtaining the trained agent strategy network, a closed-loop feedback mechanism can be established to analyze the adaptability of the scenario and optimize the scenario generation logic. First, for the scenarios that failed in the simulation test, such as the driving style that cannot match the expected goal or the unreasonable behavior, a systematic analysis is conducted to explore the reasons for their failure. Then, the scenario generation logic is readjusted based on the analysis results, such as optimizing the parameters of the reinforcement learning model or adjusting the weight distribution of the reward function. Next, the optimized scenario is input into the Carla simulation system again for testing to ensure that it can cover more boundary conditions and test scenarios. Finally, the generated scenario parameters are gradually converged through multiple iterations, making the driving style scenario generation more comprehensive and efficient. After that, the two-way migration and evaluation closed loop of simulation data and real vehicle data are carried out. Specifically, the system migrates the natural traffic environment and dynamic characteristics in the real vehicle test to the simulation platform to enhance the simulation authenticity; at the same time, the simulation results are used to optimize the real vehicle test plan, give priority to testing key boundary scenarios, and finally form a closed-loop feedback mechanism between data and model, thereby effectively improving the robustness and comprehensiveness of system performance.

[0175] This embodiment introduces risk preference and naturalness indicators of collaboration mode, and constructs fuzzy rules of driving style based on dynamic characteristics, rule compliance, psychological characteristics and collaboration mode. Then, the reward function in the reinforcement learning framework is used to train and update the agent strategy network to achieve dynamic generation of the driving behavior model. This method can generate a variety of behavior patterns covering from cautious to aggressive driving, obtain a more complete and accurate representation of behavioral characteristics, and thus generate a variety of driving behaviors, providing scientific and efficient technical support for the research and development, verification, future technical standardization, safety assessment and large-scale application of autonomous driving systems.

[0176] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A multi-dimensional natural driving scene generation and simulation test method based on reinforcement learning, characterized in that: Includes steps: S1. Based on dynamic characteristics, rule compliance, psychological characteristics and cooperation mode, the driving behavior characteristics are quantified by combining the fuzzy logic classification model to construct the driving style fuzzy rules; S2, constructing a driving environment and configuring quantitative features of different driving styles according to the driving style fuzzy rules; S3. Using a reinforcement learning algorithm, based on the quantitative characteristics of the driving environment and different driving styles, taking the driving style fuzzy rules as the target, calculating the rewards of different driving scenarios according to the reward function to train and update the intelligent agent policy network, and obtaining a trained intelligent agent policy network; the trained intelligent agent policy network is used to generate multi-dimensional natural driving scenarios and perform simulation tests.

2. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 1, characterized in that: Step S1 includes: S11, obtaining a vehicle kinematic state index, and using a fuzzy subset to define fuzzy values ​​of different vehicle kinematic state indexes to obtain a vehicle dynamic state fuzzy index; S12, obtaining a traffic compliance state index, and using a fuzzy subset to define fuzzy values ​​of different traffic compliance state indexes to obtain a traffic compliance state fuzzy index; S13. Calculate the risk aversion coefficient, risk decision threshold and risk conversion coefficient and weight them to construct a comprehensive risk preference index; S14. Calculate the collaboration initiative coefficient and information sharing degree and weight them to construct a comprehensive index of collaboration mode; S15. Construct a driving style fuzzy rule according to the vehicle dynamics state fuzzy index, the traffic compliance state fuzzy index, the risk preference comprehensive index and the cooperation mode comprehensive index.

3. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 2, characterized in that: The vehicle kinematic state index includes speed, acceleration, steering angle and yaw rate; wherein the fuzzy value of the speed includes low speed, medium speed and high speed, the fuzzy value of the acceleration includes low acceleration, medium acceleration and high acceleration, the fuzzy value of the steering angle includes low steering angle and high steering angle, and the fuzzy value of the yaw rate includes low yaw rate and high yaw rate; The traffic compliance status indicators include speeding status, following distance safety status and lane keeping standard status. The fuzzy value of the speeding status includes speeding and not speeding. The fuzzy value of the following distance safety status includes following distance safety and following distance unsafe. The lane keeping standard status includes good lane keeping, average lane keeping and poor lane keeping.

4. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 2, characterized in that: The calculation formula of the risk aversion coefficient is: Among them, a x is the longitudinal acceleration, a y is the lateral acceleration, δ is the steering angle, v is the speed; α1 and α2 are weight coefficients; The calculation formula of the risk decision threshold is: Among them, DHW, THW, TTC are following distance, headway time, and collision time, respectively, and β1, β2, and β3 are weight coefficients; The calculation formula of the risk conversion coefficient is: Among them, Δa x ,Δa y ,Δv,Δσ, They are the longitudinal acceleration change, lateral acceleration change, speed change, steering angle change and yaw rate change respectively; The comprehensive risk preference indicator is: R=w1×R a +w2×R d +w3×R RTC Among them, w1, w2, w3 are weights.

5. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 2, characterized in that: The calculation formula of the collaboration initiative coefficient is: Among them, y,d c are the lateral position and the lateral distance to the adjacent vehicle, respectively, γ1 and γ2 are weight coefficients; The calculation formula of the information sharing degree is: I s =ω1log(v)+ω2ρ Among them, v is the speed, ρ is the traffic flow density, ω1, ω2 are weight coefficients; The comprehensive indicators of the collaboration mode are: CI=μ1×C i +μ2×I s Among them, μ1, μ2 are weight coefficients.

6. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 2, characterized in that: The driving style fuzzy rules include: very defensive, high priority defensive, low priority defensive, high priority normal movement, low priority normal movement, very sporty, high priority aggressive and low priority aggressive.

7. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 1, characterized in that: Step S2 includes: A driving environment is constructed in a simulation platform and quantitative features of different driving styles are configured according to the driving style fuzzy rules, wherein the driving environment includes complex road conditions, dynamic traffic flow and multi-type vehicle models.

8. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 1, characterized in that: Step S3 includes: Acquire the real-time state information, historical state trajectory and environmental information of the vehicle, and use a Transformer encoder to encode the real-time state information, the historical state trajectory and the environmental information to obtain the vehicle state at the current moment; Based on the vehicle state at the current moment, a driving behavior action is generated using an Actor network, the driving behavior action is executed, the vehicle state at the next moment is obtained, and a reward for the current driving scenario is calculated using a multi-objective reward function; Based on the time difference method, the Critic network is used to evaluate the vehicle state and the driving behavior at the current moment; wherein the Actor network and the Critic network form the agent strategy network; Using the MAPPO reinforcement learning algorithm to update the Actor network and the Critic network; The weights in the multi-objective reward function are adjusted to optimize the agent policy network.

9. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 8, characterized in that: The real-time status information includes: vehicle speed, acceleration, yaw rate, steering angle, speeding status, following distance status and lane keeping status; the historical status trajectory includes the status of the previous time step; the environmental information includes road type and traffic flow; The driving behavior action includes continuous output and discrete output. The continuous output includes acceleration and steering angle, and the discrete output includes whether to change lanes, slow down or accelerate.

10. The method for generating and simulating multi-dimensional natural driving scenarios based on reinforcement learning according to claim 8, characterized in that: The multi-objective reward function is: R=p1·R1+p2·R2+p3·R3+p4·R4+p5·R5 Among them, R1 is the comprehensive dynamics reward, R2 is the comprehensive traffic compliance reward, R3 is the risk preference reward, R4 is the collaborative mode reward, R5 is the driving style matching reward, and p1, p2, p3, p4, and p5 are weight coefficients; R1=w speed ·R speed +w acc ·R acc +w yaw ·R Ryaw +w steering ·R steering Among them, R speed is the speed reward, R acc is the acceleration reward, R Ryaw is the yaw rate reward, R steering is the steering angle reward, w speed 、w acc 、w yaw 、w steering is the weight coefficient; R2=in overspeed ·R overspeed +in distance ·R distance +in lane ·R lane Among them, R overspeed For speeding reward, R distance is the following distance reward, R lane lane keeping reward; w overspeed 、w distance 、w lane is the weight coefficient;

Citation Information

Patent Citations

  • Adaptive cruise control method and system based on deep reinforcement learning

    CN116252791A

  • Natural automatic driving scene generation method and device based on reinforcement learning

    CN118228612A

Cited By

  • Emergency scene driving behavior data processing method based on deep learning

    CN121302215A