Electric bicycle intrusion behavior simulation method based on reinforcement learning parameter self-adaption
By deeply integrating reinforcement learning with social force models, and dynamically adjusting parameters and decision probabilities, the problems of subjective parameter calibration and rigid behavior in the simulation of electric bicycle intrusion behavior are solved, achieving simulation results with higher accuracy and generalization ability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing simulation methods for electric bicycle intrusion behavior suffer from problems such as strong subjectivity in parameter calibration, insufficient realism of behavior, and weak generalization ability in intersection traffic flow, making it difficult to adapt to dynamic traffic environments.
By deeply integrating reinforcement learning and social force models, and dynamically adjusting parameters through reinforcement learning agents, combined with multi-objective reward function optimization, accurate simulation of electric bicycle intrusion behavior can be achieved.
It achieves adaptive parameter optimization, refines the decision-making process, enhances the model's generalization ability, can adapt to changing traffic environments, and improves simulation accuracy and practicality.
Smart Images

Figure CN122046901A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent transportation simulation technology, specifically relating to a simulation method for electric bicycle intrusion behavior based on reinforcement learning parameter adaptation. Background Technology
[0002] Simulation of non-motorized traffic flow, especially the accurate reproduction of electric bicycle behavior, is a key issue in the field of intelligent transportation. Existing research scenarios mainly focus on road sections with separation facilities for motorized and non-motorized traffic, where model parameter calibration is relatively simple. However, intersections, as key nodes in urban road networks, exhibit significantly increased traffic flow complexity: on the one hand, intersections involve dynamic factors such as multi-directional traffic weaving, traffic light control, and mixed pedestrian and vehicle traffic, placing higher demands on the real-time performance and interactive realism of simulation models; on the other hand, the frequent intrusive behavior of electric bicycles at intersections is difficult to accurately describe using fixed rules or homogeneous data-driven methods, leading to significant deviations between simulation results and measured data.
[0003] Existing simulation methods are mainly divided into two categories: rule-based models and data-driven models, each with its own advantages and limitations.
[0004] Rule-based models, represented by social force models, continuously describe vehicle trajectories through mechanical equations, providing physical interpretability. However, parameter calibration relies on human experience, leading to strong subjectivity. For example, the intensity and range of repulsive force in social force models require manual adjustment; even small changes can cause significant trajectory deviations. Furthermore, their decision-making mechanisms are often preset to fixed values, failing to adapt to dynamic environments and resulting in insufficient behavioral diversity. Data-driven models, such as generative adversarial learning, learn behavioral strategies from real data, mitigating the subjectivity problem of rule-based models. However, they may homogenize vehicle behavior, making it difficult to characterize the proactive intrusion characteristics of electric bicycles in heterogeneous flows.
[0005] Although rule-based models and data-driven models have made significant progress in non-motorized traffic flow simulation, they still have the following key shortcomings that restrict the realism and practicality of the simulation: (1) the parameter calibration of rule-based models is highly subjective; (2) the behavior of data-driven models is homogeneous; and (3) dynamic decision-making is rarely applied. Therefore, the integration of reinforcement learning and simulation models has become a research focus. Its progress includes parameter adaptive optimization, innovation of dynamic decision-making mechanisms, and reward function design, which can effectively improve the simulation accuracy and generalization ability.
[0006] The main disadvantages of existing technologies are as follows: (1) High parameter sensitivity: Small changes in hyperparameters of the social force model (such as the range of repulsive force) can lead to significant deviations in simulation results, while manual calibration is inefficient and highly subjective.
[0007] (2) Insufficient realism of behavior: Fixed decision probability cannot simulate the dynamic decision-making of riders based on the environment (such as the distance between vehicles in front and behind, and the status of traffic lights), resulting in a large error between overtaking and space intrusion behaviors and the measured data.
[0008] (3) Weak generalization ability: After the model is calibrated in a specific scenario (such as a one-way road), it is difficult to transfer to complex scenarios (such as irregular intersections). Summary of the Invention
[0009] The technical problem to be solved by this invention is to provide a simulation method for electric bicycle intrusion behavior based on reinforcement learning parameter adaptation, which is applicable to simulation of non-motorized mixed traffic flow, construction of autonomous driving test platforms, and optimization of traffic management strategies.
[0010] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A simulation method for electric bicycle intrusion behavior based on reinforcement learning parameter adaptation includes the following steps: The environmental parameters are acquired and the reinforcement learning agent is initialized. The environmental parameters include vehicle parameters, traffic flow parameters, and road parameters. The action space of the reinforcement learning agent includes the hyperparameter adjustment amount of the social force model. The reinforcement learning agent is trained iteratively until the convergence condition is met to obtain the simulation model. In each iteration, the reinforcement learning agent acquires all vehicle states and updates the state space, then calculates the corresponding action space, and calculates the electric bicycle intrusion decision probability based on the state space. The social force model adjusts the hyperparameters according to the hyperparameter adjustment amount in the action space, then calculates the social force of the electric bicycle based on the electric bicycle intrusion decision probability, and finally updates the vehicle state of the electric bicycle based on the social force calculation result. The simulation model was used to simulate the intrusion behavior of electric bicycles.
[0011] Furthermore, the state space of the reinforcement learning agent includes vehicle self-state features and interaction environment features. The vehicle self-state features include vehicle position coordinates, velocity vector, and acceleration. The interaction environment features include distance between the vehicle and the target point, distance between the vehicle and the vehicle in front, distance between the vehicle and the vehicle behind, relative speed, traffic light status, and traffic density.
[0012] Furthermore, when calculating the intrusion decision probability of an electric bicycle based on the state space, the mathematical expression for the intrusion decision probability of an electric bicycle is as follows:
[0013] in, The distance between the vehicle and the vehicle in front. For relative velocity, Traffic light status. , , , All of these are weight parameters.
[0014] Furthermore, the hyperparameters of the social force model include the decision threshold. When calculating the social force of an electric bicycle based on the probability of its intrusion decision, the following parameters are included: If the probability of an electric bicycle intruding is greater than the decision threshold in the adjusted hyperparameters, then calculate the self-driving force, repulsive force, boundary force, and intrusion force experienced by the electric bicycle. If the decision probability of the electric bicycle intrusion is less than the decision threshold in the adjusted hyperparameters, then the self-driving force, repulsive force, and boundary force on the electric bicycle are calculated, and the intrusion force is set to 0.
[0015] Furthermore, the hyperparameters of the social force model also include repulsive force strength, influence range, boundary force coefficient, and attenuation distance. The calculation formula for the self-driving force is as follows:
[0016] in, For vehicle quality, For relaxation time, For the desired speed, For electric bicycles at the current moment The unit vector of the desired direction of motion. For electric bicycles at the current moment velocity vector; The formula for calculating the repulsive force is as follows:
[0017] in, The strength of the repulsive force. For vehicles With vehicles The distance between them For the scope of influence, By vehicle Pointing to the vehicle , unit vector; The formula for calculating the boundary force is as follows:
[0018] in, These are the boundary force coefficients. For vehicles With boundary distance, For attenuation distance, For vehicles A unit vector perpendicular to the lane's inner edge from the near-side boundary; The formula for calculating the intrusion force is as follows:
[0019] in, The force of intrusion is the coefficient. For electric bicycles With vehicles The distance between them The range of influence of the intrusion force. The angle between the direction of intrusion and the current direction of movement of the electric bicycle. This is a dynamically adjusted factor.
[0020] Furthermore, when updating the vehicle state of the electric bicycle based on the social force calculation results, the resultant force of the self-driving force, repulsive force, boundary force, and intrusion force is calculated, and then the vehicle state of the electric bicycle is updated using Newton's laws of motion.
[0021] Furthermore, during iterative training of the reinforcement learning agent, a proximal policy optimization algorithm is employed, and the reward function design uses a multi-objective reward function to guide the optimization direction. The mathematical expression of the reward function of the reinforcement learning agent is as follows:
[0022] in: It is a reward for trajectory error. , This refers to the total number of points or time steps used for comparison in the trajectory. This refers to the trajectory data of an electric bicycle generated through a simulation model. This data represents actual, observed electric bicycle trajectory data. It is a penalty for violating the safe distance rule. , For indicator functions, This represents the actual distance between the electric bicycle and other vehicles or obstacles in the simulation. It's an efficiency reward. , This represents the total number of time steps in the simulation or evaluation process. For electric bicycles The instantaneous velocity scalar value at a certain time step.
[0023] Furthermore, during iterative training of the reinforcement learning agent, the training parameters are adaptively optimized using different traffic flow densities. Training is conducted at low density, medium-high density, and mixed density. The low density iteration count is 1-1500, with the optimization focus on the basic social force parameters; the medium-high density iteration count is 1501-3000, with the optimization focus on the decision threshold; and the mixed density iteration coefficient is 3001-5000, with the optimization focus on generalization ability.
[0024] Furthermore, the convergence condition is that the number of consecutive iterations in which the reward fluctuation is less than a threshold reaches a specified value or the maximum number of iterations is reached.
[0025] Furthermore, the threshold is 0.1 and the specified value is 100.
[0026] Compared with the prior art, the advantages of the present invention are as follows: (1) Achieving adaptive parameter optimization: This invention dynamically adjusts parameters through reinforcement learning agents, using the state space (such as vehicle position and speed difference) as input and the action space as the parameter adjustment amount, and guides optimization through reward functions (such as the negative value of trajectory error), replacing manual debugging that relies on expert experience.
[0027] (2) Refined decision-making process: The present invention changes the decision probability to be generated in real time by a reinforcement learning agent. The agent calculates dynamic probability based on multi-dimensional states (such as the distance between vehicles, traffic light cycle, and speed difference) to simulate the random decision-making behavior of cyclists.
[0028] (3) Enhancing model generalization: This invention designs a multi-objective reward function to guide reinforcement learning agents to explore diverse scenarios. The reward function integrates trajectory error, safe distance constraints, and efficiency indicators, enabling the model to adapt to changing traffic environments. Attached Figure Description
[0029] Figure 1 This is a simplified flowchart of an embodiment of the present invention.
[0030] Figure 2 This is the overall design framework for embodiments of the present invention.
[0031] Figure 3 This is the system initialization process according to an embodiment of the present invention.
[0032] Figure 4 This is the dynamic decision-making process in an embodiment of the present invention. Detailed Implementation
[0033] The present invention will be further described below with reference to the accompanying drawings and specific preferred embodiments, but this does not limit the scope of protection of the present invention.
[0034] This invention aims to solve the core technical problem of inaccurate modeling of electric bicycles encroaching on motor vehicle driving space in intersection traffic flow simulation. Specifically, existing simulation methods suffer from the following technical bottlenecks: (1) Parameter calibration depends on human experience: Hyperparameters in the model (such as force strength parameters and decision probability parameters) need to be experimentally fitted or calibrated according to different scenarios, which makes it difficult to adapt to dynamic traffic environments, resulting in low simulation accuracy and poor generalization ability.
[0035] (2) Rigid decision-making mechanism: The decision probability of electric bicycle intrusion behavior is often preset to a fixed value, which cannot reflect the dynamic decision-making process of riders based on the real-time environment, resulting in homogenization of behavior.
[0036] (3) Insufficient quantification of uncertainty: The model does not consider parameter uncertainty and behavioral randomness, resulting in serious error accumulation in complex interaction scenarios.
[0037] To address the aforementioned issues, this implementation proposes a simulation method for electric bicycle intrusion behavior based on reinforcement learning parameter adaptation. By deeply integrating reinforcement learning with a social force model, a complete closed-loop optimization framework of "simulation-evaluation-learning" is constructed to achieve accurate simulation of electric bicycle intrusion behavior.
[0038] like Figure 1 As shown, it includes the following steps: S1) Obtain environmental parameters and initialize the reinforcement learning agent, iteratively train the reinforcement learning agent until the convergence condition is met, and obtain the simulation model: In this embodiment, the action space of the reinforcement learning agent includes the hyperparameter adjustment amount of the social force model. In each iteration of the reinforcement learning agent training, the reinforcement learning agent acquires all vehicle states and updates the state space, then calculates the corresponding action space, and simultaneously calculates the electric bicycle intrusion decision probability based on the state space. The social force model adjusts the hyperparameters according to the hyperparameter adjustment amount in the action space, then calculates the social force of the electric bicycle based on the electric bicycle intrusion decision probability, and finally updates the vehicle state of the electric bicycle based on the social force calculation result.
[0039] S2) Use the simulation model to simulate the intrusion behavior of electric bicycles.
[0040] like Figure 2 As shown, step S1 in this embodiment includes five key parts: system initialization, social force calculation and dynamic decision-making, reinforcement learning and parameter optimization, resultant force calculation and state update, and termination and output. This framework achieves parameter adaptation, dynamic decision-making, and multi-scenario adaptation through iterative optimization. The key parts are described below.
[0041] This embodiment loads environmental parameters such as road network and traffic rules during the system initialization phase and initializes the reinforcement learning agent to achieve automatic calibration of the hyperparameters of the social force model, replacing traditional manual debugging. Figure 3 As shown, it specifically includes: S101) Environmental parameter settings: including vehicle parameters, environmental parameters, and traffic flow parameters; S102) Reinforcement Learning Initialization: Reinforcement Learning Agent Initialization of State Space and action space .in: The state space includes the vehicle position. ,speed acceleration Distance between vehicles in front and behind These multi-dimensional features are divided into vehicle self-state features and interaction environment features. The vehicle self-state features specifically include: Vehicle position coordinates The vehicle's lateral and longitudinal positions within the current intersection coordinate system, with an accuracy of 0.1 meters; velocity vector Instantaneous speed of a vehicle in the lateral and longitudinal directions, measured in meters per second, reflecting its real-time motion status; acceleration The vehicle's current acceleration value is used to predict short-term motion trends. The specific features of the interactive environment include: Distance from target point The straight-line distance between the vehicle and the desired exit direction; Distance to the vehicle in front The instantaneous distance to the vehicle in front in the same direction affects the decision to follow or overtake. Distance to the car behind Maintain proper distance from vehicles traveling in the same direction to avoid rear-end collisions; relative speed The speed difference between this vehicle and the vehicle in front quantifies the urgency of the interaction. Traffic light status : Discrete variable (0 / 1 / 2), representing red / green / yellow light status; Traffic density The number of vehicles / area within the sensing range reflects the level of congestion in the scene.
[0042] The action space represents the output of the reinforcement learning agent, i.e., the amount of adjustment to the hyperparameters of the social force model. in Including repulsive force strength Scope of influence Boundary force coefficient Attenuation distance and decision threshold This embodiment designs the action space as a continuous value vector to achieve refined parameter adaptation, mainly including: Repulsive force intensity adjustment amount :scope N, optimize the collision avoidance intensity between vehicles; Adjustment of the scope of influence :scope m controls the distance at which the repulsive force acts; Decision threshold adjustment amount :scope Dynamically change the probability of triggering intrusion behavior; Boundary force coefficient adjustment :scope N, adjust the roadside boundary constraint strength; Intrusion force amplitude adjustment amount :scope N represents the degree of aggression in the behavior of crossing the line.
[0043] This embodiment calculates various forces acting on the electric bicycle during the social force calculation and dynamic decision-making stage, including self-driving force, repulsive force, and boundary force, and introduces an intrusion force module. The trigger probability of the intrusion force is generated in real time by reinforcement learning, replacing a fixed threshold. Dynamic decision-making is achieved through reinforcement learning, solving the problem of rigid behavior. Figure 4 As shown, the technical details are as follows: (1) Calculation of social force The net force on an electric bicycle Calculated using the following formula: Self-motivation: (1) in For vehicle quality, For relaxation time, For the desired speed, For electric bicycles at the current moment The unit vector of the desired direction of motion. For electric bicycles at the current moment velocity vector; Repulsive force: (2) in The intensity of the repulsive force between vehicles. For vehicles With vehicles The distance between them The distance at which the repulsive force acts. By vehicle Pointing to the vehicle , unit vector; Boundary forces: (3) in For the intensity of the boundary action, For vehicles With boundary distance, The range of action of the boundary force. For vehicles The unit vector perpendicular to the lane's inner edge from the near-side boundary.
[0044] Intrusiveness: From dynamic probability Decide.
[0045] (2) Dynamic Intrusion Decision Calculation The electric bicycle intrusion decision probability is generated in real time by the reinforcement learning agent based on the state space. The mathematical expression for the electric bicycle intrusion decision probability is as follows: (4) in For the distance to the vehicle in front, For relative velocity, Traffic light status (0 / 1), weight parameters , , , Adjusted by the agent. When > At that time, calculate the intrusion force: (5) in, The force of intrusion is the coefficient. For electric bicycles With vehicles The distance between them The range of influence of the intrusion force. The angle between the direction of intrusion and the current direction of movement of the electric bicycle. This is a dynamically adjusted factor. Otherwise... (when hour).
[0046] Therefore, in step S1, when calculating the social force of the electric bicycle based on the intrusion decision probability, it includes: If the probability of an electric bicycle intruding is greater than the decision threshold in the adjusted hyperparameters, then calculate the self-driving force, repulsive force, boundary force, and intrusion force experienced by the electric bicycle. If the decision probability of the electric bicycle intrusion is less than the decision threshold in the adjusted hyperparameters, then the self-driving force, repulsive force, and boundary force on the electric bicycle are calculated, and the intrusion force is set to 0.
[0047] During the reinforcement learning and parameter optimization phase, the reinforcement learning agent adjusts the hyperparameters of the social force model (such as the intensity of repulsion force and the decision threshold) according to the environmental state, and guides the optimization through the reward function.
[0048] Specifically, during iterative training of the reinforcement learning agent, the Proximal Policy Optimization (PPO) algorithm is used, and the reward function design employs a multi-objective reward function to guide the optimization direction. The mathematical expression of the reward function for the reinforcement learning agent is as follows: (6) in: It is a trajectory error reward: (7) In the above formula, This refers to the total number of points or time steps used for comparison in the trajectory. This refers to the trajectory data of an electric bicycle generated through a simulation model. This data represents actual, observed electric bicycle trajectory data. Penalties for violating safe distance rules: (8) In the above formula, For indicator functions, This represents the actual distance between the electric bicycle and other vehicles or obstacles in the simulation. It's an efficiency reward: (9) In the above formula, This represents the total number of time steps in the simulation or evaluation process. For electric bicycles The instantaneous velocity scalar value at a certain time step.
[0049] Furthermore, reinforcement learning agents are based on reward functions. The parameters of the social force model are iteratively adjusted, among which... These represent the historical optimal values of the parameters for the social force model. The optimization range of the key parameters of the social force model is shown in Table 1.
[0050] Table 1. Optimization Range of Social Force Model Parameters
[0051] In the resultant force calculation and state update stage, this embodiment calculates the resultant force by integrating all forces and updates the vehicle's position, speed, and other states. Specifically, in step S1, when updating the electric bicycle's vehicle state based on the social force calculation results, the following is included: Calculate the resultant force of the self-driving force, repulsive force, boundary force, and intrusion force as the resultant force acting on the electric bicycle: (10) Update the vehicle status of the electric bicycle using Newton's laws of motion: (11) (12) (13) in, The simulation step size is typically 0.1s. After the state is updated, the social force calculation step is returned to form a closed loop. In the next iteration, the reinforcement learning agent updates the state space based on these new states and calculates the new electric bicycle intrusion decision probability and the new action space. After the social force model adjusts the hyperparameters based on the new action space, it calculates the social force of the electric bicycle based on the new electric bicycle intrusion decision probability and updates the vehicle state of the electric bicycle again.
[0052] In the termination and output phase, when the training in step S1 converges (the reward fluctuation is less than the threshold or the maximum number of iterations is reached), the optimized simulation model is output.
[0053] The effectiveness of the method in this embodiment will be verified through experiments below: (1) Simulation scene construction SUMO 1.15.0 can be used as the simulation engine, Python 3.8 as the control scripting language, and TensorFlow 2.0 to implement the reinforcement learning algorithm.
[0054] Taking a standard crossroads as an example, the environment is built using the NETEDIT tool. Key parameter settings can be selected within the range of intersection diameter (30-50 meters), approach lane width (5-7 meters), and signal control cycle (90-120 seconds). Parameters should be determined according to actual needs. Traffic flow is generated based on a negative binomial distribution, with vehicle arrival rates ranging from 800 to 2400 vehicles per hour, covering low, medium, and high density scenarios.
[0055] (2) Reinforcement learning initialization The Proximal Policy Optimization (PPO) algorithm is adopted, and the network structure is an Actor-Critic architecture. The hyperparameter settings can be found in the table below.
[0056] Table 2 Hyperparameter Settings
[0057] (3) Create Individual List Module Used to initialize individual attributes of motor vehicles and non-motor vehicles, including lists of motor vehicles and electric bicycles, including attributes such as number, coordinates, mass, speed, and acceleration.
[0058] (4) Install electric bicycle sensing module It is used to monitor the status of surrounding vehicles in real time. The sensing range is set to a fan shape with a radius of 15 meters (it can also be set to an ellipse or rectangle). The monitoring elements include features such as relative position, velocity vector, acceleration, and spacing, and further calculate the time difference of the expected conflict point between vehicles.
[0059] (5) Calculate the probability of electric bicycle intrusion decision. Decision thresholds are dynamically optimized using reinforcement learning. (Range 0.1-0.9) and substitute into formula (4) to calculate the probability of electric bicycle intrusion decision.
[0060] (6) Strengthen learning and training When iteratively training reinforcement learning agents, the training parameters are adaptively optimized using different traffic flow densities, which can be divided into low density, medium-high density, and mixed density training. For low density, the number of iterations is 1-1500, with the optimization focus on basic social force parameters; for medium-high density, the number of iterations is 1501-3000, with the optimization focus on decision thresholds; and for mixed density, the number of iterations is 3001-5000, with the optimization focus on generalization ability.
[0061] The Proximal Policy Optimization (PPO) algorithm is adopted, and a multi-objective reward function is designed according to formulas (6) to (9) to guide the optimization direction. The convergence condition is selected that the reward fluctuation is less than 0.01 for 100 consecutive iterations. The optimized parameters are used for the final simulation output.
[0062] (7) Status update and effect verification Perform the resultant force calculation according to formula (10), and update the vehicle status according to formulas (11) to (13) based on the resultant force calculation results.
[0063] (8) Data output The output data is in CSV format and includes trajectory, velocity, acceleration, etc. The model performance is evaluated by calculating metrics such as trajectory error and number of collisions.
[0064] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0065] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A simulation method for electric bicycle intrusion behavior based on reinforcement learning parameter adaptation, characterized in that, Includes the following steps: The environmental parameters are acquired and the reinforcement learning agent is initialized. The environmental parameters include vehicle parameters, traffic flow parameters, and road parameters. The action space of the reinforcement learning agent includes the hyperparameter adjustment amount of the social force model. The reinforcement learning agent is trained iteratively until the convergence condition is met to obtain the simulation model. In each iteration, the reinforcement learning agent acquires all vehicle states and updates the state space, then calculates the corresponding action space, and calculates the electric bicycle intrusion decision probability based on the state space. The social force model adjusts the hyperparameters according to the hyperparameter adjustment amount in the action space, then calculates the social force of the electric bicycle based on the electric bicycle intrusion decision probability, and finally updates the vehicle state of the electric bicycle based on the social force calculation result. The simulation model was used to simulate the intrusion behavior of electric bicycles.
2. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 1, characterized in that, The state space of the reinforcement learning agent includes the vehicle's own state features and the interaction environment features. The vehicle's own state features include the vehicle's position coordinates, velocity vector, and acceleration. The interaction environment features include the distance between the vehicle and the target point, the distance between the vehicle and the vehicle in front, the distance between the vehicle and the vehicle behind, relative speed, traffic light status, and traffic density.
3. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 1, characterized in that, When calculating the intrusion decision probability of an electric bicycle based on the state space, the mathematical expression for the intrusion decision probability of an electric bicycle is as follows: in, The distance between the vehicle and the vehicle in front. For relative velocity, Traffic light status. , , , All of these are weight parameters.
4. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 1, characterized in that, The hyperparameters of the social force model include the decision threshold. When calculating the social force of an electric bicycle based on the probability of its intrusion decision, the following parameters are included: If the probability of an electric bicycle intruding is greater than the decision threshold in the adjusted hyperparameters, then calculate the self-driving force, repulsive force, boundary force, and intrusion force experienced by the electric bicycle. If the decision probability of the electric bicycle intrusion is less than the decision threshold in the adjusted hyperparameters, then the self-driving force, repulsive force, and boundary force on the electric bicycle are calculated, and the intrusion force is set to 0.
5. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 4, characterized in that, The hyperparameters of the social force model also include repulsive force strength, influence range, boundary force coefficient, and attenuation distance. The calculation formula for the self-driving force is as follows: in, For vehicle quality, For relaxation time, For the desired speed, For electric bicycles at the current moment The unit vector of the desired direction of motion. For electric bicycles at the current moment velocity vector; The formula for calculating the repulsive force is as follows: in, The strength of the repulsive force. For vehicles With vehicles The distance between them For the scope of influence, By vehicle Pointing to the vehicle , unit vector; The formula for calculating the boundary force is as follows: in, These are the boundary force coefficients. For vehicles With boundary distance, For attenuation distance, For vehicles A unit vector perpendicular to the lane's inner edge from the near-side boundary; The formula for calculating the intrusion force is as follows: in, The force of intrusion is the coefficient. For electric bicycles With vehicles The distance between them The range of influence of the intrusion force. The angle between the direction of intrusion and the current direction of movement of the electric bicycle. This is a dynamically adjusted factor.
6. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 4, characterized in that, When updating the vehicle state of an electric bicycle based on the social force calculation results, the resultant force of the self-driving force, repulsive force, boundary force, and intrusion force is calculated, and then the vehicle state of the electric bicycle is updated using Newton's laws of motion.
7. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 1, characterized in that, When iteratively training the reinforcement learning agent, a proximal policy optimization algorithm is used. The reward function design employs a multi-objective reward function to guide the optimization direction. The mathematical expression of the reward function of the reinforcement learning agent is as follows: in: It is a reward for trajectory error. , This refers to the total number of points or time steps used for comparison in the trajectory. This refers to the trajectory data of an electric bicycle generated through a simulation model. This data represents actual, observed electric bicycle trajectory data. It is a penalty for violating the safe distance rule. , For indicator functions, This represents the actual distance between the electric bicycle and other vehicles or obstacles in the simulation. It's an efficiency reward. , This represents the total number of time steps in the simulation or evaluation process. For electric bicycles The instantaneous velocity scalar value at a certain time step.
8. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 1, characterized in that, During iterative training of the reinforcement learning agent, the training parameters are adaptively optimized using different traffic flow densities. Training is conducted at low density, medium-high density, and mixed density. The low density iteration count is 1-1500, with the optimization focus on the basic social force parameters; the medium-high density iteration count is 1501-3000, with the optimization focus on the decision threshold; and the mixed density iteration count is 3001-5000, with the optimization focus on generalization ability.
9. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 1, characterized in that, The convergence condition is that the number of consecutive iterations in which the reward fluctuation is less than a threshold reaches a specified value or the maximum number of iterations is reached.
10. The method for simulating electric bicycle intrusion behavior based on reinforcement learning parameter adaptation according to claim 9, characterized in that, The threshold is 0.1 and the specified value is 100.