Urban traffic reinforcement learning parallel training method based on loose synchronization
By adopting loose synchronization and shadowed edge/node methods in reinforcement learning, long training time and synchronization problems in large-scale traffic networks are solved, and efficient traffic information integration and simulation accuracy are achieved.
Patent Information
- Application Number
- CN202510297260.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The existing reinforcement learning has a long training time in traffic flow control, especially in large-scale traffic networks. Each intersection requires independent decision-making, resulting in a significant increase in the total interaction time. The parallel processing technology has synchronization and data consistency problems in traditional frameworks, which affects training efficiency.
A parallel training method for urban traffic reinforcement learning based on loose synchronization is proposed. Through road network partitioning and initialization, preliminary simulation and data recording, loose synchronization communication, boundary processing, parallel reward calculation, model training and strategy update, convergence verification and other steps, the synchronous communication frequency is reduced, and shadow edges and shadow nodes are introduced to ensure the fusion of neighborhood information.
It significantly reduces synchronization overhead, improves training speed, ensures the integrity of traffic information and the effectiveness of neighborhood cooperation, and improves the accuracy of traffic condition simulation.
Smart Images

Figure CN120219137A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reinforcement learning, and particularly to a parallel training method for urban traffic reinforcement learning based on loose synchronization. Background Art
[0002] Traffic control strategies are an important part of urban management. The use of advanced information technology and intelligent algorithms, especially the application of reinforcement learning technology to traffic management, has become a new strategy for effectively alleviating urban traffic problems. Reinforcement learning is a machine learning technology that optimizes decision-making strategies through interaction with the environment. In the field of traffic management, the application of reinforcement learning includes traffic signal control, road pricing, and public transport system scheduling, etc., showing broad application prospects. The control strategy of reinforcement learning depends on the data generated by interacting with the traffic environment, and explores to make correct decisions under different results. However, the real-world urban traffic cannot provide enough interactive data to train these policies, because the exploration of policies may have a negative impact on urban traffic, such as causing traffic congestion. Therefore, traffic simulators were born as an alternative, allowing researchers to test and evaluate their RL strategies without disturbing the actual traffic flow. These simulators obtain traffic action from the decisions made by traffic control strategies and simulate the state of road networks and vehicles in the simulator. Currently widely used traffic simulators include SUMO, CityFlow, etc.
[0003] Although the application of reinforcement learning in traffic flow control performs well, its training time is relatively long because it needs to go through multiple rounds of training to converge to the optimal strategy. When the scale of the traffic network expands, especially in multi-agent reinforcement learning, each intersection needs to make independent decisions, and the total interaction time will increase significantly.
[0004] In the field of traffic simulation, to address the high computational demands of large-scale traffic network simulation, researchers have developed various parallel processing-based methods to accelerate the simulation speed. However, applying parallel processing techniques to reinforcement learning training also brings a series of new challenges: In traditional parallel traffic simulation frameworks, multi-threading or distributed computing methods are usually adopted. Due to the limitations of thread synchronization and conflict handling, their acceleration effects in large-scale traffic environments are not ideal. Similarly, methods based on multi-machine parallelism also face many challenges. In traditional distributed computing frameworks, multiple simulation instances or agents execute independently on different computing nodes, introducing complex synchronization and data consistency issues, which may affect the efficiency of the learning process and the final performance of the model. In addition, multi-machine parallelism incurs additional communication overhead when solving synchronization problems. Given that reinforcement learning inherently requires multiple rounds of training to achieve model convergence, if the cost of synchronous communication is too high, it may seriously slow down the overall training efficiency. At the same time, parallel-based processing methods affect the integrity of road network and vehicle data in reinforcement learning training, thereby further affecting the neighborhood cooperation of reinforcement learning for traffic control.
[0005] Therefore, we propose a parallel training method for urban traffic reinforcement learning based on loose synchronization. Summary of the Invention
[0006] The present invention mainly solves the technical problems existing in the above-mentioned prior art and provides a parallel training method for urban traffic reinforcement learning based on loose synchronization.
[0007] To achieve the above object, the present invention adopts the following technical solution. A parallel training method for urban traffic reinforcement learning based on loose synchronization includes the following steps:
[0008] S1. Road network partitioning and initialization: Before performing reinforcement learning training, divide the traffic network into multiple sub-road network partitions, and each sub-road network partition is processed by an independent process; at the same time, initialize the traffic simulator and reinforcement learning model within each partition to prepare for subsequent simulation and training;
[0009] S2. Preliminary simulation and data recording: In the first round of reinforcement learning training, the multi-threaded traffic simulators in each sub-road network partition perform preliminary simulations simultaneously; accurately record all vehicle information expected to cross the sub-road network partition boundary, including vehicle identity, driving trajectory, and the predetermined time step to enter the adjacent partition;
[0010] S3. Loose Synchronous Communication: After the first round of reinforcement learning training, information synchronization is carried out for different sub-road network zones; using the initially simulated vehicle information crossing the boundary of the sub-road network zone as the starting data, and the results of each subsequent round of reinforcement learning training as iterative data to update the time step information of cross-boundary vehicles; the traffic simulation process of the sub-road network adds vehicles to the sub-road network zone at the correct time step according to the updated time step information of cross-boundary vehicles.
[0011] S4. Boundary Processing: Copy the edges and nodes adjacent to the partition before the partition at the boundary of each sub-road network zone to form shadow edges and shadow nodes; at each time step during the training process of the reinforcement learning model, send vehicle information to the shadow nodes of the relevant sub-road network zones to keep the partition state synchronized and ensure the fusion of neighborhood information.
[0012] S5. Parallel Reward Calculation: Only the parts related to interacting with the traffic environment and calculation in reinforcement learning are processed in parallel; each parallel processing process is responsible for obtaining the information of the sub-road network zone and calculating the pressure value of each intersection. After each process finishes the calculation, it returns a list array to the main process. The main process merges and removes the shadow nodes, and finally obtains the pressure value of each intersection and the final reward value.
[0013] S6. Model Training and Policy Update: The intelligent agent updates the decision-making policy using the reinforcement learning algorithm according to the traffic state information obtained from the traffic simulator, including the number of vehicles, traffic density, and signal light status; transmits the decision result to the traffic simulator, and the simulator updates the environment state accordingly and calculates the reward to feedback to the intelligent agent, forming a closed-loop learning process.
[0014] S7. Convergence Verification: Define the state s, action a, and reward r; according to the Bellman equation
[0015]
[0016] where the state s is the new state reached after taking the action a, γ is the discount factor of future rewards, θ - is the parameter of the target network, Q(s,a) is the state-action value function, representing the expected return obtained by taking the action a in the state s, Q target is the target Q value;
[0017] Value update formula
[0018]
[0019] Define the error function
[0020] δ = Q target (s,a) - Q(s,a;θ);
[0021] By analyzing the updated error
[0022]
[0023] Verify whether the training result tends to converge under the conditions that the learning rate is small enough, the gradient has a sufficient magnitude and is stable. When it tends to converge, the training is completed.
[0024] Preferably, in the loose synchronous communication step S3, when updating the time step information of the cross-border vehicle, it is adjusted based on the differences between the actual driving speed and distance of the vehicle during the current round of training and the expected situation. If the actual speed is faster than the expected speed and the driving distance exceeds 10% of the original path, the time step for the cross-border vehicle to enter the adjacent partition is advanced accordingly.
[0025] Preferably, in the boundary processing step S4, when the reinforcement learning model sends vehicle information to the shadow node, the information includes the vehicle speed and driving direction, so as to more accurately simulate the vehicle behavior at the boundary of the sub-road network partition and maintain the integrity of the neighborhood information.
[0026] Preferably, in the parallel reward calculation step S5, when each parallel processing process calculates the pressure value, in addition to considering the vehicle density of the incoming and outgoing lanes, the waiting time of the vehicle is also combined to comprehensively calculate the pressure value:
[0027]
[0028] where ρ represents the vehicle density, t wait represents the vehicle waiting time, t total is the total duration of this time step, w ρ is the weight of the vehicle density, w t is the weight of the waiting time.
[0029] Preferably, in the model training and policy update step S6, the agent adjusts the ratio of exploration to exploitation in an exponentially decaying manner according to the reward value feedback by the traffic simulator, exploring more at the beginning of training and gradually increasing the ratio of using the existing policy as the training progresses:
[0030] P explore = P0 × α n ,
[0031] where, P explore is the current exploration probability, P0 is the initial exploration probability, α is the decay factor, and n is the number of training rounds.
[0032] Preferably, in the convergence verification step S7, in addition to satisfying the conditions that the learning rate is small enough, the gradient has a sufficient magnitude and is stable, it is also required that the difference between the error function values of two adjacent rounds of training converges during multiple consecutive rounds of training to determine that the training result tends to converge.
[0033] Beneficial effects
[0034] The present invention provides a parallel training method for urban traffic reinforcement learning based on loose synchronization, which has the following beneficial effects:
[0035] (1) In the traditional parallel traffic simulation and reinforcement learning training of the parallel training method for urban traffic reinforcement learning based on loose synchronization, frequent synchronous communication occupies a large amount of performance. The present invention proposes to perform cross-regional synchronization of information only at the end of each round of reinforcement learning training, replacing the traditional method of performing information synchronization once at each time step of each round, significantly reducing the communication frequency. This avoids the problem of slowing down the overall training efficiency due to the too high cost of synchronous communication. In large-scale traffic network simulation and reinforcement learning training, it can effectively reduce the synchronous overhead and improve the training speed.
[0036] (2) In the parallel training method for urban traffic reinforcement learning based on loose synchronization, the parallel processing method will affect the integrity of road network and vehicle data, and thus affect the neighborhood cooperation of reinforcement learning. The present invention introduces shadow edges and shadow nodes to replicate adjacent edges and nodes at the sub-road network partition boundary. During the training process, the reinforcement learning model sends vehicle information to the shadow nodes to ensure the integration of neighborhood information and maintain the integrity of traffic information at the sub-road network partition boundary. Taking traffic signal control as an example, it ensures the co-optimization of signal timing, avoids the negative impact of a single signal decision on the traffic flow of adjacent signals, and improves the accuracy of the overall traffic condition simulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and those of ordinary skill in the art can also obtain other implementation drawings according to the provided drawings without creative efforts.
[0038] Figure 1 It is the framework diagram of the present invention;
[0039] Figure 2 It is the comparison diagram of loose synchronization and time step synchronization of the present invention;
[0040] Figure 3 It is the schematic diagram of boundary processing of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] Embodiment 1:
[0043] Regarding the performance problems existing in the parallel traffic simulation for reinforcement learning training, a parallel traffic simulation framework for reinforcement learning training is provided to solve the performance problems. The framework consists of two parts: (1) A parallel training mechanism for urban traffic reinforcement learning based on loose synchronization that ensures the correctness of reinforcement learning training through loose synchronization and boundary data structure design. (2) A parallel traffic simulation acceleration method that can improve the training speed of traffic simulation for reinforcement learning under the parallel mechanism through a load-balanced road network partitioning algorithm. However, there are still two difficulties in ensuring the performance and accuracy of this framework:
[0044] 1) How to reduce the synchronization overhead between partitions during the training process under data parallelism. Reinforcement learning essentially requires multiple rounds of training to achieve model convergence. If the cost of synchronous communication is too high, it may seriously slow down the overall training efficiency.
[0045] 2) How to ensure the correctness and effectiveness under the parallel mechanism. Environment simulation and interactive calculation are key factors in accelerating the reinforcement learning training process, and these two parts need to be optimized in the parallel computing framework. At the same time, the parallel processing method will affect the integrity of the road network and vehicle data in the reinforcement learning training, thus posing challenges to the neighborhood cooperation of reinforcement learning.
[0046] Therefore, a parallel training method for urban traffic reinforcement learning based on loose synchronization is provided, including the following technical points:
[0047] 1. Propose loose synchronization for reinforcement learning to solve the problem of too high cost of synchronous communication;
[0048] 2. Introduce shadow nodes to ensure the neighborhood cooperation of reinforcement learning;
[0049] 3. Design a parallel interactive calculation method to optimize the time overhead of this part;
[0050] 4. Prove the convergence of the proposed loose synchronization mechanism under reinforcement learning.
[0051] Specifically, before elaborating on the technical key points, it should be noted that in traditional traffic simulation reinforcement learning, each reinforcement learning training episode, that is, the complete process where the agent starts from the initial state and continuously interacts with the environment until it reaches a certain termination state. Each reinforcement learning training episode consists of multiple time steps. Before conducting parallel traffic simulation, the entire traffic network needs to be divided into multiple sub-road network partitions, and each sub-road network partition is handed over to a training process for distributed training. In each time step of the parallel traffic simulation reinforcement learning training episode, a synchronous communication is performed. The synchronous communication is to synchronously integrate the training information of multiple processes to prepare for the training of the next time step.
[0052] Among them, the time step is the basic time unit in the discrete time flow mechanism adopted by the traffic simulator. The traffic simulator divides the entire simulation time into equal-sized intervals, and these intervals are the time step lengths. Each time step usually corresponds to one second in the real world. In each time step, the traffic simulator simulates the behaviors of each vehicle in that time step length and updates the states of the vehicles and roads in the road network. At the end of each time step, the reinforcement learning agent obtains the necessary traffic state information from the simulator, including the number of vehicles, traffic density, the current state of traffic lights, etc.
[0053] As Figure 2 shown, the loose synchronization scheme proposed by the present invention adjusts the communication method and frequency of parallel traffic simulation in the prior art. In particular, considering the characteristic that reinforcement learning is an iterative process, that is, information synchronization is performed in each time step of each round, which occupies a large amount of performance. The present invention performs cross-region synchronization of information only at the end of each round of reinforcement learning training, significantly reducing the communication frequency.
[0054] Specifically, after partitioning the traffic network, a round of reinforcement learning training is carried out simultaneously. In the first round of training, all vehicles that are expected to cross the boundary of the sub-road network partition and their scheduled time steps to enter the adjacent partition are accurately recorded. Such preprocessing ensures that all data is complete and accurate in the first round of training of reinforcement learning.
[0055] Subsequently, after each round of training, the time step information of these cross-boundary vehicles is updated according to the results of the simulation in that round. In actual training, the driving conditions of the vehicles in each round of training are different. For example, in a certain training, vehicle A was expected to travel at a speed of 30 km / h, but the actual speed reached 35 km / h, and the driving distance within this sub-region exceeded 10% of the expected distance. At this time, according to the rules we set, based on the speed increase ratio and the situation of the driving distance exceeding, the time step for it to enter the adjacent sub-region is advanced accordingly. Suppose it was originally expected to enter the adjacent sub-region at the 10th time step, and after adjustment, it is advanced to the 9th time step. This can more accurately simulate the transfer of vehicles between sub-regions and ensure the accuracy of information synchronization. The training process of simulating the traffic of each sub-road network will add vehicles to it at the correct time step according to the latest boundary vehicle information, so as to complete the efficient communication and synchronization between sub-regions.
[0056] After dividing the road network into multiple sub-road network partitions, the issue of the sub-road network boundary is also involved. The parallel processing method will affect the integrity of the sub-road network and vehicle data in the reinforcement learning training, thereby further affecting the neighborhood cooperation of the reinforcement learning for traffic control. Taking traffic signal control as an example, in an urban environment, traffic signals are usually not far apart, which requires joint optimization of signal timing. Usually, this process is called "coordinated signal timing". Since many scenarios in multi-agent reinforcement learning use neighborhood information for fusion, the division of the road network will affect the integrity of the road network, and thus affect the accuracy of the state in the reinforcement learning training. Therefore, when considering parallelism, it is necessary to process the boundaries of the sub-road network partitions.
[0057] As Figure 3 shown, in the present invention, the method of shaded edges and shaded nodes is adopted.
[0058] Specifically, in the overall road network, there is an edge connection between node J1 and node J2. After the sub-road network partition, the connection between J1 and J2 is broken. To ensure that the sub-road network partition does not affect the integrity of the neighborhood information of nodes J1 and J2, the shaded node J2' is replicated in the partition P1 containing node J1, and then the shaded edge J1J2' connecting J1 and J2' is generated. At the same time, the shaded node J1' is replicated in the partition P2 containing node J2, and then the shaded edge J2J1' connecting J2 and J1' is generated. Among them, the lengths of the shaded edges J1J2' and J2J1' are the same as the length of the edge J1J2. At each time step of the training process, the reinforcement learning model will send decisions on nodes J1 and J2 to partitions P1 and P2 simultaneously to keep the states of these two partitions synchronized.
[0059] For a vehicle, when a vehicle v moves from intersection J0 to J1 and continues to move towards intersection J2 through edge J1J2, in this case, the vehicle v will appear on the shaded edge J1J2' in partition P1 and move towards the replicated shaded node J2'. According to the synchronization mechanism in the parallel framework mentioned above, the vehicle v will be inserted onto the shaded edge J1J2' in partition P2 at the end of time step t according to the time t when it entered road J1J2 in the previous round of reinforcement learning training. In this way, both partitions will perform local updates on the vehicle v, thus ensuring the integrity of vehicle information within both partitions. v , and will be inserted onto the shaded edge J1J2' in partition P2 at the end of time step t v-1 .
[0060] During the training process, the reinforcement learning model sends vehicle speed and driving direction information to the shaded nodes, which is crucial for accurately simulating traffic conditions. For example, when a vehicle moves from partition P1 towards the shaded node J2', its speed information enables partition P2 to prepare in advance for the vehicle's entry, predict the vehicle's arrival time based on the speed, and reasonably arrange the traffic flow within this partition. The driving direction information helps to determine the driving path of the vehicle after it enters partition P2, ensuring that the simulation of vehicle behavior in both partitions remains consistent and maintaining the integrity of neighborhood information. For instance, if a vehicle approaches J2' at a specific angle, partition P2 can determine from which direction the vehicle will enter the roads in this partition, avoiding traffic simulation chaos.
[0061] Based on the above, the present invention makes the interaction and calculation parts with the traffic environment in the reinforcement learning training parallel. Taking traffic signal control as an example, at the end of each simulation time step, the traffic simulator needs to update the environmental state according to the decision result of the reinforcement learning agent and calculate the reward value (Reward) under this decision and feedback it to the agent. The calculation of the reward value usually involves global state interaction with the traffic simulator and the calculation of relevant reward value definitions, and this part will become complex as the road network and vehicle scale increase.
[0062] Without modifying the reinforcement learning training framework, this paper performs parallel processing on information acquisition and calculation in reinforcement learning. Taking "pressure", which is widely used in the field of traffic signal control, as an example, the pressure P of intersection i i represents the degree of imbalance in vehicle density between incoming and outgoing lanes. The formula for defining the reward value R using pressure is: R = -∑P i .
[0063] Therefore, calculating the reward requires obtaining the pressure values of all intersections and summing them up. In a parallel framework, the road network is divided into several regions, and each process is responsible for obtaining the information on this part of the road network and calculating the pressure value of the intersections, optimizing the calculation time of this part through parallel processing. After each process has completed the calculation, it returns a list array to the main process. The main process merges them and removes the shaded nodes, and finally obtains the pressure values of each intersection and the final reward value.
[0064] Taking a sub-road network partition as an example, when calculating the pressure value, in addition to considering the vehicle density of the incoming and outgoing lanes, the vehicle waiting time is also combined as follows:
[0065]
[0066] where ρ represents the vehicle density, t wait represents the vehicle waiting time, T total is the total duration of this time step, w ρ is the weight of the vehicle density, w t is the weight of the waiting time. We first count the waiting time of vehicles in each lane of each intersection within this partition. Assume that for a certain lane of intersection K, the total waiting time of vehicles within one time step is 20 seconds, and at the same time, the vehicle density of this lane is calculated to be 0.5 (assumed density calculation value). According to the calculation method we set, weights are assigned to the vehicle density and the waiting time. For example, the weight of the vehicle density is 0.6, and the weight of the waiting time is 0.4. Then the pressure value of this lane at this intersection is calculated as: Pressure value = 0.5×0.6 + (20÷total duration of this time step). Assume it is 60 seconds, then here it is (20÷60)×0.4. The pressure value calculated in this comprehensive way can better reflect the actual traffic congestion situation, making the calculation of the reward value more in line with the actual traffic scenario.
[0067] In the model training and policy update stage, the agent adjusts the exploration and exploitation ratio using an exponential decay method
[0068] P explore = P0×α n ,
[0069] where, P explore is the current exploration probability, P0 is the initial exploration probability, α is the decay factor, and n is the number of training rounds. At the beginning of training, the exploration probability is set to 0.9. For example, in the first 10 rounds of training, the agent has a 90% probability of randomly selecting actions and trying different signal control strategies to explore more traffic states and potential effective strategies. As the number of training rounds increases, the exploration probability gradually decreases according to the exponential decay formula (exploration probability = 0.9×0.95 训练轮数 ). When training reaches the 50th round, the exploration probability is approximately 0.9×0.95 50≈0.07, at this time, the agent makes more use of existing successful strategies, makes decisions based on the previously accumulated experience, improves the decision-making efficiency, and accelerates the model convergence.
[0070] For the effectiveness of the results of this training mechanism, the present invention proves it based on single-step convergence. Taking the traffic signal control problem as an example, each intersection can be regarded as an environment for reinforcement learning, where:
[0071] State s: includes information such as the number of vehicles and waiting time in each direction of the intersection.
[0072] Action a: represents the on / off state of the traffic lights in each direction or the change in the duration.
[0073] Reward r: is usually related to reducing the travel time and preventing traffic congestion. For example, for every one-second reduction in the total travel time, the reward increases by one unit.
[0074] Define the Bellman equation used in the reinforcement learning training as follows:
[0075]
[0076] where the state s is the new state reached after taking the action a, γ is the discount factor for future rewards, θ - is the parameter of the target network, Q(s,a) is the state-action value function, representing the expected return obtained by taking the action a in the state s, and Q target is the target Q value.
[0077] The Q value is updated using the following formula:
[0078]
[0079] where α is the learning rate and θ is the parameter of the current network.
[0080] Therefore, the error function δ is defined as follows in this paper. Our goal is to reduce this error through the training process:
[0081] δ = Q target (s,a) - Q(s,a;θ)..
[0082] To prove the single-step convergence, what we are concerned about is the updated error δ′, that is:
[0083]
[0084] If is a suitable small positive number, this indicates that θ′ is less than θ, meaning that the error decreases after each update, that is, the training result tends to converge. The following conditions need to be considered and satisfied:
[0085] 1) The learning rate α should be small enough to ensure that the updates are gradual and do not cause over-updating of the parameters (i.e., avoid instability caused by overly large step sizes);
[0086] 2) The gradient should be large enough (non-zero) and stable to ensure an effective learning process.
[0087] When judging whether the training converges, in addition to paying attention to the learning rate and gradient conditions, it is also necessary to look at the difference between the error function values of two adjacent rounds of training in multiple consecutive rounds of training. For example, during the training process, record the error function value after each round of training. Suppose the error value in the nth round is 0.5 and the error value in the (n + 1)th round is 0.498, and the difference between the two is 0.002. When the differences in multiple consecutive rounds (such as 10 consecutive rounds) are all less than 0.001, combined with the conditions that the learning rate is small enough, the gradient is large enough and stable, we can more reliably determine that the training result tends to converge and ensure that the model training reaches a stable state.
[0088] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A parallel training method for urban traffic reinforcement learning based on loose synchronization, characterized in that: The following steps are involved: S1. Road network partitioning and initialization: Before reinforcement learning training, the traffic network is divided into multiple sub-road network partitions, each of which is processed by an independent process; at the same time, the traffic simulator and reinforcement learning model in each partition are initialized to prepare for subsequent simulation and training; S2. Preliminary simulation and data recording: During the first round of reinforcement learning training, a multi-threaded traffic simulator for each sub-network partition simultaneously performs preliminary simulations; accurately records all vehicle information that is expected to cross the sub-network partition boundary, including vehicle identity, driving trajectory, and scheduled time step to enter the adjacent partition; S3, loosely synchronized communication: after the first round of reinforcement learning training, information synchronization is performed between sub-network partitions; The initial simulated vehicle information crossing the sub-network partition boundary is used as the starting data, and the subsequent reinforcement learning training results are used as iterative data to update the time step information of the cross-border vehicle; the traffic simulation process of the sub-network adds vehicles to the sub-network partition at the correct time step based on the updated cross-border vehicle time step information; S4, boundary processing: copy the adjacent edges and nodes before the partition at the boundary of each sub-network partition to form shadow edges and shadow nodes; At each time step of the training process, the reinforcement learning model sends vehicle information to the shadow nodes of the relevant sub-network partitions to keep the partition status synchronized and ensure the fusion of neighborhood information; S5, parallel reward calculation: In reinforcement learning, only the interaction with the traffic environment and the calculation part are processed in parallel; each parallel processing process is responsible for obtaining the information of the sub-road network partition and calculating the pressure value of each intersection. After each process is calculated, a list array is returned to the main process. After the main process merges and removes the shadow nodes, the pressure value and the final reward value of each intersection are finally obtained; S6, model training and strategy update: The agent uses the reinforcement learning algorithm to update the decision strategy based on the traffic status information obtained from the traffic simulator, including the number of vehicles, traffic density, and signal light status; the decision results are passed to the traffic simulator, which updates the environment status and calculates rewards to feed back to the agent, forming a closed-loop learning process; S7, Convergence verification: define state s, action a, reward r; according to the Bellman equation where state s is the new state reached after taking action a, γ is the discount factor for future rewards, and θ - is the parameter of the target network, Q(s,a) is the state-action value function, which represents the expected return of taking action a in state s, and Q target is the target Q value; Q value update formula Where α is the learning rate and θ is the parameter of the current network; Define the error function δ δ=Q target (s,a)-Q(s,a;θ); By analyzing the updated error δ′ Verify that the learning rate α is small enough and the gradient Under the condition of sufficient size and stability, the training results tend to converge. When they tend to converge, the training is completed.
2. The urban traffic reinforcement learning parallel training method based on loose synchronization according to claim 1 is characterized by: In step S3, when updating the time step information of the cross-border vehicle, adjustments are made based on the difference between the actual vehicle speed and distance in the current round of training and the expected situation. If the actual speed is faster than the expected speed and the distance exceeds 10% of the original path, the time step for the cross-border vehicle to enter the adjacent partition is advanced accordingly.
3. The urban traffic reinforcement learning parallel training method based on loose synchronization according to claim 1 is characterized in that: In step S4, when the reinforcement learning model sends vehicle information to the shadow node, the information includes the speed and driving direction of the vehicle, so as to more accurately simulate the behavior of the vehicle at the boundary of the sub-road network partition and maintain the integrity of the neighborhood information.
4. The urban traffic reinforcement learning parallel training method based on loose synchronization according to claim 1 is characterized in that: In step S5, when each parallel processing process calculates the pressure value, in addition to considering the density of vehicles entering and exiting the lane, it also combines the waiting time of the vehicle to comprehensively calculate the pressure value: Where ρ represents the vehicle density, t wait Represents the vehicle waiting time, T total is the total duration of the time step, w ρ is the weight of vehicle density, w t is the weight of the waiting time.
5. The urban traffic reinforcement learning parallel training method based on loose synchronization according to claim 1 is characterized by: In step S6, the agent adjusts the ratio of exploration and utilization in an exponential decay manner according to the reward value fed back by the traffic simulator, performing more exploration in the early stages of training and gradually increasing the ratio of utilizing existing strategies as training progresses: P explore =P0×α n , Among them, P explore is the current exploration probability, P0 is the initial exploration probability, α is the decay factor, and n is the number of training rounds.
6. The urban traffic reinforcement learning parallel training method based on loose synchronization according to claim 1 is characterized by: In step S7, in addition to satisfying the conditions that the learning rate is small enough and the gradient is large enough and stable, it is also required that the difference between the error function values of two adjacent rounds of training converges in multiple consecutive rounds of training before it is determined that the training result tends to converge.