Traffic Signal Control Method, Device, Electronic Device and Storage Medium
Through the ETS-RL algorithm, the traffic signal control model is optimized, the upstream and downstream pressure at the intersection is coordinated, and the model complexity and delay problems in the existing methods are solved, and traffic pressure reduction and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202310183083.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-02-28
AI Technical Summary
The existing adaptive traffic signal control method improves model complexity while only slightly reducing traffic delays, and ignores the basic traffic state representation, resulting in reduced model usability and driver action.
Using the ETS-RL algorithm, by coordinating the effective pressure upstream and downstream of the intersection and the effective driving vehicles within a defined range, a new traffic state representation is designed, phase probability is calculated, signal phase timing scheme is optimized, traffic signal control model is established and trained.
Effectively reduce traffic pressure, reduce vehicle waiting time, improve traffic efficiency, and alleviate traffic congestion. At the same time, the model is simple and highly usable.
Smart Images

Figure CN116189454B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic signal control, and in particular to a traffic signal control method, device, electronic equipment and storage medium. Background Art
[0002] With the development of social economy and urbanization, people's travel methods have also changed, resulting in a rapid increase in road vehicles, and the resulting traffic problems are also increasing. The existing road traffic management system can no longer adapt to today's traffic pressure. Many problems such as traffic jams, traffic accidents, environmental pollution and energy waste not only affect the development of the country and social progress, but also bring a lot of inconvenience to daily travel. Reasonable control and guidance of traffic flow at intersections is an inevitable requirement for improving traffic efficiency and alleviating traffic congestion, and it is also the only way to ensure traffic safety and maintain ecological sustainable development.
[0003] Existing adaptive traffic signal control usually models traffic movement as a queuing system of vehicle storage and release, and achieves good results in the method by greedily improving the throughput of the traffic network. However, the traffic signal control algorithm based on reinforcement learning focuses on the diverse combination of traffic states, ignoring the most basic traffic state representation, and only slightly reduces traffic delay while greatly increasing the complexity of the model, while reducing the usability of the model. Summary of the invention
[0004] In view of the above-mentioned shortcomings of the prior art, the present invention provides a traffic signal control method, device, electronic device and storage medium to solve the technical problems that the above-mentioned model is complex and the effect of reducing traffic delay is not good.
[0005] A traffic signal control method provided by the present invention comprises: obtaining traffic status information, wherein the traffic status information comprises signal light information of an intersection and vehicle information of different travel directions, wherein the vehicle information comprises vehicle queue length, vehicle position and speed; mapping the position and speed of a moving vehicle to the vehicle queue length to obtain the weight of an effective moving vehicle, and calculating the phase effective pressure of different phases of an intersection based on the weight of the effective moving vehicle and the vehicle queue length, wherein the vehicle comprises the moving vehicle, and the phase comprises a group of non-conflicting travel directions; establishing a traffic signal control model according to the traffic status information and the phase effective pressure, and inputting new traffic status information into the traffic signal control model for training; and controlling the traffic signal of a target intersection based on the trained traffic signal control model.
[0006] In an embodiment of the present invention, the position and speed of a new moving vehicle are mapped to a new vehicle queue length to obtain the weight of the new effective moving vehicle; based on the weight of the new effective moving vehicle and the new vehicle queue length, the new phase effective pressure of different phases at an intersection is calculated; a decision is made according to the new signal light information and the new phase effective pressure of different phases to obtain an optimal phase timing plan.
[0007] In an embodiment of the present invention, for each traffic direction, based on the current phase duration, a preset road speed threshold, and the total length of the upstream lane, the farthest effective position of the upstream lane is determined, and the effective driving distance of the upstream lane is calculated according to the farthest effective position and the upstream congestion length of the upstream lane. The signal light information includes the current phase duration, and the traffic state information further includes the total length of the upstream lanes in different traffic directions. The upstream congestion length is obtained based on the positions of the queuing vehicles, and the vehicles further include the queuing vehicles; the farthest effective position is compared with the position of the moving vehicle, and the effective moving vehicle is determined according to the comparison result, and the weight of the effective moving vehicle in the traffic direction is calculated according to the effective driving distance and the speed of the effective moving vehicle.
[0008] In an embodiment of the present invention, if the lane saturation of the traffic direction is greater than or equal to a preset saturation threshold, the traffic moving pressure of the traffic direction is calculated according to the upstream vehicle queue length and the downstream vehicle queue length of the traffic direction; if the lane saturation of the traffic direction is less than the preset saturation threshold, the traffic moving pressure of the traffic direction is calculated according to the upstream vehicle queue length of the traffic direction to obtain the traffic moving pressures of different traffic directions. The vehicle queue length includes the upstream vehicle queue length and the downstream vehicle queue length, and the vehicle information further includes the lane saturation; for each phase, the sum of the traffic moving pressures of each traffic direction of the phase is used as the phase queue pressure of the phase, and the phase effective pressure of the phase is calculated based on the sum of the weights of the effective moving vehicles in each traffic direction of the phase and the phase queue pressure of the phase.
[0009] In an embodiment of the present invention, the sum of the upstream vehicle queue lengths of each traffic direction of the phase is used as the phase queue length of the phase to obtain the phase queue lengths of different phases; based on a preset weight parameter, the phase queue length of the phase, and the phase waiting time of the phase, the reward value of the phase is determined to obtain the reward values of different phases. The signal light information includes the phase waiting times of different phases, and the preset weight parameter increases as the phase waiting time increases; the traffic signal control model is converged based on the reward values of different phases.
[0010] In an embodiment of the present invention, a plurality of initial phase timing schemes are determined according to the new phase waiting time of different phases and the new phase effective pressure of different phases. Each initial phase timing scheme includes a phase duration, a probability, and a set of phase actions. If the phase duration in an initial phase timing scheme satisfies a preset time interval, the initial phase timing scheme is used as a candidate phase timing scheme. The probabilities of each candidate phase timing scheme are compared, and the candidate phase timing scheme corresponding to the maximum probability is used as the preferred phase timing scheme.
[0011] In an embodiment of the present invention, the number of training times for training the traffic signal control model is counted. If the number of training times is equal to a preset threshold, the traffic signal control model is determined as the trained traffic signal control model.
[0012] In an embodiment of the present invention, a traffic signal control device is further provided, including: an acquisition module, configured to acquire traffic state information, where the traffic state information includes signal lamp information of an intersection and vehicle information in different passing directions, and the vehicle information includes a vehicle queue length, a position and a speed of a vehicle; a processing module, configured to map the position and speed of a traveling vehicle to the vehicle queue length to obtain a weight of an effectively traveling vehicle, and calculate a phase effective pressure of different phases of the intersection based on the weight of the effectively traveling vehicle and the vehicle queue length, where the vehicle includes the traveling vehicle, and the phase includes a set of non-conflicting passing directions; a training module, configured to establish a traffic signal control model according to the traffic state information and the phase effective pressure, and input new traffic state information into the traffic signal control model for training; and a control module, configured to control traffic signals of a target intersection based on the trained traffic signal control model.
[0013] In an embodiment of the present invention, an electronic device is further provided. The electronic device includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the traffic signal control method as described above.
[0014] In an embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is enabled to execute the traffic signal control method as described above.
[0015] Advantages of the present invention: The present invention provides a traffic signal control method, apparatus, electronic device, and storage medium. The traffic signal control method establishes a traffic signal control model based on basic traffic state information and trains it. Based on the trained traffic signal control model, the signal phase timing scheme is optimized to control the traffic signals at the target intersection, which can effectively reduce traffic pressure, reduce vehicle waiting time, improve traffic efficiency, alleviate traffic congestion, and at the same time, the model is simple and highly usable. Description of the Drawings
[0016] Figure 1 FIG. is a schematic diagram of the implementation environment of a traffic signal control method shown in an exemplary embodiment of the present invention;
[0017] Figure 2 FIG. is a flowchart of a traffic signal control method shown in an exemplary embodiment of the present invention;
[0018] Figure 3 FIG. is a schematic diagram of a brief intersection shown in an exemplary embodiment of the present invention;
[0019] Figure 4 FIG. is a schematic diagram of a brief traffic movement shown in an exemplary embodiment of the present invention;
[0020] Figure 5 FIG. is a schematic diagram of a four-phase shown in an exemplary embodiment of the present invention;
[0021] Figure 6 FIG. is a schematic diagram of an eight-phase shown in an exemplary embodiment of the present invention;
[0022] Figure 7 FIG. is a schematic diagram of the traffic conditions at intersection γ shown in an exemplary embodiment of the present invention;
[0023] Figure 8 FIG. is a schematic diagram of the training process of a traffic signal control model shown in an exemplary embodiment of the present invention;
[0024] Figure 9 FIG. is a schematic diagram of the Jinan road network simulation shown in an exemplary embodiment of the present invention;
[0025] Figure 10 FIG. is a block diagram of a traffic signal control apparatus shown in an exemplary embodiment of the present invention. Detailed Embodiments
[0026] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0027] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0028] It should be noted that in the present invention, "first", "second", etc. are only used to distinguish similar objects, and are not intended to limit the order or sequence of similar objects. The described "including", "having", etc. are deformations, indicating that the scope covered by the subject of the word excludes the examples shown by the word and is not exclusive.
[0029] It can be understood that the various numerical numbers, step numbers, etc. recorded in the present invention are for the convenience of description and are not used to limit the scope of the present invention. The size of the reference numbers in the present invention does not mean the order of execution. The execution order of each process should be determined by its function and internal logic.
[0030] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0031] It should be noted that existing adaptive traffic signal control usually models traffic movement as a queuing system for vehicle storage and release, and has achieved good results in methods by greedily improving the throughput of the traffic flow network. These methods model traffic signal control as an agent that observes various traffic-related features, such as queue length, vehicle speed, average waiting time, etc., and optimizes its scheme based on the rewards feedback from the traffic environment after phase actions selection (i.e., the change of traffic lights), such as the number of vehicles and vehicle passing rate, to learn how to take the next action. For example, six states are used for representation, including queue length, number of vehicles, current phase, next phase, vehicle image, and updated waiting time, and six rewards, including queue length, delay, total updated waiting time, light change indicator, number of vehicles passing through, and total travel time. Another example is to use a simpler state representation, including the current phase and an image-like representation, but requires complex rewards, including delay, emergency stop, light change indicator, and vehicle waiting time.
[0032] However, traffic signal control algorithms based on reinforcement learning focus on the diverse combination usage of traffic states, ignoring the most basic traffic state representation. While slightly reducing traffic delay, they instead significantly increase model complexity and reduce usability. At the same time, the main purpose of previous solutions is to maximize the traffic capacity of intersections, while ignoring the longest red light time for a single lane, leading to aggressive driving behavior of drivers. In addition, although the maximum pressure method is highly representative in road pressure representation, none of them can vividly express the impact of vehicles traveling in the traffic network on phase adjustment and are difficult to represent complex traffic states.
[0033] To solve the above problems, embodiments of the present invention respectively propose a traffic signal control method, a traffic signal control device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail below.
[0034] It should be noted that the traffic signal control method proposed in the embodiments of the present invention is an ETS-RL algorithm (Excellent Traffic State Reinforcement Learning, a more superior traffic state representation algorithm based on reinforcement learning). It introduces the traveling vehicles and road saturation into phase competition, designs a new traffic state (Excellent Traffic State, abbreviated as ETS) representation by coordinating the effective pressure upstream and downstream of the intersection and the effective traveling vehicles within a defined range, and calculates the corresponding phase sequence probability and selects the maximum one for output by presetting the intersection traffic state division rule - vehicle driving routes (traffic directions) that do not conflict with each other can be divided into the same phase.
[0035] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the implementation environment of a traffic signal control method shown in an exemplary embodiment of the present invention.
[0036] As Figure 1 shown, the implementation environment may include a traffic state perception terminal 101, a computer device 102, and a signal lamp control terminal 103. Among them, the computer device 102 may be at least one of a microcomputer, an embedded computer, a neural network computer, etc. The traffic state perception terminal 101 is used to collect traffic state information and provide it to the computer device 102. The computer device 102 is used to establish a traffic signal control model based on the traffic state information and train it, optimize the signal phase timing plan based on the trained traffic signal control model, and send the phase timing plan to the signal lamp control terminal 103 to control the traffic signals at the target intersection.
[0037] First, the traffic state perception terminal 101 composed of multiple sensors collects real-time traffic state information, including vehicle position and speed, and lane queue length, and statistically processes it with the lane ID (Identity document) as the identifier. Second, at a specific moment, the computer device 102 initiates a request to obtain real-time traffic state information. Then, the ETS-RL algorithm calculates the corresponding phase sequence probability according to the preset intersection traffic state division rules and selects the maximum output, which can achieve better traffic signal adaptive control. Finally, on the premise of ensuring safety, the signal lamp control terminal 103 executes signal control according to the intersection phase timing plan. The traffic signal control model adjusts the phase and optimizes the timing by learning the phase rules of traffic signal lamps, aiming to reduce the waiting time of vehicles and effectively improve the passing capacity of intersections, and relieve urban traffic congestion. It has feasibility and certain advantages both in terms of the difficulty of scale-up and the improvement of urban traffic, and has good practical value for building a smart city, so it has a good development prospect.
[0038] Exemplarily, the computer device 102 obtains traffic state information, where the traffic state information includes traffic signal information at intersections and vehicle information in different driving directions. The vehicle information includes the vehicle queue length, the position and speed of the vehicle. The position and speed of the driving vehicle are mapped to the vehicle queue length to obtain the weight of the effective driving vehicle. Based on the weight of the effective driving vehicle and the vehicle queue length, the phase effective pressure of different phases at the intersection is calculated. The vehicle includes the driving vehicle, and the phase includes a group of non-conflicting driving directions. A traffic signal control model is established according to the traffic state information and the phase effective pressure, and the new traffic state information is input into the traffic signal control model for training. The traffic signals at the target intersection are controlled based on the trained traffic signal control model. It can be seen that the technical solution of the embodiment of the present invention can establish a traffic signal control model through basic traffic state information and perform training, optimize the signal phase timing scheme based on the trained traffic signal control model to control the traffic signals at the target intersection, effectively reduce traffic pressure, reduce vehicle waiting time, improve traffic efficiency, relieve traffic congestion, and at the same time, the model is simple and highly usable.
[0039] It should be noted that the traffic signal control method provided by the embodiment of the present invention is generally executed by the traffic state perception end 101, the computer device 102, and the signal lamp control end 103. The traffic signal control device is generally set in the computer device 102.
[0040] Please refer to Figure 2 , Figure 2 is a flowchart of a traffic signal control method shown in an exemplary embodiment of the present invention. This method can be applied to Figure 1 the shown implementation environment and is specifically executed by the traffic state perception end 101, the computer device 102, and the signal lamp control end 103 in this implementation environment. It should be understood that this method can also be applicable to other exemplary implementation environments and be specifically executed by devices in other implementation environments. The embodiment does not limit the implementation environment applicable to this method.
[0041] As Figure 2 shown, in an exemplary embodiment, the traffic signal control method at least includes steps S210 to S240, which are introduced in detail as follows:
[0042] Step S210, obtain traffic state information.
[0043] In one embodiment of the present invention, traffic state information can be obtained through at least one of a simulation simulator, a traffic state perception terminal, a shared experience pool, etc., without limitation here. Among them, the traffic state information includes signal light information at intersections and vehicle information in the passing direction. The vehicle information includes the vehicle queue length, the position and speed of the vehicle. The vehicle queue length refers to the number of currently queuing vehicles, which can be greater than or equal to 0. Vehicles include moving vehicles and queuing vehicles. The signal light information includes the phase waiting time and the current phase duration. The phase waiting time refers to the time interval between the signal phase (abbreviated as phase) corresponding to each passing direction and the last time it obtained the green light right of way, that is, the red light duration of the phase. The current phase duration is the green light duration of the currently obtained right of way phase.
[0044] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an intersection shown in an exemplary embodiment of the present invention. As Figure 3 shown, a traffic intersection (intersection) is composed of several groups of incoming roads (L in ) and corresponding outgoing roads (L out ) that intersect or cross each other, denoted by the symbol I. Each road is composed of several lanes that determine the driving path of the lane and are the basic components in the road network. Each traffic network is composed of multiple intersections (I1... I N ) connected to each other through a group of roads (R1... R M ), where N represents the total number of traffic intersections and M represents the total number of roads. The reasonable driving trajectory of a vehicle from the incoming lane (upstream lane) l through the intersection to the outgoing lane (downstream lane) m is called a traffic movement (passing direction), denoted as (l, m).
[0045] Please refer to Figure 4 , Figure 4 which is a schematic diagram of a traffic movement shown in an exemplary embodiment of the present invention. As Figure 4 shown, a 4-way intersection includes 4 "left turns", 4 "going straight", and 4 "right turns", a total of 12 traffic movement methods. According to the traffic rules of most intersections, vehicles can turn right regardless of the signal. Therefore, usually only 8 traffic movements need to be considered for coordination, such as Figure 4 the 1#-8# traffic movements in
[0046] Please refer to Figure 5 and Figure 6 , Figure 5 which is a schematic diagram of a four-phase shown in an exemplary embodiment of the present invention, Figure 6 and Figure 5 which is a schematic diagram of an eight-phase shown in an exemplary embodiment of the present invention. As Figure 6 shown,Figure 5 Four groups of commonly used traffic movement matching schemes are described, which respectively constitute four signal phases A, B, C, and D (abbreviated as phases). Figure 6 Eight groups of commonly used traffic movement matching schemes are described, which respectively constitute eight signal phases A, B, C, D, E, F, G, and H. And each signal phase contains a group of non-conflicting traffic movements. The phase is usually represented by s.
[0047] Step S220: Map the position and speed of the moving vehicle to the vehicle queue length to obtain the weight of the effective moving vehicle, and calculate the phase effective pressure of different phases at the intersection based on the weight of the effective moving vehicle and the vehicle queue length.
[0048] In an embodiment of the present invention, for the lanes in each traffic direction, since moving vehicles may continuously join the current vehicle queue length, the directly obtained vehicle queue length cannot truly reflect the vehicle queuing situation of the lane to be passed. Correspondingly, the accuracy of the phase pressure calculated only based on the vehicle queue length is not high. Therefore, it is necessary to determine the vehicle queuing situation that can truly reflect the traffic direction. It can be judged which moving vehicles can merge into the vehicle queue length before the end of the current phase duration of the current green light phase according to the position and speed of the moving vehicle, and then map the position and speed of the moving vehicle to the vehicle queue length to obtain the weight of the effective moving vehicle. Calculate the phase effective pressure of different phases at the intersection based on the weight of the effective moving vehicle and the vehicle queue length, which improves the accuracy of the phase pressure. Wherein, the vehicle includes the moving vehicle. Correspondingly, the position and speed of the vehicle include the position and speed of the moving vehicle, and each phase includes a group of non-conflicting traffic directions.
[0049] In an embodiment of the present invention, mapping the position and speed of the moving vehicle to the vehicle queue length to obtain the weight of the effective moving vehicle includes the following:
[0050] For each traffic direction, determine the farthest effective position of the upstream lane based on the current phase duration, the preset road speed threshold, and the total length of the upstream lane. Calculate the effective driving distance of the upstream lane according to the farthest effective position and the upstream blockage length of the upstream lane. The signal lamp information includes the current phase duration, and the traffic state information also includes the total length of the upstream lanes in different traffic directions. The upstream blockage length is obtained based on the position of the queuing vehicle. The vehicle also includes the queuing vehicle.
[0051] Compare the farthest effective position with the position of the moving vehicle, determine the effective moving vehicle according to the comparison result, and calculate the weight of the effective moving vehicle in the traffic direction according to the effective driving distance and the speed of the effective moving vehicle.
[0052] In this embodiment, the number of sub-lanes in the upstream lane in each traffic direction is greater than or equal to 1, and the number of effective driving vehicles in each upstream lane is greater than or equal to 0. Taking the weight of an effective driving vehicle in the upstream lane l of traffic movement (l, m) as an example, the calculation process of the weight of the effective driving vehicle is as follows:
[0053] 1) Determine the effective range (the farthest effective position). This effective range refers to the farthest effective position where a driving vehicle can pass through the intersection within the current phase duration, and the farthest effective position cannot exceed the total length of the upstream lane. First, calculate the initial value of the farthest effective position through the current phase duration and the preset road speed threshold, and the calculation method is as follows:
[0054] L = V max × t duration Equation (1)
[0055] where L is the initial value of the farthest effective position, V max is the maximum speed allowed on the road (preset road speed threshold), and t duration is the current phase duration.
[0056] Then, compare the initial value of the farthest effective position with the total length of the upstream lane. If the initial value of the farthest effective position is less than or equal to the total length of the upstream lane, take the initial value of the farthest effective position as the farthest effective position. If the initial value of the farthest effective position is greater than the total length of the upstream lane, take the total length of the upstream lane as the farthest effective position.
[0057] 2) Determine the effective driving distance. The effective driving distance refers to the effective driving distance of a driving vehicle in the upstream lane for each traffic direction. The calculation method of the effective driving distance of the upstream lane l is as follows:
[0058] L surplus = L - X(l) - spaceHeadway Equation (2)
[0059] where L surplus is the effective driving distance on the upstream lane l of traffic movement (l, m), X(l) is the upstream congestion length of the upstream lane l of traffic movement (l, m), and spaceHeadway is the preset headway. The preset headway can be 2.5m, or 3m, or other lengths set by those skilled in the art. If the number of sub-lanes in the upstream lane l is greater than 1, the number of upstream congestion lengths is also greater than 1, and the effective driving distance of each sub-lane of the upstream lane l needs to be calculated separately.
[0060] Among them, the upstream blockage length can be obtained according to the position of the last vehicle in the queuing vehicles in the upstream lane. Since the vehicles also include queuing vehicles, correspondingly, the position of the vehicles also includes the position of the queuing vehicles. When there are no queuing vehicles, the upstream blockage length is 0, and at this time, the farthest valid position is used as the effective driving distance.
[0061] 3) Determine the weight of the effectively driving vehicles. Effectively driving vehicles refer to the driving vehicles at the entrance (passing through the intersection) within the effective range of the intersection. By using traffic state information such as vehicle position, speed, and upstream blockage length, it is estimated whether the driving vehicle can be in a stopped or driving state within the future phase duration. The estimated value represents the influence degree of the phase switch on the driving vehicle, that is, the weight of the effectively driving vehicle. Compare the farthest valid position with the positions of the driving vehicles in the upstream lane. If the position of a driving vehicle does not exceed the farthest valid position, then this driving vehicle is determined as an effectively driving vehicle. Calculate the weight of this effectively driving vehicle according to the effective driving distance of the upstream lane l of the traffic movement (l, m) and the speed of an effectively driving vehicle in this upstream lane l. The calculation method is as follows:
[0062]
[0063] Among them, r(l, m) is the weight of an effectively driving vehicle in the upstream lane l of the traffic movement (l, m), and v is the speed of this effectively driving vehicle. If the number of effectively driving vehicles of the traffic movement (l, m) is multiple, the weights of all effectively driving vehicles of the traffic movement (l, m) can be calculated by formula (3).
[0064] And so on, the weights of all effectively driving vehicles of the traffic movement (k, v) can also be determined by the calculation methods of formula (1), formula (2), and formula (3).
[0065] In an embodiment of the present invention, calculating the phase effective pressure of different phases of the intersection based on the weight of the effectively driving vehicles and the vehicle queue length includes the following steps:
[0066] Step S221, if the lane saturation degree of the passing direction is greater than or equal to the preset saturation threshold, then calculate the traffic movement pressure of the passing direction according to the upstream vehicle queue length and the downstream vehicle queue length of the passing direction. If the lane saturation degree of the passing direction is less than the preset saturation threshold, then calculate the traffic movement pressure of the passing direction according to the upstream vehicle queue length of the passing direction, and obtain the traffic movement pressures of different passing directions.
[0067] In one embodiment of the present invention, the vehicle information further includes lane saturation, which refers to the lane saturation of the downstream lane corresponding to each traffic direction and can be obtained by collecting the current traffic volume of the downstream lane corresponding to each traffic direction and preprocessing the current traffic volume. Exemplarily, for the current stage, Q is used m to represent the lane saturation of the downstream lane m of the traffic movement (l,m), and the calculation method is as follows:
[0068]
[0069] where C now is the current traffic volume of the downstream lane m, and C max is the preset maximum traffic volume of the downstream lane m.
[0070] The lane saturations of different downstream lanes are different, and the delay congestion phenomena caused by the upcoming upstream vehicles are also different. Therefore, a conditional function needs to be used for classification processing when calculating the traffic pressure. The lane saturation is used as a judgment of the possible vehicle congestion degree. The lane saturation reflects the lane service level, while the phase queue pressure represents the vehicle passing needs.
[0071] Please refer to Figure 7 , Figure 7 which is a schematic diagram of the traffic situation at intersection γ shown in an exemplary embodiment of the present invention. As Figure 7 shown, the north-south straight phase is released, but the east-west traffic demand is stronger. Obviously, it is undoubtedly unreasonable to calculate the phase queue pressure only using the upstream and downstream vehicle queue lengths.
[0072] In this embodiment, the vehicle queue length includes the upstream vehicle queue length and the downstream vehicle queue length. Before calculating the traffic movement pressure, it is necessary to compare the preset saturation threshold with the lane saturation of each traffic direction, and determine the calculation method of the traffic movement pressure according to the comparison result. The calculation method of the traffic movement pressure is as follows:
[0073]
[0074] where p q (l,m) is the traffic movement pressure on the traffic movement (l,m) formed by the upstream lane l and the downstream lane m, x(l i ) is the upstream vehicle queue length of the sub-lane l i of the upstream lane l, M is the number of sub-lanes of the upstream lane l, x(m j ) is the downstream vehicle queue length of the sub-lane m j of the downstream lane, N is the number of sub-lanes of the downstream lane m, and Q mis the lane saturation of the downstream lane m, and W1 is the preset saturation threshold. It should be noted that the upstream lane or the downstream lane is a general term, which does not solely represent one lane. The upstream lane or the downstream lane includes at least one sub-lane.
[0075] Step S222: For each phase, take the sum of the traffic movement pressures in all passing directions of the phase as the phase queue pressure of the phase, and calculate the phase effective pressure of the phase based on the sum of the weights of the effective driving vehicles in all passing directions of the phase and the phase queue pressure of the phase.
[0076] In an embodiment of the present invention, after calculating the traffic movement pressures in different passing directions, take the sum of the traffic movement pressures in all non-conflicting passing directions in a phase as the phase queue pressure of the phase to obtain the phase queue pressures of different phases. The calculation method of the phase queue pressure is as follows:
[0077] p q (s) = p q (l,m) + p q (k,v) Equation (6)
[0078] Wherein, p q (s) is the phase queue pressure of phase s. Phase s includes a group of non-conflicting traffic movements (l,m) and traffic movement (k,v), p q (k,v) is the traffic movement pressure on the traffic movement (k,v) composed of the upstream lane k and the downstream lane v, and p q (k,v) is calculated in the same way as the above p q (l,m), and will not be elaborated here.
[0079] Then, calculate the phase effective pressures of different phases. The phase effective pressure of each phase is the sum of the phase queue pressure of the phase and the sum of the weights of all effective driving vehicles in all passing directions of the phase. For example, the calculation method of the phase effective pressure of phase s is as follows:
[0080] d(s) = ∑r(l,m) + ∑r(k,v) + p q (s) Equation (7)
[0081] Wherein, d(s) is the phase effective pressure of phase s, ∑r(l,m) is the sum of the weights of all effective driving vehicles in traffic movement (l,m), and ∑r(k,v) is the sum of the weights of all effective driving vehicles in traffic movement (k,v).
[0082] Similarly, through the above calculation method, the phase effective pressures of other phases can be calculated.
[0083] Step S230: Establish a traffic signal control model based on traffic state information and phase effective pressure, and input the new traffic state information into the traffic signal control model for training.
[0084] In an embodiment of the present invention, establishing a traffic signal control model based on traffic state information and phase effective pressure includes defining a state space and presetting a traffic signal phase action space (abbreviated as phase action space). Among them, the traffic signal control model can be at least one of a deep learning network model, a convolutional neural network model, and a fully connected network model. Obtain new traffic state information, and use the new traffic state information as an input value to train the traffic signal control model to obtain a trained traffic signal control model.
[0085] It should be noted that the flexibility of the phase action space has an obvious impact on the performance of the traffic signal control model. The phase action space design of the present invention mainly considers two situations. First, the signal phases are combined in pairs on the premise of lane turning and non-conflict. Based on real-time traffic flow information (traffic state information), the signal light can jump to any green light phase, and the right-turn direction is set to always green. The phase action space can be expressed as Figure 5 the four-phase in Figure 6 and the eight-phase in
[0086] In addition, an input interface can also be defined for the traffic signal control model so that the input interface converts the traffic state information into a state matrix.
[0087] In an embodiment of the present invention, inputting the new traffic state information into the traffic signal control model for training includes the following:
[0088] Map the position and speed of the new traveling vehicles to the new vehicle queue length to obtain the weight of the new effective traveling vehicles;
[0089] Calculate the new phase effective pressure of different phases at the intersection based on the weight of the new effective traveling vehicles and the new vehicle queue length;
[0090] Make a decision based on the new signal light information and the new phase effective pressure of different phases to obtain an optimal phase timing plan.
[0091] In this embodiment, the traffic signal control model learns to process traffic state information, including mapping the positions and speeds of new moving vehicles to new vehicle queue lengths, obtaining the weights of new effective moving vehicles, and calculating the new traffic pressure at the intersection based on the weights of new effective moving vehicles and new vehicle queue lengths. The traffic signal control model also learns to make decisions based on new signal light information and new phase effective pressures, and outputs the optimal phase timing plan for the next moment.
[0092] Schematically, a simulation simulator can be configured to obtain the currently simulated traffic state information through the simulation simulator, input the currently simulated traffic state information into the traffic signal control model, so that the traffic signal control model outputs the optimal phase timing plan for the next moment, control the simulation simulator to execute the optimal phase timing plan and obtain the new simulated traffic state information, and then input the new simulated traffic state information into the traffic signal control model for training.
[0093] In an embodiment of the present invention, making a decision based on new signal light information and new phase effective pressures of different phases to obtain an optimal phase timing plan includes the following:
[0094] Determine a plurality of initial phase timing plans according to the new phase waiting times of different phases and the new phase effective pressures of different phases. Each initial phase timing plan includes a phase duration, a probability, and a set of phase actions;
[0095] If the phase duration in an initial phase timing plan satisfies a preset time interval, then use the initial phase timing plan as a candidate phase timing plan;
[0096] Compare the probabilities of each candidate phase timing plan, and use the candidate phase timing plan corresponding to the maximum probability as the optimal phase timing plan.
[0097] In this embodiment, the traffic signal control model determines the phase duration and probability of each set of phase actions in the phase action space according to the new phase waiting time and new phase effective pressure of each phase, obtains a plurality of initial phase timing plans, uses the initial phase timing plan whose phase duration of the phase action satisfies the preset time interval as a candidate phase timing plan, and uses the candidate phase timing plan with the maximum probability of the phase action as the optimal phase timing plan and outputs it.
[0098] By stipulating the minimum green light time and the maximum green light time to limit the adopted action plan, it can prevent the situation that the green light time of a single lane is too long and the other lanes cannot bear it, and can ensure the driving safety at the intersection.
[0099] In another embodiment of the present invention, inputting the new traffic state information into the traffic signal control model for training further includes the following:
[0100] The sum of the lengths of the upstream vehicle queues in each traffic direction of a phase is used as the phase queue length of the phase, and the phase queue lengths of different phases are obtained;
[0101] Based on a preset weight parameter, the phase queue length of a phase, and the phase waiting time of the phase, the reward value of the phase is determined, and the reward values of different phases are obtained. The signal lamp information includes the phase waiting times of different phases, and the preset weight parameter increases as the phase waiting time increases;
[0102] Based on the reward values of different phases, the traffic signal control model is converged.
[0103] In this embodiment, during the reinforcement learning process, the reward function can provide a learning direction for the traffic signal control model and determine the convergence speed of the traffic signal control model. For the definition of the reward function, the present invention mainly considers from two directions: regarding the phase queue length as the delay time and regarding the phase waiting time as a competition term. First, the delay time of a vehicle can be approximated as the phase queue length, which can also reflect the traffic demand of the road. Second, in order to balance the traffic flows in all directions and prevent a phase from falling into a long waiting state, the red light duration (phase waiting time) of a phase is used as a competition term. The reward function can be defined by the following formula:
[0104]
[0105] where R i is the reward value of phase i, q j is the length of the upstream lane queue in traffic direction j of phase i, W waiting is the red light duration after the end of the last green light of phase i, that is, the phase waiting time, and α is a preset weight coefficient that increases as the phase waiting time increases, indicating that the lane with a longer waiting time has a higher priority.
[0106] The Bellman equation is used for model update, which is expressed as follows:
[0107] Q(s t ,a t ) = R(s t ,a t ) + γmaxQ(s t+1 ,a t+1 ) Equation (9)
[0108] where Q(s t ,a t ) is the action value under the optimal policy at the current time t, s is a finite state set, and a is a finite action set. R(s t ,a t ) is the reward obtained in the state s at the current time t tUnder the following circumstances, take action a t The obtained reward value, maxQ(s t+1 , a t+1 ) is the expectation of the future action value that can be obtained by acting according to the optimal strategy, and γ is a discount relationship of the future action value.
[0109] The parameters of the traffic signal control model are continuously adjusted through the reinforcement learning algorithm, and the finally output optimal phase timing plan can regulate the traffic flow to the greatest extent and improve traffic efficiency.
[0110] In another embodiment of the present invention, when new traffic state information is input into the traffic signal control model for training, the following is further included:
[0111] Count the number of training times for training the traffic signal control model. If the number of training times is equal to the preset threshold, determine the traffic signal control model as the trained traffic signal control model.
[0112] Illustratively, the preset threshold can be 100, or other values set by those skilled in the art.
[0113] Step S240, control the traffic signals at the target intersection based on the trained traffic signal control model.
[0114] In one embodiment of the present invention, obtain the current traffic state information, where the current traffic state information includes the current signal light information at the target intersection and the current vehicle information in the passing direction. The current vehicle information includes the current vehicle queue length, the position and speed of the current vehicle. Input the current traffic state information into the trained traffic signal control model to obtain the optimal phase timing plan for the next moment, and send the optimal phase timing plan to the signal light control end corresponding to the target intersection, so that the signal light control end controls the traffic signals at the target intersection according to the optimal phase timing plan.
[0115] Generally speaking, the technical solution of the embodiment of the present invention maps the vehicles in motion and the traffic signal states to the waiting queue to represent the overall traffic demand of the road in the recent period. At the same time, different downstream lanes have a delay and congestion phenomenon for the upcoming upstream vehicles due to their different lane saturations. Therefore, a conditional function needs to be used for classification processing when calculating the phase pressure. Finally, this traffic state representation method is combined with reinforcement learning to develop an algorithm template based on reinforcement learning, and the phase adjustment and timing optimization are learned through environmental feedback, so as to perform better.
[0116] Please refer to Figure 8 , Figure 8 is a schematic diagram of the training process of a traffic signal control model shown in an exemplary embodiment of the present invention. As shown in the figure, the training process is as follows:
[0117] 1) Configure a simulation emulator, build a reinforcement learning network, and define a multi-intersection control model
[0118] The present invention is based on the Windows (an operating system) system, and uses the traffic simulation software SUMO (a simulation emulator) as a test platform. Configure the simulation intersection environment and traffic flow data through SUMO, and extract simulation data and traffic signal control through the API interface (Application Programming Interface) and the TraCI interface (Traffic Control Interface). Define the action space, define the reward function, build a reinforcement learning network, use the vehicle speed, vehicle position, vehicle queue length, and current phase waiting time at each intersection, etc. as the traffic state representation, and combine the signal light information as the input parameters of the model. According to the data characteristics, use a convolutional network and a fully connected network for feature extraction. Regard each intersection as an agent, and the agent selects and executes the next action to maximize the expected reward, and adjusts its own strategy according to the environmental feedback.
[0119] The present invention uses a total of 5 real-world traffic data sets from Jinan and Hangzhou to configure the simulation road files and traffic flow files to describe the traffic road network and vehicle status. Among them, 3 are from Jinan and 2 are from Hangzhou. Please refer to Figure 9 , Figure 9 which is a schematic diagram of the Jinan road network simulation shown in an exemplary embodiment of the present invention. As Figure 6 shown, the road network of the Jinan data set has 12 crossroads (3×4). Each crossroad is a four-way crossroad, with two 400-meter (east-west) long sections and two 800-meter (north-south) long sections. The road network of the Hangzhou data set has 16 crossroads (4×4). Each crossroad is a four-way crossroad, with two 800-meter (east-west) long sections and two 600-meter (north-south) long sections. The maximum allowable speed of all lanes is 40 km / h.
[0120]
[0121] Table 1
[0122] Please refer to Table 1. Table 1 is the vehicle arrival rate table of the data set in a specific embodiment of the present invention. As shown in Table 1, these data sets have different vehicle arrival rates and can simulate traffic conditions in different situations, which are sufficient to meet the experimental requirements.
[0123] 2) Obtain the intersection traffic state information, and generate the signal timing plan (preferably the phase timing plan) for the next moment based on the control model
[0124] During the experiment, traffic state information such as the position and speed of vehicles, the length of vehicle queues, and signal light information is obtained from SUMO in real time through the TraCI interface. After processing all vehicle state information, it is converted into a matrix as the input of the convolutional network. Finally, the signal timing plan for the next moment is output, and the signal timing plan includes the probability values of a set of action spaces (phase actions) and the green light duration of the phase (phase duration).
[0125] 3) The simulation simulator executes the timing plan and obtains a new traffic state
[0126] The agent executes the selected action, updates the traffic state, and enters the next state according to the traffic state information of the simulation environment obtained from the simulation simulator. Collect the traffic state information of the intersection in the simulation simulator and continuously control and update. At the same time, store the historical data (historical traffic state information) in the shared experience pool to accelerate the training speed and update the model parameters in a timely manner. At the same time, count the number of training times and determine whether the preset number of training times (preset threshold) is reached. If the preset number of training times is reached, the final control model, that is, the trained traffic signal control model, is output, otherwise continue training.
[0127] The traffic signal control method provided by the present invention is a signal timing optimization method based on traffic state representation. Through the underlying traffic state representation, it can effectively reduce traffic pressure, reduce vehicle waiting time, improve traffic efficiency, and alleviate traffic congestion. By considering the connection between queuing vehicles and moving vehicles, a traffic signal control method based on the maximum pressure algorithm - ETS is designed, which can be flexibly applied to different adaptive traffic signal control models. Further experiments prove that integrating ETS with a reinforcement learning-based method can bring better model effects.
[0128] Please refer to Figure 10 , Figure 10 which is a block diagram of a traffic signal control device shown in an exemplary embodiment of the present invention. This device can be applied to Figure 1 the implementation environment shown, and is specifically configured in the computer device 102. This device can also be applicable to other exemplary implementation environments and is specifically configured in other devices. This embodiment does not limit the implementation environment applicable to this device.
[0129] As Figure 10 shown, this exemplary traffic signal control device includes:
[0130] An acquisition module 1010 is configured to acquire traffic state information, where the traffic state information includes signal light information at intersections and vehicle information in different driving directions. The vehicle information includes the vehicle queue length, the position and speed of the vehicle. A processing module 1020 is configured to map the position and speed of the driving vehicle to the vehicle queue length to obtain the weight of the effective driving vehicle, and calculate the phase effective pressure of different phases at the intersection based on the weight of the effective driving vehicle and the vehicle queue length. The vehicle includes the driving vehicle, and the phase includes a set of non-conflicting driving directions. A training module 1030 is configured to establish a traffic signal control model according to the traffic state information and the phase effective pressure, and input new traffic state information into the traffic signal control model for training. A control module 1040 is configured to control the traffic signal at the target intersection based on the trained traffic signal control model.
[0131] It should be noted that the traffic signal control device provided in the above embodiment and the traffic signal control method provided in the above embodiment belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment, and will not be elaborated here. In practical applications, the traffic signal control device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.
[0132] This embodiment also provides an electronic device, including: one or more processors; a storage device configured to store one or more programs. When the one or more programs are executed by the one or more processors, the electronic device implements the traffic signal control method provided in each of the above embodiments.
[0133] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of the computer, the computer is made to execute the traffic signal control method as described above. The computer-readable storage medium may be included in the electronic device described in the above embodiment, or may exist separately without being assembled into the electronic device.
[0134] This embodiment also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the traffic signal control method provided in each of the above embodiments.
[0135] The electronic device provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication therebetween. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program so that the electronic device executes each step of the above method.
[0136] In this embodiment, the memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0137] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0138] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROM (Read Only Memory), RAM (Random Access Memory), magnetic disk, or optical disk and other media that can store program codes.
[0139] The above embodiments are only used to exemplarily illustrate the principles and effects of the present invention, rather than to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A traffic signal control method, characterized in that, The method includes: Obtaining traffic state information, where the traffic state information includes signal light information at intersections and vehicle information in different traffic directions. The vehicle information includes vehicle queue length, the position and speed of vehicles; Mapping the position and speed of the driving vehicles to the vehicle queue length to obtain the weight of the effective driving vehicles. For each traffic direction, based on the current phase duration, a preset road speed threshold, and the total length of the upstream lane, the farthest effective position of the upstream lane is determined. According to the farthest effective position and the upstream congestion length of the upstream lane, the effective driving distance of the upstream lane is calculated. The signal light information includes the current phase duration. The traffic state information also includes the total length of the upstream lanes in different traffic directions. The upstream congestion length is obtained based on the position of the queuing vehicles, and the vehicles also include the queuing vehicles. Comparing the farthest effective position with the position of the driving vehicles, determining the effective driving vehicles according to the comparison result, and calculating the weight of the effective driving vehicles in the traffic direction according to the effective driving distance and the speed of the effective driving vehicles; Calculating the phase effective pressure of different phases at the intersection based on the weight of the effective driving vehicles and the vehicle queue length. If the lane saturation of a traffic direction is greater than or equal to a preset saturation threshold, calculating the traffic movement pressure of the traffic direction according to the upstream vehicle queue length and the downstream vehicle queue length of the traffic direction. If the lane saturation of the traffic direction is less than the preset saturation threshold, calculating the traffic movement pressure of the traffic direction according to the upstream vehicle queue length of the traffic direction to obtain the traffic movement pressure of different traffic directions. The vehicle queue length includes the upstream vehicle queue length and the downstream vehicle queue length. The vehicle information also includes the lane saturation. For each phase, taking the sum of the traffic movement pressures of each traffic direction of the phase as the phase queue pressure of the phase, and calculating the phase effective pressure of the phase based on the sum of the weights of the effective driving vehicles of each traffic direction of the phase and the phase queue pressure of the phase. The vehicles include the driving vehicles, and the phase includes a set of non-conflicting traffic directions; Establishing a traffic signal control model according to the traffic state information and the phase effective pressure, and inputting new traffic state information into the traffic signal control model for training; Controlling the traffic signals at the target intersection based on the trained traffic signal control model.
2. The traffic signal control method according to claim 1, wherein Inputting new traffic state information into the traffic signal control model for training includes: Mapping the position and speed of the new driving vehicles to the new vehicle queue length to obtain the weight of the new effective driving vehicles; Calculating the new phase effective pressure of different phases at the intersection based on the weight of the new effective driving vehicles and the new vehicle queue length; Making a decision according to the new signal light information and the new phase effective pressure of different phases to obtain an optimal phase timing plan.
3. The traffic signal control method according to claim 1, wherein Inputting new traffic state information into the traffic signal control model for training further includes: Taking the sum of the upstream vehicle queue lengths in each traffic direction of the phase as the phase queue length of the phase, the phase queue lengths of different phases are obtained; Based on preset weight parameters, the phase queue length of the phase, and the phase waiting time of the phase, the reward value of the phase is determined, and the reward values of different phases are obtained. The signal light information includes the phase waiting times of different phases, and the preset weight parameter increases as the phase waiting time increases; Based on the reward values of different phases, the traffic signal control model is converged.
4. The traffic signal control method according to claim 3, wherein Making a decision according to the new signal light information and the new phase effective pressure of different phases to obtain an optimal phase timing plan, including: Determining a plurality of initial phase timing plans according to the new phase waiting times of different phases and the new phase effective pressures of different phases. Each initial phase timing plan includes a phase duration, a probability, and a set of phase actions; If the phase duration in an initial phase timing plan satisfies a preset time interval, then taking the initial phase timing plan as a candidate phase timing plan; Comparing the probabilities of each candidate phase timing plan, and taking the candidate phase timing plan corresponding to the maximum probability as the optimal phase timing plan.
5. The traffic signal control method according to any one of claims 1 to 4, characterized in that Inputting new traffic state information into the traffic signal control model for training, further including: Counting the number of training times for training the traffic signal control model. If the number of training times is equal to a preset threshold, then determining the traffic signal control model as the trained traffic signal control model.
6. A traffic signal control device, characterized in that, The device includes: An acquisition module, configured to acquire traffic state information, where the traffic state information includes signal light information of an intersection and vehicle information in different traffic directions, and the vehicle information includes vehicle queue length, vehicle position, and speed; A processing module, configured to map the position and speed of a moving vehicle to the vehicle queue length to obtain the weight of the effectively moving vehicle. For each traffic direction, based on the current phase duration, a preset road speed threshold, and the total length of the upstream lane, determine the farthest effective position of the upstream lane. Calculate the effective driving distance of the upstream lane according to the farthest effective position and the upstream congestion length of the upstream lane. The signal light information includes the current phase duration. The traffic state information further includes the total lengths of the upstream lanes in different traffic directions. The upstream congestion length is obtained based on the positions of the queuing vehicles, and the vehicle also includes the queuing vehicles. Compare the farthest effective position with the position of the moving vehicle, determine the effectively moving vehicle according to the comparison result, and calculate the weight of the effectively moving vehicle in the traffic direction according to the effective driving distance and the speed of the effectively moving vehicle. Calculate the phase effective pressure of different phases at the intersection based on the weight of the effectively moving vehicle and the vehicle queue length. If the lane saturation degree of a traffic direction is greater than or equal to a preset saturation threshold, calculate the traffic movement pressure of the traffic direction according to the upstream vehicle queue length and the downstream vehicle queue length of the traffic direction. If the lane saturation degree of the traffic direction is less than the preset saturation threshold, calculate the traffic movement pressure of the traffic direction according to the upstream vehicle queue length of the traffic direction to obtain the traffic movement pressures of different traffic directions. The vehicle queue length includes the upstream vehicle queue length and the downstream vehicle queue length, and the vehicle information further includes the lane saturation degree. For each phase, use the sum of the traffic movement pressures of all traffic directions of the phase as the phase queue pressure of the phase, and calculate the phase effective pressure of the phase based on the sum of the weights of the effectively moving vehicles in all traffic directions of the phase and the phase queue pressure of the phase. The vehicle includes the moving vehicle, and the phase includes a set of non-conflicting traffic directions; A training module, configured to establish a traffic signal control model according to the traffic state information and the phase effective pressure, and input new traffic state information into the traffic signal control model for training; A control module, configured to control the traffic signals at a target intersection based on the trained traffic signal control model.
7. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, which when executed by the one or more processors, cause the electronic device to implement the traffic signal control method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which when executed by a processor of a computer, causes the computer to execute the traffic signal control method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Traffic signal lamp control method and system based on reinforcement learning
CN113380054A
Signal lamp control method, model training method, system and device and storage medium
CN113643528A