Traffic signal control system and method
Through multi-source perception equipment and hybrid modeling technology, combined with queue theory and deep reinforcement learning, and dynamically adjusting signal control strategies, the problem that signal lights cannot be adjusted in real time in the existing technology is solved, and efficient management and safety improvement of traffic flow are achieved.
Patent Information
- Application Number
- CN202510691131.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing road signal control equipment cannot dynamically adjust the signal duration according to real-time road conditions, resulting in frequent traffic congestion problems.
Multi-source sensing equipment is used to collect traffic flow data, combine queue theory, cellular automata and deep reinforcement learning through a hybrid modeling framework, build a dynamic intersection model, dynamically adjust signal control strategies using the dual closed-loop feedback mechanism, and link variable lane control equipment.
It improves the accuracy of traffic status assessment and the real-time signal control, reduces congestion, and improves traffic efficiency and safety.
Smart Images

Figure CN120472687A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic control, and in particular to a traffic signal control system and method. Background Art
[0002] In current road traffic control systems, especially at the entrances and exits of highways in major cities, traffic congestion, slow-moving vehicles, illegal parking, and other traffic jams often occur, as well as traffic jams caused by sudden traffic accidents. This causes great inconvenience to vehicles traveling on the highway and the relevant departments' highway traffic control.
[0003] Therefore, it is necessary to conduct real-time detection and control of traffic conditions on the road. When poor traffic conditions are detected, such as road congestion, slow-moving vehicles, illegal parking, pedestrian intrusion, etc., an alarm will be recorded and notified to the relevant departments immediately, so that nearby vehicles can be notified in time, which can alleviate the road congestion to a certain extent. However, the road conditions are also closely related to the signal control equipment on the road. For example, the duration of traffic lights is generally fixed, or the duration is set according to the time period, and it cannot be adjusted in real time according to the road conditions, and road congestion often occurs. Summary of the Invention
[0004] In order to overcome the above-mentioned shortcoming of the road signal control equipment in the prior art that it cannot adjust the road signal duration at any time according to road vehicle information, the main purpose of the present invention is to provide a traffic signal control system and method.
[0005] To achieve the above object, the present invention adopts the following technical solution, a traffic signal control method, comprising the following steps:
[0006] The system collects vehicle count, instantaneous speed, and steering intention data through multi-source sensing devices, and uses V2X communication to obtain the driving status of the fleet and obtain traffic flow data.
[0007] Preprocess the collected traffic flow data, including abnormal data cleaning and spatiotemporal alignment of multi-source data. Combined with the predicted lane-level flow, a structured traffic feature matrix is constructed.
[0008] Based on a hybrid modeling framework, a dynamic intersection model is constructed by integrating a queue theory model, a cellular automaton model, and a deep reinforcement learning agent. Historical traffic flow data and a structured traffic characteristic matrix are input into the queue theory model and the cellular automaton model to obtain the queue state evaluation index Sq and the microscopic traffic flow dynamic distribution Sca. Combined with the deep reinforcement learning agent, the integrated traffic state evaluation results are obtained, and the optimal signal control strategy is obtained through a multi-objective optimization algorithm.
[0009] A dual closed-loop feedback mechanism is used to dynamically adjust the signal control strategy. The inner loop fine-tunes the vehicle phase based on the PID controller, and the outer loop updates the parameters of the deep reinforcement learning agent through online learning to obtain the adjusted signal control strategy.
[0010] After obtaining the adjusted signal control strategy, the method further includes:
[0011] Execute the adjusted signal control strategy and link the variable lane control equipment, and at the same time push the signal light status information and dynamic vehicle speed guidance instructions in the adjusted signal control strategy to the on-board terminal.
[0012] The multi-source perception equipment includes a video camera, a laser radar, a millimeter-wave radar, a microwave detector, and a geomagnetic detector deployed at the intersection and surrounding areas;
[0013] The traffic flow data includes vehicle quantity data of each import lane, instantaneous speed data of vehicles in each import lane, turning intention data of vehicles in each import lane and driving status data of the fleet.
[0014] The pretreatment comprises the following steps:
[0015] Identifying the traffic flow data and eliminating erroneous data to obtain cleaned data, wherein the erroneous data includes speed abnormality data, impossible turning intention data, and erroneous data caused by sensor failure;
[0016] Performing outlier detection on the cleaned data through a rule base to obtain abnormal cleaned data;
[0017] Integrate the anomaly-cleaned data into a unified spatiotemporal frame to obtain aligned data;
[0018] Based on historical traffic flow data and real-time traffic flow data, traffic flow forecasting is performed through time series forecasting models to obtain predicted traffic flow;
[0019] Based on the aligned data and predicted traffic flow, a matrix of vehicle density, speed, and turning rate is constructed as a structured traffic feature matrix.
[0020] Obtaining the integrated traffic status assessment result includes the following steps:
[0021] Obtain historical traffic data for cleaning and feature extraction to obtain historical traffic flow data;
[0022] Using the historical traffic flow data and the structured traffic feature matrix as standardized input data sets;
[0023] The standardized input data set is respectively input into a queue theory model, a cellular automaton model, and a reinforcement learning agent. An M / M / 1 / K queue model is established using the queue theory model to obtain a queue state evaluation index Sq, which includes the maximum queue length Qmax of each lane and the average waiting time Wavg. The cellular automaton model is used to set the cell size and vehicle behavior rules to obtain a microscopic traffic flow dynamic distribution Sca. Based on the queue state evaluation index Sq and the microscopic traffic flow dynamic distribution Sca, the phase state, queue length, and arrival rate are obtained to construct a state space. The action space and reward function are defined using the reinforcement learning agent to obtain a strategic action value evaluation result Q(s,a).
[0024] A weighted fusion algorithm is used to perform weighted fusion processing on the queue state evaluation index Sq, the microscopic traffic flow dynamic distribution Sca and the strategic action value evaluation result Q(s,a) to obtain the fused comprehensive traffic state evaluation result.
[0025] The method of obtaining the optimal signal control strategy by a multi-objective optimization algorithm includes the following steps:
[0026] Construct the objective function, the formula is as follows:
[0027] min(P1·S q +A 2 Number of stops + max(P i ·u i ))
[0028] The constraints are:
[0029] T min ≤tW max ≤T max +ΔT
[0030] Among them, P1 weight coefficient is used to weigh the comprehensive state evaluation S q Importance in the objective function; A 2 The weight coefficient is used to weigh the importance of the number of stops in the objective function. The number of stops represents the number of times a vehicle stops and waits at an intersection. u i represents the i-th specific indicator, P i Represents the weight coefficient, which is used to weigh the i-th specific indicator u i Importance in the objective function, max(P i ·u i ) means that among all specific indicators, the specific indicator with the largest product of weight and indicator value is selected;
[0031] T minrepresents the minimum green light time, tWmax represents the maximum allowed waiting time of a vehicle at an intersection, T max represents the maximum green light time, and ΔT represents the allowed green light time adjustment range; obtain the optimal green light time combination scheme Popt;
[0032] The fused comprehensive traffic state evaluation result is input into the objective function to obtain the optimal phase duration combination solution P opt ;
[0033] Virtual deduction is performed in the digital twin system to obtain signal control effect data, and the signal control effect data is selected according to set key evaluation indicators to obtain the optimal signal control strategy, where the set key evaluation indicators include the number of conflict points and the maximum queue length.
[0034] The optimal signal control strategy includes the timing plan of each signal light and the control instructions of the lane change control device. The timing plan includes the phase sequence, green light time, yellow light time, and red light time; the control instructions of the lane change control device change the lane direction and adjust the lane function.
[0035] The inner loop adjusts the vehicle phase based on the PID controller, including:
[0036] The real-time detected traffic flow data is compared with the traffic flow target value preset in the optimal signal control strategy, and a phase adjustment amount is obtained using a PID control algorithm as feedback for the inner loop;
[0037] The outer loop updates the parameters of the deep reinforcement learning agent through online learning, including:
[0038] Using real-time traffic flow data and signal control effect data, the neural network weights of the deep reinforcement learning agent in the hybrid modeling framework are updated through online gradient descent as feedback for the outer loop;
[0039] Based on the feedback of the inner loop and the outer loop, the optimal signal control strategy is dynamically adjusted to generate an adjusted signal control strategy.
[0040] A communication signal control system, comprising:
[0041] Data acquisition module, used to collect road vehicle data, including vehicle quantity, speed, and driving direction;
[0042] The data processing module is used to pre-process the collected traffic flow data, including abnormal data cleaning, spatiotemporal alignment of multi-source data, lane-level flow prediction, and construction of a structured traffic feature matrix;
[0043] A decision module is used to construct a dynamic intersection model based on a hybrid modeling framework, input historical traffic data and real-time traffic data into the hybrid modeling framework integrating a queue theory model, a cellular automation model, and a deep reinforcement learning agent, obtain a fused comprehensive traffic state assessment result, and obtain the optimal signal control strategy through a multi-objective optimization algorithm;
[0044] A dual closed-loop feedback mechanism is used to dynamically adjust the signal control strategy. The inner loop fine-tunes the vehicle phase using a PID controller, while the outer loop updates the parameters of the deep reinforcement learning agent through online learning to obtain the adjusted signal control strategy.
[0045] The execution module is used to execute the adjusted signal control strategy and link the variable lane control equipment, and at the same time push the signal light status information and dynamic vehicle speed guidance instructions in the adjusted signal control strategy to the on-board terminal.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. This invention improves the comprehensiveness and accuracy of traffic flow data by collecting data from multiple sources, ensuring that traffic signal control systems make decisions based on comprehensive, real-time traffic information. By cleaning erroneous data, it reduces interference from abnormal data. Spatiotemporal alignment and flow prediction enable traffic control systems to integrate historical and real-time data, reducing misjudgments caused by data fragmentation, improving the effectiveness of signal strategies, and achieving more accurate traffic status assessments.
[0048] 2. The fusion of multiple models leverages their respective strengths: queue theory models assess queue states, cellular automata describe microscopic traffic flows, and deep reinforcement learning agents optimize signal control strategies. Queuing theory, cellular automata, and DRL complement each other, balancing theoretical rigor with dynamic adaptability. This multi-model fusion approach improves the accuracy and comprehensiveness of traffic state assessment. The multi-objective optimization algorithm effectively balances multiple traffic indicators, ensuring the efficiency and fairness of the signal control strategy while addressing the needs of different lanes and traffic conditions.
[0049] 3. The inner-loop PID controller provides fine-grained real-time signal control adjustments, ensuring high-precision signal timing. The outer-loop online learning continuously optimizes the parameters of the deep reinforcement learning agent based on real-time data. The inner-loop PID quickly responds to traffic bursts, while the outer-loop online learning adapts to long-term changes, demonstrating adaptability and continuous optimization.
[0050] 4. Dynamically adjust lane functions and directions, and use variable lanes, signal lights, and on-board induction to achieve global collaborative optimization, which can flexibly respond to changes in traffic flow, reduce congestion, and improve intersection capacity. Digital twin technology can verify and optimize signal control strategies in real time, virtually deduce conflict points, and screen risk-free strategies, avoiding the uncertainty in traditional methods, reducing safety hazards in actual deployment, and ensuring the reliability of strategies in practical applications. The present invention adopts advanced technologies such as multi-source perception, hybrid modeling, and deep reinforcement learning, fully considering multiple factors, improving the intelligence, real-time nature, and adaptability of traffic signal control, and helping to alleviate urban traffic congestion and improve traffic efficiency and safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the process structure of the present invention;
[0052] Figure 2 1 is a flow chart of obtaining the integrated traffic status evaluation result according to the present invention;
[0053] Figure 3 It is a flow chart of obtaining the strategic action value evaluation result of the present invention. DETAILED DESCRIPTION
[0054] In order to ensure that road traffic congestion is minimized, it is necessary to conduct real-time detection and control of traffic conditions on the road. When poor traffic conditions are detected, such as road congestion, slow-moving vehicles, illegal parking, pedestrian intrusion, etc., an alarm will be recorded and notified to the relevant departments immediately so that nearby vehicles can be notified in time, which alleviates the road congestion to a certain extent. However, the road conditions are also closely related to the signal control equipment on the road. For example, the duration of traffic lights is generally fixed, or the duration is set according to the time period, and cannot be adjusted in real time according to the road conditions, and road congestion often occurs.
[0055] In order to overcome the above-mentioned shortcomings of the prior art road signal control device that cannot adjust the road signal duration at any time according to road vehicle information, refer to Figure 1 ,The main purpose of the present invention is to provide a traffic signal ,control method.
[0056] To achieve the above object, the present invention adopts the following technical solution, a traffic signal control method, comprising the following steps:
[0057] The system collects vehicle quantity, instantaneous speed, and steering intention data through multi-source sensing devices, and uses V2X communication to obtain the driving status of the fleet and traffic flow data.
[0058] Preprocessing of collected traffic flow data, including abnormal data cleaning, spatiotemporal alignment of multi-source data, and lane-level flow prediction, to construct a structured traffic feature matrix;
[0059] Based on a hybrid modeling framework, a dynamic intersection model is constructed. Historical and real-time traffic data are input into the hybrid modeling framework, which integrates a queue theory model, a cellular automation model, and a deep reinforcement learning agent. The integrated traffic state assessment results are obtained, and the optimal signal control strategy is obtained through a multi-objective optimization algorithm.
[0060] A dual closed-loop feedback mechanism is used to dynamically adjust the signal control strategy. The inner loop adjusts the vehicle phase based on the PID controller, and the outer loop updates the parameters of the deep reinforcement learning agent through online learning to obtain the adjusted signal control strategy.
[0061] And if further executed, the adjusted signal control strategy will be executed and the variable lane control equipment will be linked, and the signal light status information and dynamic vehicle speed guidance instructions in the adjusted signal control strategy will be pushed to the vehicle terminal.
[0062] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0063] Example 1:
[0064] This embodiment applies the traffic signal control method to a large intersection in a city center. This embodiment will explain in detail how to apply the method at this intersection, provide relevant data values, and describe the specific implementation of each step.
[0065] The intersection is located in the city center and has four entrance lanes, each with different traffic flows. The following sensing devices are deployed around the intersection:
[0066] Video cameras: used to monitor lane and traffic conditions.
[0067] LiDAR: Used to accurately measure the distance and speed of vehicles.
[0068] Millimeter-wave radar: used to detect traffic density and lane usage.
[0069] Microwave detector: used to detect the real-time flow of vehicles.
[0070] Geomagnetic detector: used to monitor vehicle detection information on the lane and determine traffic flow direction based on the vehicle's dwell time.
[0071] Through these sensing devices, the following traffic flow data (unit: vehicles / hour) is obtained:
[0072] Lane 1 (Import 1): Traffic volume is 300 vehicles / hour, average speed is 40 km / hour, left-turning vehicles account for 20% and right-turning vehicles account for 10%.
[0073] Lane 2 (Import 2): Traffic volume is 400 vehicles / hour, average speed is 35 km / hour, left-turning vehicles account for 25% and right-turning vehicles account for 15%.
[0074] Lane 3 (Import 3): Traffic volume is 350 vehicles / hour, average speed is 38 km / hour, left-turning vehicles account for 15% and right-turning vehicles account for 20%.
[0075] Lane 4 (Import 4): Traffic volume is 450 vehicles / hour, average speed is 30 km / hour, left-turning vehicles account for 10% and right-turning vehicles account for 25%.
[0076] Then, traffic flow data is preprocessed, including data cleaning, data alignment and flow prediction, specifically:
[0077] The first step is to eliminate invalid data, especially erroneous data caused by sensor failure or weather interference, such as excessively high or low speed data. For example, the speed data for lane 2 was abnormal, exceeding 50 kilometers per hour. This erroneous data was filtered out using the rule base.
[0078] All sensor data is then aligned in time and space to ensure that data collected by different devices are consistent at the same timestamp. For example, traffic flow data in lane 1 and lane 2 are synchronized using the spatiotemporal alignment model.
[0079] Next, based on historical traffic data and real-time traffic data, the ARIMA model is used to predict traffic for the next 30 minutes. The prediction results are as follows:
[0080] Lane 1: Traffic flow is predicted to increase to 350 vehicles per hour.
[0081] Lane 2: The predicted flow rate increases to 420 vehicles per hour.
[0082] Lane 3: The predicted flow rate increases to 380 vehicles per hour.
[0083] Lane 4: Traffic is forecast to increase to 470 vehicles per hour.
[0084] Based on the above collected and preprocessed data, see Figure 2 and Figure 3 , enter the hybrid modeling framework and integrate queue theory models, cellular automata models, and deep reinforcement learning agent models:
[0085] In the queue theory model, the M / M / 1 / K queue model is used to evaluate the queue status of each import lane, and the following results are obtained:
[0086] The maximum queue length in lane 1 is 8 vehicles, and the average waiting time is 1.5 minutes.
[0087] The maximum queue length in lane 2 is 10 vehicles, and the average waiting time is 2 minutes.
[0088] The maximum queue length in lane 3 is 9 vehicles, and the average waiting time is 1.8 minutes.
[0089] The maximum queue length in lane 4 is 12 vehicles, and the average waiting time is 2.5 minutes.
[0090] In the cellular automaton model, the dynamic distribution of microscopic traffic flow, that is, the distribution of vehicles, is simulated by the cellular automaton model to obtain the traffic density of each lane.
[0091] For the deep reinforcement learning agent, the current traffic signal control strategy is evaluated based on the deep reinforcement learning model, and the impact of signal timing on traffic flow is predicted.
[0092] The objective function is constructed using a multi-objective optimization algorithm, and the objective function is solved to obtain the optimal signal control strategy.
[0093] Construct the objective function, the formula is as follows:
[0094] min(P1·S q +A 2 Number of stops + max(P i ·u i ))
[0095] The constraints are:
[0096] T min ≤tW max ≤T max +ΔT
[0097] Among them, P1 weight coefficient is used to weigh the importance of comprehensive state evaluation Sq in the objective function; A 2 The weight coefficient is used to weigh the importance of the number of stops in the objective function. The number of stops represents the number of times a vehicle stops and waits at an intersection. u i represents the i-th specific indicator, P i Represents the weight coefficient, which is used to weigh the i-th specific indicator u i Importance in the objective function, max(P i ·u i ) means that among all specific indicators, the specific indicator with the largest product of weight and indicator value is selected;
[0098] T minrepresents the minimum green light time, tWmax represents the maximum allowed waiting time of a vehicle at an intersection, T max Maximum green light time, ΔT represents the allowed green light time adjustment range; obtain the optimal green light time combination scheme Popt;
[0099] The fused comprehensive traffic state evaluation result is input into the objective function to obtain the optimal phase duration combination solution P opt ;
[0100] Virtual deduction is performed in the digital twin system to obtain signal control effect data, and the signal control effect data is selected according to set key evaluation indicators to obtain the optimal signal control strategy, where the set key evaluation indicators include the number of conflict points and the maximum queue length.
[0101] The number of stops represents the number of times a vehicle stops and waits at an intersection. Through virtual simulation, the optimal signal control strategy is determined as follows:
[0102] Lane 1: Green light duration is 20 seconds, yellow light duration is 5 seconds, and red light duration is 30 seconds.
[0103] Lane 2: Green light time is 22 seconds, yellow light time is 5 seconds, and red light time is 28 seconds.
[0104] Lane 3: Green light time is 18 seconds, yellow light time is 5 seconds, and red light time is 32 seconds.
[0105] Lane 4: Green light duration is 25 seconds, yellow light duration is 5 seconds, and red light duration is 25 seconds.
[0106] A dual closed-loop feedback mechanism with dynamic adjustments. For the inner loop, real-time traffic flow data is used to fine-tune the PID controller. Lane 2's traffic flow rate reached 450 vehicles per hour, deviating from the target flow rate of 420 vehicles per hour. Therefore, the PID controller adjusted the green light duration for that lane to 24 seconds. For the outer loop, based on real-time data and the adjusted signal control results, the parameters of the deep reinforcement learning agent are updated using online gradient descent to better predict traffic flow dynamics.
[0107] The signal control strategy finally implemented includes: dynamic signal light timing plan and lane change control instructions.
[0108] Lane change control instructions adjust lane directions during peak hours, allowing vehicles in certain lanes to turn through dedicated lanes to optimize lane usage.
[0109] This intelligent traffic signal control method effectively manages traffic flow at the intersection, reducing congestion and vehicle waiting times. After implementation, the maximum queue length in lane 1 decreased by 20%, the average waiting time in lane 2 decreased by 10%, and traffic efficiency at the entire intersection increased by 15%.
[0110] Example 2
[0111] In this embodiment, at the intersection of a main road and a secondary road in a certain city, the traffic volume is heavy during the morning rush hour (7:30-9:00), with left-turning vehicles from the east entrance accounting for 40% and straight-going vehicles from the west entrance accounting for 60%.
[0112] The configured multi-source sensing devices are:
[0113] LiDAR, detection range 100m, accuracy ±0.1m;
[0114] Video camera, 4K resolution, 30fps frame rate;
[0115] Geomagnetic sensor, lane-level vehicle counting error <2%;
[0116] V2X roadside unit; RSU coverage radius 300m
[0117] There are variable lanes, a tidal lane is set up at the east entrance, and a dedicated left-turn lane is used during the morning rush hour from 7:00 to 9:00.
[0118] Traffic flow characteristics of the intersection:
[0119] Real-time vehicle count: 200 vehicles / hour for east import and 300 vehicles / hour for west import.
[0120] Average speed: 15km / h for the east entrance (congested), 25km / h for the west entrance.
[0121] Turning intention: 40% for east entrance turning left and 60% for going straight; 70% for west entrance going straight and 30% for turning right.
[0122] First, multi-source data was collected. The lidar detected that the queue length at the east entrance was 85 meters. The camera identified that the instantaneous speed of vehicles entering the west was 22 km / h. The V2X obtained the fleet status as a three-vehicle formation with a coordinated speed of 20 km / h. Traffic flow data was obtained.
[0123] Traffic flow data was then preprocessed. First, anomaly cleaning was performed to remove three data points with speeds greater than 150 km / h, which were falsely detected by the radar. Two data points were also corrected where the camera misjudged a "left turn" as "straight ahead." Temporal and spatial alignment was then performed. Using a 90-second signal cycle as a window, the timestamps of the data from each device were aligned with an error of less than 0.1 seconds. This achieved temporal and spatial alignment. The camera coordinates (pixel point (1200, 800)) were converted to physical coordinates (east entrance lane 2, 50 meters from the stop line) to achieve spatial alignment. Historical traffic data, representing the same time period over the previous seven days, was then fed into the LSTM model. The prediction for the east entrance traffic flow over the next five minutes was 220 vehicles / hour (with a 5% error rate). After completing this process, the structured traffic feature matrix was obtained, as shown in Table 1:
[0124] Table 1 Example of structured traffic feature matrix
[0125]
[0126]
[0127] Next, hybrid modeling and multi-objective optimization, model input and calculation:
[0128] Queue theory model (M / M / 1 / K): The maximum queue length at the east entrance, Qmax, is 85 m, and the average waiting time, Wavg, is 120 seconds.
[0129] Cellular automaton model: The cell size is 2m×5m, simulating vehicle lane-changing behavior, outputting microscopic congestion hotspots and two conflict points on the west entrance through lane.
[0130] Deep Reinforcement Learning (DQN) agent: The state space is: the current phase is a green light at the east entrance, the queue length is 85 meters, and the arrival rate is 220 vehicles / hour. The reward function is to minimize the delay, and the action space is "extend the green light at the east entrance by 10 seconds."
[0131] The next step is weight allocation. In this example, the queuing theory is 0.4, the cellular automation is 0.3, and the DRL is 0.3. The comprehensive state evaluation result is a congestion index of 0.75. The threshold value > 0.6 triggers optimization.
[0132] Next, a multi-objective optimization (NSGA-II) was performed, with the objective functions minimizing delays (weight P1 = 0.6) and the number of stops (weight A2 = 0.4), and the constraints of green light durations Tmin = 15 seconds and Tmax = 60 seconds. The Pareto optimal solution was achieved by increasing the green light duration for the east entrance from 30 seconds to 45 seconds and reducing the green light duration for the west entrance from 40 seconds to 30 seconds. In the subsequent digital twin verification, the number of simulated conflict points was reduced from 2 to 0, and the maximum queue length was shortened from 85 meters to 50 meters. The optimal signal control strategy was ultimately obtained, as shown in Table 2:
[0133] Table 2 Optimal signal control strategy
[0134]
[0135] And variable lane control instructions, specifically: the east entrance tidal lane during the morning rush hour is switched to a dedicated left-turn lane.
[0136] Dynamic adjustment is also required for dual closed-loop feedback. The inner loop utilizes PID control, monitoring the queue length (50m) at the east entrance in real time. The target value is 40m, and the deviation ΔQ = 10m. The resulting PID output is: proportional gain Kp = 0.5, integral time Ti = 10 seconds, and the green light duration is adjusted by +5 seconds, from 45 to 50 seconds. For the outer loop, online learning is used to collect actual delay data every 10 minutes, reducing it from 120 seconds to 90 seconds. The DRL Q network weights are updated using gradient descent with a learning rate of 0.001. The NSGA-II weight coefficient is adjusted, increasing the delay weight P1 from 0.6 to 0.7. This adjustment prioritizes congestion relief. The final dynamically adjusted policy results in a 50-second green light for the east entrance and a 25-second green light for the west entrance. Furthermore, the model prediction error is reduced by 2% after online learning.
[0137] The final signal control strategy was implemented and a phase plan was issued. The east entrance had a green light of 50 seconds, with the tidal lane dedicated to left turns. The RSU sent a notification to the vehicle terminal: "The green light at the east entrance has 20 seconds left. We recommend accelerating to 25 km / h." Vehicles entering from the west received a notification: "The red light has 40 seconds left. We recommend slowing down to 15 km / h."
[0138] Implementation of the final signal control strategy improved traffic efficiency, with the following results: average delays at the east entrance were reduced from 120 seconds to 75 seconds, a 37.5% reduction in delay rate. The number of stops at the west entrance was reduced from 8 per cycle to 5 per cycle, a 37.5% reduction in traffic frequency. Furthermore, optimized driving behavior reduced the number of sudden braking incidents by queued vehicles by 50% and the probability of secondary queues by 30%.
[0139] For some occasional events that require regulation of the intersection of main and secondary roads, this has also been verified in scenarios such as sudden accidents and severe weather:
[0140] In an emergency, an accident at the west entrance caused the lane to be closed. The PID control reallocated the green light duration within 3 cycles, and the queue length was reduced from 100m to 60m.
[0141] In severe weather conditions, such as heavy rain causing low visibility, online learning is used to adjust the weights of the DRL reward function, prioritizing safe passage and eliminating conflict points.
[0142] This example verifies the effectiveness of this method in complex urban traffic scenarios and quantifies the technical advantages through specific data, such as a 37.5% reduction in delays.
[0143] It should be noted that, in the present invention, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0144] The above embodiments are merely examples of the present invention and do not limit the scope of protection of the present invention. Any designs that are identical or similar to the present invention fall within the scope of protection of the present invention.
Claims
1. A traffic signal control method, characterized in that: The following steps are involved: The system collects vehicle count, instantaneous speed, and steering intention data through multi-source sensing devices, and uses V2X communication to obtain the driving status of the fleet and obtain traffic flow data. Preprocess the collected traffic flow data, including abnormal data cleaning and spatiotemporal alignment of multi-source data. Combined with the predicted lane-level flow, a structured traffic feature matrix is constructed. Based on a hybrid modeling framework, a dynamic intersection model is constructed by integrating a queue theory model, a cellular automaton model, and a deep reinforcement learning agent. Historical traffic flow data and a structured traffic characteristic matrix are input into the queue theory model and the cellular automaton model to obtain the queue state evaluation index Sq and the microscopic traffic flow dynamic distribution Sca. Combined with the deep reinforcement learning agent, the integrated traffic state evaluation results are obtained, and the optimal signal control strategy is obtained through a multi-objective optimization algorithm. A dual closed-loop feedback mechanism is used to dynamically adjust the signal control strategy. The inner loop adjusts the vehicle phase based on the PID controller, and the outer loop updates the parameters of the deep reinforcement learning agent through online learning to obtain the adjusted signal control strategy.
2. A traffic signal control method according to claim 1, characterized in that: After obtaining the adjusted signal control strategy, the method further includes: Execute the adjusted signal control strategy and link the variable lane control equipment, and at the same time push the signal light status information and dynamic vehicle speed guidance instructions in the adjusted signal control strategy to the on-board terminal.
3. A traffic signal control method according to claim 1, characterized in that: The multi-source perception equipment includes a video camera, a laser radar, a millimeter-wave radar, a microwave detector, and a geomagnetic detector deployed at the intersection and surrounding areas; The traffic flow data includes vehicle quantity data of each import lane, instantaneous speed data of vehicles in each import lane, turning intention data of vehicles in each import lane and driving status data of the fleet.
4. A traffic signal control method and system according to claim 1, characterized in that: The pretreatment comprises the following steps: Identifying the traffic flow data and eliminating erroneous data to obtain cleaned data, wherein the erroneous data includes speed abnormality data, impossible turning intention data, and erroneous data caused by sensor failure; Performing outlier detection on the cleaned data through a rule base to obtain abnormal cleaned data; Integrate the anomaly-cleaned data into a unified spatiotemporal frame to obtain aligned data; Based on historical traffic flow data and real-time traffic flow data, traffic flow forecasting is performed through time series forecasting models to obtain predicted traffic flow; Based on the aligned data and predicted traffic flow, a matrix of vehicle density, speed, and turning rate is constructed as a structured traffic feature matrix.
5. A traffic signal control method according to claim 4, characterized in that: Obtaining the integrated traffic status assessment result includes the following steps: Obtain historical traffic data for cleaning and feature extraction to obtain historical traffic flow data; Using the historical traffic flow data and the structured traffic feature matrix as standardized input data sets; The standardized input data set is respectively input into a queue theory model, a cellular automaton model, and a reinforcement learning agent. An M / M / 1 / K queue model is established using the queue theory model to obtain a queue state evaluation index Sq, which includes the maximum queue length Qmax of each lane and the average waiting time Wavg. The cellular automaton model is used to set the cell size and vehicle behavior rules to obtain a microscopic traffic flow dynamic distribution Sca. Based on the queue state evaluation index Sq and the microscopic traffic flow dynamic distribution Sca, the phase state, queue length, and arrival rate are obtained to construct a state space. The action space and reward function are defined using the reinforcement learning agent to obtain a strategic action value evaluation result Q(s,a). A weighted fusion algorithm is used to perform weighted fusion processing on the queue state evaluation index Sq, the microscopic traffic flow dynamic distribution Sca and the strategic action value evaluation result Q(s,a) to obtain the fused comprehensive traffic state evaluation result.
6. A traffic signal control method according to claim 5, characterized in that: The method of obtaining the optimal signal control strategy by a multi-objective optimization algorithm includes the following steps: Construct the objective function, the formula is as follows: min(P1·S q +A 2 Number of stops + max(P i ·u i )) The constraints are: T min ≤tW max ≤T max +ΔT Among them, P1 weight coefficient is used to weigh the comprehensive state evaluation S q Importance in the objective function; A 2 The weight coefficient is used to weigh the importance of the number of stops in the objective function. The number of stops represents the number of times a vehicle stops and waits at an intersection. u i represents the i-th specific indicator, P i Represents the weight coefficient, which is used to weigh the i-th specific indicator u i Importance in the objective function, max(P i ·u i ) means that among all specific indicators, the specific indicator with the largest product of weight and indicator value is selected; T min represents the minimum green light time, tWmax represents the maximum allowed waiting time of a vehicle at an intersection, T max Indicates the maximum green light time, ΔT indicates the allowed green light time adjustment range; obtain the optimal green light time combination solution; The fused comprehensive traffic state evaluation result is input into the objective function to obtain the optimal phase duration combination solution P opt ; Virtual deduction is performed in the digital twin system to obtain signal control effect data, and the signal control effect data is selected according to set key evaluation indicators to obtain the optimal signal control strategy, where the set key evaluation indicators include the number of conflict points and the maximum queue length.
7. A traffic signal control method according to claim 6, characterized in that: The optimal signal control strategy includes the timing plan of each signal light and the control instructions of the lane change control device. The timing plan includes the phase sequence, green light time, yellow light time, and red light time; the control instructions of the lane change control device change the lane direction and adjust the lane function.
8. A traffic signal control method according to claim 6, characterized in that: The inner loop adjusts the vehicle phase based on the PID controller, including: The real-time detected traffic flow data is compared with the traffic flow target value preset in the optimal signal control strategy, and a phase adjustment amount is obtained using a PID control algorithm as feedback for the inner loop; The outer loop updates the parameters of the deep reinforcement learning agent through online learning, including: Using real-time traffic flow data and signal control effect data, the neural network weights of the deep reinforcement learning agent in the hybrid modeling framework are updated through online gradient descent as feedback for the outer loop; Based on the feedback of the inner loop and the outer loop, the optimal signal control strategy is dynamically adjusted to generate an adjusted signal control strategy.
9. A traffic signal control system, characterized in that: include: Data acquisition module, used to collect road vehicle data, including vehicle quantity, speed, and driving direction; The data processing module is used to pre-process the collected traffic flow data, including abnormal data cleaning, spatiotemporal alignment of multi-source data, and construction of a structured traffic feature matrix based on the predicted lane-level flow. The decision-making module is used to integrate the queue theory model, cellular automaton model, and deep reinforcement learning agent based on a hybrid modeling framework to build a dynamic intersection model. Historical traffic flow data and structured traffic characteristic matrix are input into the queue theory model and cellular automaton model to obtain the queue state evaluation index Sq and the microscopic traffic flow dynamic distribution Sca. Combined with the deep reinforcement learning agent, the fused comprehensive traffic state evaluation result is obtained, and the optimal signal control strategy is obtained through a multi-objective optimization algorithm. A dual closed-loop feedback mechanism is used to dynamically adjust the signal control strategy. The inner loop fine-tunes the vehicle phase based on the PID controller, and the outer loop updates the parameters of the deep reinforcement learning agent through online learning to obtain the adjusted signal control strategy.
10. The traffic signal control system according to claim 9, characterized in that: Also includes: The execution module is used to execute the adjusted signal control strategy and link the variable lane control equipment, and at the same time push the signal light status information and dynamic vehicle speed guidance instructions in the adjusted signal control strategy to the on-board terminal.
Citation Information
Cited By
Vehicle-mounted communication signal control method based on deep learning
CN121056905A