A single intersection traffic signal control method and system based on video detection results

By collecting and mapping video data in single-intersection traffic signal control, and using reinforcement learning models for decision-making and action constraints, the problems of traffic flow fluctuations and directional imbalances are solved, enabling real-time adaptive traffic signal control and improving the real-time performance and stability of the control.

CN122050168BActive Publication Date: 2026-07-31CENT SOUTH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-04-16
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing single-intersection traffic signal control methods are unable to reflect the random fluctuations and directional imbalances of traffic flow in real time, leading to increased vehicle queues and decreased traffic efficiency. Furthermore, the results of multi-channel video perception are difficult to directly unify and use for control decisions.

Method used

By collecting video data from four directions at the intersection, perception mapping and state construction are performed. A reinforcement learning control model is used to make decisions and apply action constraints. Finally, phase change control is executed to form a complete technical closed loop.

Benefits of technology

It enables real-time, adaptive control of traffic signals at a single intersection, improving the real-time performance, stability, and adaptability of intersection control, reducing the average queue length, and increasing traffic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050168B_ABST
    Figure CN122050168B_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent traffic signal control technology, and discloses a single-intersection traffic signal control method and system based on video detection results. The method includes: collecting video data from four directions of the intersection; extracting traffic parameters based on the detection results of each video stream; constructing a unified intersection state vector; inputting the unified intersection state vector into a reinforcement learning control model to obtain candidate phase actions; applying constraints to the candidate phase actions to obtain target phase actions; executing phase change control according to the target phase actions; and implementing an anomaly fallback strategy in case of video anomalies, control anomalies, or program anomalies to maintain the continuity of single-intersection traffic signal control. This solves the problems of existing single-intersection traffic signal control, such as the difficulty in uniformly converting multi-video detection results into control states, easy oscillation during phase switching, and insufficient continuous control capability under abnormal conditions, thus improving the real-time performance, stability, and adaptability of single-intersection traffic signal control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent traffic control technology, specifically to a single-intersection traffic signal control method and system based on video detection results. Background Technology

[0002] Traffic signal control is a core component of urban road traffic management. At-grade intersections, as critical nodes in the road network, directly impact vehicle throughput, delay levels, and the stability of the road network. Existing single-intersection signal control methods primarily rely on fixed timing, time-based timing, or simple sensor control, typically depending on pre-set release durations, detector triggering conditions, or empirical rules. When traffic flow exhibits significant random fluctuations, directional imbalances, or periodic clustering, traditional control methods struggle to reflect the actual operational status of the intersection in a timely manner, easily leading to increased queues at some approach lanes, increased vehicle dwell times, and a decline in overall traffic efficiency.

[0003] With the development of video sensing and intelligent analysis technologies, acquiring vehicle operation information using intersection surveillance videos has become an important technical approach for traffic status collection. While existing technologies can obtain traffic information such as vehicle numbers, queue conditions, and traffic trajectories through video analysis, most solutions still focus on independent statistics from a single perspective or local area. Especially in single-intersection control scenarios, existing methods generally lack a complete and suitable real-time decision-making mechanism for transforming discrete traffic information from different approach directions into control data that reflects the overall operational characteristics of the intersection. Summary of the Invention

[0004] This invention aims to solve the problem that the results of multi-channel video perception are difficult to directly input into the signal control model in the prior art. It provides a single-intersection traffic signal control method and system based on video detection results, which forms a complete technical closed loop of four-channel video perception, state modeling, reinforcement learning decision-making, action constraints, phase state machine execution and abnormal backoff.

[0005] To achieve the above objectives, the first aspect of the present invention provides a single-intersection traffic signal control method based on video detection results, comprising the following steps:

[0006] S1 collects video data and signal phase information from the four directions of the intersection, and performs perception mapping on the video data of each lane to obtain lane-level traffic parameters for each lane. S2, based on the lane-level traffic parameters of each lane and the correspondence between each lane and its respective approach direction, aggregate according to the approach direction and combine with the signal phase information to construct a unified intersection state vector. The unified intersection state vector includes at least the direction-level queue length and the current signal phase of each approach direction. S3, input the unified intersection state vector into the pre-trained reinforcement learning control model, and the reinforcement learning control model outputs candidate phase actions. The reinforcement learning control model is a model that outputs phase actions based on the state vector. S4, apply action constraints to the candidate phase actions to obtain the target phase action; S5, perform phase change control according to the target phase action.

[0007] A second aspect of the present invention provides a single-intersection traffic signal control system based on video detection results, comprising: The data acquisition and processing module collects video data and signal phase information from the four directions of the intersection, and performs perception mapping on the video data of each lane to obtain lane-level traffic parameters for each lane. The state construction module, based on the lane-level traffic parameters of each lane and the correspondence between each lane and its respective approach direction, aggregates them according to the approach direction and combines them with the signal phase information to construct a unified intersection state vector. The unified intersection state vector includes at least the direction-level queue length and the current signal phase of each approach direction. The decision module inputs the unified intersection state vector into a pre-trained reinforcement learning control model to obtain candidate phase actions; The constraint processing module applies action constraints to the candidate phase actions to obtain the target phase action; The phase execution module performs phase change control according to the target phase action.

[0008] Based on the above technical solution, the present invention has the following beneficial effects: This invention unifies the acquisition, perception mapping, state construction, reinforcement learning decision-making, and phase execution control of video data from four directions at an intersection. This transforms traffic detection results, originally scattered across different video feeds, into a unified intersection state representation suitable for signal control. Based on this, it outputs target phase actions that meet control constraints, thereby achieving real-time, adaptive control of traffic signals at a single intersection. This effectively solves the problem of the difficulty in directly unifying the results of multiple video detections for control decisions and helps improve the real-time performance, stability, and adaptability of intersection control. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1The schematic diagram illustrates a flowchart of a single-intersection traffic signal control method based on video detection results according to an embodiment of the present invention. Figure 2 The diagram illustrates the structure of a single-intersection traffic signal control system based on video detection results according to an embodiment of the present invention. Detailed Implementation

[0011] To facilitate understanding of the present invention, the present invention will be described more fully and in detail below with reference to the accompanying drawings and preferred embodiments, but the scope of protection of the present invention is not limited to the following specific embodiments.

[0012] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those skilled in the art. The technical terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the invention.

[0013] Please see Figure 1 One embodiment of the present invention provides a single-intersection traffic signal control method based on video detection results, comprising the following steps: S1 collects video data and signal phase information from the four directions of the intersection, and performs perception mapping on the video data of each lane to obtain lane-level traffic parameters for each lane. The signal phase information specifically includes the current signal phase and the duration of the current phase; the perception mapping of each video data is specifically to map the traffic information reflecting the vehicle's operating status in the video data to their respective lanes, and to perform statistics according to each lane to obtain the lane-level traffic parameters corresponding to each lane. S2, based on the lane-level traffic parameters of each lane and the correspondence between each lane and its respective approach direction, aggregates them according to the approach direction and combines them with the signal phase information to construct a unified intersection state vector. The unified intersection state vector includes at least the direction-level queue length and the current signal phase of each direction. Specifically, the aggregation is performed according to the direction of entry. This involves merging the lane-level traffic parameters of each lane under the same direction of entry to form the direction-level traffic state information for the corresponding direction of entry. Then, the direction-level traffic state information is combined with the signal phase information at the current moment to construct a unified intersection state vector that can characterize the overall traffic operation and control state of the intersection. S3, input the unified intersection state vector into the pre-trained reinforcement learning control model, and the reinforcement learning control model outputs candidate phase actions; S4, apply action constraints to the candidate phase actions to obtain the target phase action; Among them, action constraints are used to restrict the executability of candidate phase actions and correct the control logic so that the control actions after constraint processing meet the operational requirements of the traffic signal control process; the target phase action is the phase action finally determined under the condition of satisfying the action constraints. S5, perform phase change control according to the target phase action.

[0014] This embodiment unifies the acquisition, perception mapping, state construction, reinforcement learning decision-making, and phase execution control of video data from four directions at an intersection. This transforms traffic detection results, originally scattered across different video feeds, into a unified intersection state representation suitable for signal control. Based on this, it outputs target phase actions that meet control constraints, thereby achieving real-time, adaptive control of traffic signals at a single intersection. These steps effectively solve the problem of directly unifying multi-channel video detection results for control decisions and contribute to improving the real-time performance, stability, and adaptability of intersection control.

[0015] Specifically, in step S1, perceptual mapping is performed on the video data from each path to obtain lane-level traffic parameters for each lane, including: Define the video data from the four directions of the intersection as a set of video streams;

[0016] in, This represents the set of video streams from the four directions of the intersection. These represent the video streams corresponding to the four import directions; And establish a set of directions:

[0017] in, This represents the set of directions of the intersection's entrance. These represent the four directions of entry: east, south, west, and north. And the mapping relationship between video and direction For each video stream at any time t Image frames Preprocessing may be performed; preprocessing may include one or more of the following: resolution unification, frame rate adaptation, timestamp alignment, and region of interest cropping, in order to improve the consistency and stability of subsequent detection and tracking processing.

[0018] After preprocessing, a deep learning-based object detection model is used to analyze the image frames of each video stream. Perform target detection to obtain a set of vehicle target detection results:

[0019] in, For the target bounding box, The target category (which may specifically include cars, trucks, buses, and motorcycles, etc.) To test the confidence level; Multi-target tracking is performed on the target detection results, specifically by using a trajectory association-based multi-target tracking model to generate continuous target trajectories across frames. ;

[0020] in, For the set of tracking results, For vehicle trajectory identification.

[0021] Based on pre-defined lane areas, stop line areas, and cross-line judgment lines, the trajectories of each vehicle are... The target is mapped to the corresponding lane, and its lane affiliation and lane crossing status are determined based on whether its center point enters the corresponding area or the overlap between the target's bounding box and the corresponding area. This establishes a correspondence between vehicle targets in the original video and actual lanes, achieving a mapping from video space to lane traffic semantic space, where the lane mapping satisfies:

[0022] in, Represents the target trajectory At any moment Lane number, Indicates the target vehicle At any moment The trajectory bounding box, Indicates the first k The first in the road video Area masking for lanes, The target vehicle's identification number; The intersection-union ratio (IUU) function is used to characterize the degree of overlap between the target bounding box and the lane region. Within a preset time period Δt, assuming the same trajectory marker is counted only once, the real-time number of vehicles, the number of queued vehicles, the waiting time, and the number of vehicles crossing the line are statistically analyzed for each lane. The traffic flow for each lane is then calculated based on the number of vehicles crossing the line. In other words, the lane-level traffic parameters for each lane include the real-time number of vehicles, the number of queued vehicles, the waiting time, and the traffic flow for the preset time period.

[0023]

[0024] in, Indicates the first lane at time The approximate number of vehicles in the queue (i.e., the number of vehicles in the queue). For the first The car at any time instantaneous speed, To determine the threshold for queuing, With the target center point, For the first The lane area corresponding to each lane; Indicates the preset duration Inner Traffic flow of each lane The moment when the m-th vehicle crosses the j-th detection line.

[0025] The statistical results for each lane can be reported via queue messages and written to a local persistent file, serving as the basic input for subsequent directional aggregation and unified intersection status construction.

[0026] S1 can extract lane-level traffic parameters with a unified structure and clear semantics from four video feeds, providing a basic input for the subsequent construction of a unified intersection state vector.

[0027] In a basic implementation, a unified intersection state vector is used. The signal is composed of the number of vehicles queuing in four directions and the current signal phase, satisfying the following:

[0028] in, These represent the number of vehicles queuing in the four directions: east, south, west, and north, respectively. Indicates the current phase number, superscript This represents the transpose of a vector.

[0029] Furthermore, for the input modeling scenario, the number of vehicles queuing in each direction can also be obtained by aggregating the number of vehicles at the lane level, for example:

[0030] in, and These represent the number of vehicles queuing in the two eastbound input lanes, respectively. and These represent the number of vehicles queuing in the two southbound input lanes, respectively. and These represent the number of vehicles queuing in the two westbound input lanes, respectively. and These represent the number of vehicles queuing in the two northbound input lanes, respectively.

[0031] In an extended embodiment, the unified intersection state vector in S2 is constructed by aggregating lane-level traffic parameters according to the approach direction and concatenating them with the current signal phase, specifically: For any direction ,and ,set up This represents the set of lanes corresponding to that direction, and the direction-level queue length for that direction. Directional waiting time Directional unit time flow and directional flow change trend They respectively satisfy:

[0032]

[0033]

[0034]

[0035] in, For a moment No. The number of vehicles queuing in each lane; For a moment No. Waiting time statistics for each lane Preset duration Inner Traffic flow in each lane; Unified intersection state vector satisfy:

[0036] in, This is the current phase number or phase code. The duration of the current phase. For video validity marking, and:

[0037]

[0038] When the video in the corresponding direction is valid, When the video in the corresponding direction fails, .

[0039] In one embodiment, during the training process, the reinforcement learning control model in S3 abstracts the single-intersection signal control problem into a Markov decision process for signal control, expressed as:

[0040] in, For state space, For the action space, This is a state transition relation. For the reward function, Discount factor; The state space is composed of a unified intersection state vector obtained from multiple video perceptions. The construction of the unified intersection state vector includes at least a lane-level statistical layer, a direction-level aggregation layer, and a control state layer. The lane-level statistics layer receives the current number of vehicles, the number of vehicles in queue, and the traffic flow for each lane; the direction-level aggregation layer aggregates data according to the direction of entry. The lane-level traffic parameters of each lane collected in the lane-level statistics layer are aggregated to obtain the direction-level traffic state quantities for each direction; the control state layer is used to combine the direction-level traffic state quantities with the current signal phase and the duration of the current phase to construct a unified intersection state vector. The action space consists of signal phase actions, including at least maintaining the current phase and switching to the next candidate phase (switching to the next phase in the preset phase sequence). The state transition relation satisfies:

[0041] in, This is the traffic environment state transition function at the intersection. For a moment The target phase action, This represents the random disturbance term in traffic flow.

[0042] The reward function is used to measure the combined impact of the current action on intersection traffic efficiency, average delay, and control stability.

[0043] In a specific implementation, the action space is preferably a binary action space. ,in This indicates that the current phase is maintained. This indicates a switch to the next phase in the preset sequence; in a further improved implementation, the action space can be expanded to:

[0044] in, Indicates switching to the first One candidate phase, This means extending the current green light while satisfying the maximum green constraint; all of the above actions must be executed after being filtered by constraints such as minimum / maximum green, conflict phase, switching anti-shake and protection mechanisms.

[0045] In one embodiment, the reinforcement learning control model is preferably a deep Q-network-based reinforcement learning model, specifically DQN or DDQN. In extended embodiments, other reinforcement learning models such as PPO can also be used.

[0046] When using DDQN, the network parameters are evaluated and denoted as follows: The target network parameters are denoted as The system is based on the current unified intersection state vector Output each candidate action Value, and select The action with the largest value is the candidate phase action. :

[0047] If DDQN is used, the target value Represented as:

[0048] in, This is a discount factor, used to characterize the weight of future returns in the current target value, taking a value of 0 ≤ <1 (for example, it can be taken) = 0.99); For a moment The reward value obtained after performing an action.

[0049] loss function Represented as:

[0050] in, This is the expected value; During the training phase, an ε-greedy strategy is preferred to balance exploration and exploitation, and an experience replay mechanism is used to process state transition samples. Perform random sampling training. If the current phase is... When a candidate action indicates a switch, the target phase can be updated according to a preset phase sequence:

[0051] in, This represents the next phase mapping function. For a four-phase control scenario, the preset phase sequence can preferably be 0→2→1→3→0 or its equivalent safe release sequence.

[0052] In one embodiment, the reward function can be in the form of a negative sum of squared queue values ​​to quickly drive the model to suppress queue backlog, i.e.:

[0053] in, For a moment The reward value obtained after performing the action. These represent the number of vehicles queuing in the four directions: east, south, west, and north, respectively.

[0054] In one embodiment, the reward function adopts a multi-objective weighted form, including at least a queue length penalty, a waiting time penalty, a throughput reward, and a phase switching cost penalty; as well as at least one of a multi-directional flow balancing term, a pressure term, or an anomaly penalty term; In one embodiment, the reward function of the reinforcement learning control model in S3 during the training process adopts a multi-objective weighted form, including at least a queue length penalty, a waiting time penalty, a throughput reward, a phase switching cost penalty, a multi-directional flow balancing penalty, and an anomaly penalty. The reward function satisfies:

[0055] Where α, β, χ, δ, μ, and λ are non-negative weighting coefficients. Indicates time Instant reward value, This is a penalty for queue length. For waiting time penalty items, This is a traffic volume reward item; Phase switching cost term satisfy:

[0056] Multi-directional flow balancing item satisfy:

[0057] Abnormal penalty items The value is 1 when an abnormality, illegal action, or continuous oscillation is detected; otherwise, it is 0.

[0058] In a preferred embodiment, the reward function includes at least a queue length penalty, a waiting time penalty, a throughput reward, a phase switching cost penalty, a multi-directional flow balancing penalty, a pressure penalty, and an anomaly penalty; the reward function satisfies:

[0059] Among them, α, β, χ, δ, μ, λ, and These are non-negative weighting coefficients; This indicates the number of vehicles passing the stop line within the statistics window in direction d; The phase switching cost term satisfies:

[0060] Multi-directional flow balancing terms satisfy:

[0061] in, It is a variance operator used to characterize the degree of dispersion of queue length in each direction.

[0062] Furthermore, when considering the balance between the inlet and outlet channels, the pressure term can be defined as:

[0063] in, Indicates direction At any moment Traffic pressure on the import lanes. Indication and direction The corresponding traffic pressure at the exit lane can be represented by the number of vehicles queuing at the corresponding entrance and exit lanes, respectively.

[0064] And adopt total pressure (The sum of pressure terms in all inlet directions) in the form of:

[0065] Abnormal penalty items The value is 1 when an abnormality, illegal action, or continuous oscillation is detected; otherwise, it is 0.

[0066] No release instruction quantity If a vehicle is not allowed to pass during the control cycle due to a transition phase such as a yellow light or a completely red light, then the value is 1; otherwise, the value is 0.

[0067] In a preferred embodiment, the basic DQN comparison group can use parameters α=1.0, β=0.2, χ=1.0, δ=5.0, μ=0.2, λ=20.0, κ=0.0, and η=0.0, without enabling anti-oscillation and protection mechanisms; while the improved method of the present invention can use α=1.0, β=0.15, χ=3.0, δ=8.0, μ=0.1, λ=10.0, κ=0.4, and η=12.0, and combine minimum / maximum green constraints, anti-oscillation constraints, and protection mechanisms to jointly determine the final target phase action. Through the above design, this method achieves a better trade-off between congestion suppression, traffic efficiency, control stability, and directional balance.

[0068] In one embodiment, the motion constraints in S4 include at least: minimum green light duration constraint, maximum green light duration constraint, conflict phase constraint, and switching debounce constraint; The minimum green light duration constraint is satisfied as follows:

[0069] in, The minimum green light duration threshold; The maximum green light duration constraint is satisfied as follows:

[0070] in, This is the maximum green light duration threshold. Conflict phase constraints are satisfied: At that time, determine the phase of the candidate target. These are illegal actions, among which, To pre-determine the conflict phase matrix, The current phase; The switching stabilization constraint satisfies: like And the time interval between two adjacent switching events is less than the threshold. If so, the current phase will be maintained or the protection timing will be switched.

[0071] in, For candidate phase actions, For target phase action, This refers to the current phase (the phase code of the traffic lights at the current intersection). For candidate phase action The corresponding candidate target phase, The phase of the previous control cycle. This is the phase for the next control cycle.

[0072] In one embodiment, the phase change control in S5 is implemented using a finite state machine, which includes at least a green light hold state G, a yellow light transition state Y, a full red clear state AR, and a target phase activation state G′.

[0073] When the target phase corresponding to the target phase action is consistent with the current phase, the state machine remains in the green light hold state (i.e., the state of allowing passage, not referring to the state of a single traffic light, but the state of all current lights), that is:

[0074] in, For the target phase; Indicates time The state of a finite state machine.

[0075] When the target phase corresponding to the target phase action is inconsistent with the current phase, the state machine first transitions from the green light hold state to the yellow light transition state, that is:

[0076] The yellow light transition lasts for a preset duration T_y before entering a fully red clear state, i.e.:

[0077] After the all-red state is cleared for a preset all-red duration T_r, the system switches to the target phase active state, i.e.:

[0078] After the target phase takes effect, the phase duration counter is reset, satisfying the following conditions:

[0079] Finite state machines are suitable for two-way four-lane four-phase control scenarios and can be extended to control scenarios with multiple entrance lanes, multiple phases, or dedicated steering phases.

[0080] In one embodiment, the present invention further includes the step: S6, when an abnormal object such as video abnormality, control abnormality or program abnormality occurs, an abnormal rollback strategy is executed. The abnormal rollback strategy is used to generate an alternative state based on the historical state, switch to the protection timing or restore the control state before the interruption, so as to maintain the continuity of traffic signal control at a single intersection.

[0081] In a specific embodiment, the abnormal rollback strategy in S6 includes perceptual abnormal rollback, decision abnormal rollback and program abnormal rollback; perceptual abnormal rollback is used to generate an alternative state when the video detection result is distorted, missing or unstable, and the video detection result distortion, missing or unstable includes video stream interruption, frame loss, false detection or missed detection or state change. Decision anomaly rollback is used to switch to the protection timing scheme when the reinforcement learning model outputs illegal actions, empty actions, or continuous oscillations. Illegal actions, empty actions, or continuous oscillations include output actions that do not belong to the preset action space, output results that are empty, or actions that repeatedly switch between two phases within adjacent control cycles. Program anomaly rollback is used to restore the historical state and model parameters and continue to execute the control flow after program interruption, restart, or abnormal model termination. Program interruption, restart, or model anomaly includes abnormal process exit, system restart, model parameter loading failure, or abnormal termination of the inference process.

[0082] When the video detection confidence level is below the threshold Frame drop rate higher than the threshold If the statistical results suddenly exceed the threshold, an alternative state is used. Replace the current state ,in:

[0083] in, For smoothing coefficients, This is the length of the historical sliding window; When the reinforcement learning model outputs an illegal action, a no-action, or more than K consecutive oscillations, it switches to a preset protection timing scheme, which has a fixed period. Set of phase green light durations:

[0084] in, These represent the green light duration for each phase in the protection timing scheme. To protect the number of phases in the timing scheme.

[0085] When a program is abnormally interrupted, the state restored from persistent storage includes at least the following:

[0086] in, The phase state at the moment of interruption. The duration of the current phase. To enhance the learning model parameters, For experience replay caching, This is historical statistical data.

[0087] Through the above-mentioned three-level anomaly fallback mechanism, this embodiment can maintain the continuity, stability and recoverability of traffic signal control at a single intersection under various abnormal conditions.

[0088] In one specific embodiment, to verify the control effect of the aforementioned single-intersection traffic signal control method based on video detection results in a single-intersection scenario, a single-intersection traffic signal control simulation environment corresponding to steps S1 to S5 is constructed. This simulation environment does not directly access real video, but generates equivalent lane-level traffic parameters consistent with the output of step S1 based on the arrival, queuing, and release processes of vehicles from each approach direction, in order to simulate the input results after video perception mapping.

[0089] In this specific embodiment, the intersection is modeled according to the four entrance directions of east, south, west and north, and a dual-lane input modeling method is adopted; two equivalent lanes are set for each entrance direction, one of which is a straight and right turn equivalent lane and the other is a left turn equivalent lane, forming a total of 8 equivalent lane queues.

[0090] The system operates iteratively in discrete control cycles. Each control cycle executes the following processes sequentially: First, the arriving vehicles in each approach direction are updated; then, the number of queuing vehicles, waiting time, and traffic flow of a preset duration in each equivalent lane are counted to form corresponding lane-level traffic parameters; next, the traffic parameters of each lane are aggregated according to the approach direction, and a unified intersection state vector is constructed by combining the current signal phase and the current phase duration; then, the unified intersection state vector is input into the reinforcement learning control model to obtain candidate phase actions; then, minimum green light duration constraints, maximum green light duration constraints, conflict phase constraints, and switching anti-jitter constraints are applied to the candidate phase actions to obtain the target phase action; finally, based on the target phase action, the green light is maintained, yellow light transition, all-red light clearing, and target phase activation are executed through a finite state machine, and the queue state of each lane is updated according to the phase release result, and the reward value of the current control cycle is calculated for model training or performance evaluation.

[0091] With the above settings, the state perception process in the simulation environment corresponds to steps S1 and S2, the action decision process corresponds to step S3, the action constraint process corresponds to step S4, and the phase execution process corresponds to step S5; when the exception handling mechanism is enabled in the experiment, the exception rollback process corresponds to step S6.

[0092] In the experimental setup, a four-phase cyclic sequence 0→2→1→3→0 was used. Phase 0 corresponds to allowing straight-ahead and right-turn traffic in east-west directions, phase 2 corresponds to allowing straight-ahead and right-turn traffic in north-south directions, phase 1 corresponds to allowing left-turn traffic in north-south directions, and phase 3 corresponds to allowing left-turn traffic in east-west directions. The transition phases were the yellow light and the all-red light phase. The yellow light duration was set to... =33 steps, all-red duration set to =2 steps; Fixed timing comparison group fixed green light duration set to 2 steps. =20 steps. The inflow from the four directions is generated independently and randomly at each time step. The traffic volume from the east, south, west, and north directions is a uniformly random integer within the interval [1, 3]. The green light capacity is set to... =4 vehicles / step, with 0 vehicles allowed during yellow light and all-red light phases.

[0093] To ensure reproducibility, both training and evaluation were performed using a fixed set of random seeds. Preferably, the training steps were set to 12,000 steps, the evaluation steps to 600 steps, and the evaluation random seed was {1, 2, 3}. The base DQN network used 5-dimensional state features and a 2-dimensional action space, with a learning rate of 0.01, an experience replay capacity of 2000, a target network update cycle of 100, a discount factor of 0.99, and a batch size of 32.

[0094] To objectively demonstrate the improvement of the method of this invention compared to traditional schemes and basic reinforcement learning schemes, this embodiment sets up three comparison groups: The first is a fixed timing scheme. This scheme uses a fixed green light plus yellow light plus all red light, without considering the real-time queue status; The second is based on the DQN scheme. This scheme only learns the "hold / switch" action, and phase execution uses a basic finite state machine, but does not enable anti-oscillation constraints, protection mechanisms, and advanced reward terms; The third is an improved method of the present invention, which introduces phase controller constraints, compound rewards, state smoothing and anomaly detection mechanisms on the basis of DQN.

[0095] Evaluation metrics include: average queue length Average number of vehicles passing through Number of phase switching and optional anomaly count Among these, a smaller average queue length is better, a larger average number of vehicles passing through is better, and the number of phase switching is used to evaluate the control stability while ensuring efficiency.

[0096] Under the conditions of 600 evaluation steps and a random seed of {1, 2, 3}: The average queue length for the fixed timing scheme is 274.419±6.832, the average number of vehicles passing through is 7.323±0.038, and the number of phase switching is 25.000±0.000. The average queue length of the basic DQN scheme is 358.460±6.969, the average number of vehicles passing through is 7.409±0.043, the number of phase switching is 21.000±0.000, and the anomaly count is 0.000±0.000. The improved method of this invention has an average queue length of 219.916±9.174, an average number of vehicles passing through of 7.589±0.063, a phase switching count of 27.667±0.471, and an anomaly count of 0.000±0.000.

[0097] The results above show that, compared with the traditional fixed timing scheme, the improved method of this invention reduces the average queue length by about 19.86%, increases the average number of vehicles passing by about 3.65%, and significantly improves directional balance. Compared with the basic DQN scheme, the improved method of this invention reduces the average queue length by about 38.65%, increases the average number of vehicles passing by about 2.43%, and significantly reduces the queuing variance between directions, indicating that the method of this invention can achieve a better trade-off between efficiency, congestion, and stability.

[0098] The experimental results above demonstrate that, by unifying the detection results from four video streams into a unified intersection control state, and by introducing a multi-objective reward function, action constraints, finite state machine phase execution, and an anomaly backoff mechanism, the method of this invention can effectively suppress queue growth at single intersections and improve traffic efficiency per unit time. Simultaneously, it avoids the degradation phenomenon that occurs in basic reinforcement learning schemes under insufficient constraints or imperfect rewards. Therefore, this invention is not only feasible but also possesses clear and verifiable optimization effects in single-intersection traffic signal control scenarios.

[0099] In one embodiment, such as Figure 2 As shown, the present invention also provides a single-intersection traffic signal control system based on video detection results, comprising: The data acquisition and processing module collects video data and signal phase information from the four directions of the intersection, and performs perception mapping on the video data of each lane to obtain lane-level traffic parameters for each lane. The state construction module, based on the lane-level traffic parameters of each lane and the correspondence between each lane and its respective approach direction, aggregates them according to the approach direction and combines them with the signal phase information to construct a unified intersection state vector. The unified intersection state vector includes at least the direction-level queue length and the current signal phase of each approach direction. The decision module inputs the unified intersection state vector into a pre-trained reinforcement learning control model to obtain candidate phase actions; The constraint processing module applies action constraints to the candidate phase actions to obtain the target phase action; The phase execution module performs phase change control according to the target phase action.

[0100] In a further embodiment, it also includes: The abnormal rollback module is used to execute abnormal rollback strategies when video, control, or program abnormalities occur, in order to maintain the continuity of traffic signal control at a single intersection. The storage module is used to store reinforcement learning control model parameters, lane calibration information, historical statistical data, and phase status.

[0101] The above are merely preferred embodiments of the present invention. It should be noted that the present invention is not limited to the above embodiments. For those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A single intersection traffic signal control method based on video detection results, characterized by, Including the following steps: S1, collect video data and signal phase information from the four directions of the intersection, and perform perception mapping on the video data of each lane to obtain lane-level traffic parameters for each lane; wherein, the lane-level traffic parameters include at least the number of queuing vehicles, waiting time, and traffic flow for a preset duration; S2, based on the lane-level traffic parameters of each lane and the correspondence between each lane and the associated import direction, aggregate according to the import direction and combine with the signal phase information to construct a unified intersection state vector ; the unified intersection state vector at least includes the direction-level queue length, the direction-level waiting time, the direction-level traffic per unit time and the direction-level traffic trend of change, and the current signal phase of each import direction. S3, the unified intersection state vector Input a pre-trained reinforcement learning control model and output candidate phase actions. The reinforcement learning control model is a model that outputs phase actions based on state vectors. Among them, when the video detection confidence of the video data is lower than the confidence threshold The frame drop rate is higher than the frame drop threshold. If the statistical results show a mutation exceeding the mutation threshold, an alternative state is used. Replace the current unified intersection state vector ,in: in, For smoothing coefficients, This is the length of the historical sliding window; When the reinforcement learning model outputs an illegal action, a no-action, or more than K consecutive oscillations, it switches to a preset protection timing scheme, which has a fixed period. Set of phase green light durations: in, These represent the green light duration for each phase in the protection timing scheme. To protect the number of phases in the timing scheme; S4, apply action constraints to the candidate phase action to obtain the target phase action; wherein, the target phase action is a phase action determined under the action constraints; the action constraints include at least: minimum green light duration constraint, maximum green light duration constraint, conflict phase constraint, and handover anti-shake constraint; S5, perform phase change control according to the target phase action.

2. The control method according to claim 1, characterized in that, The process of performing perception mapping on the video data from each path in S1 to obtain lane-level traffic parameters for each lane specifically includes: Define the video data from the four directions of the intersection as a set of video streams; in, This represents the set of video streams from the four directions of the intersection. These represent the video streams corresponding to the four import directions; And establish a set of directions: in, This represents the set of directions of the intersection's entrance. These represent the four directions of entry: east, south, west, and north. And the mapping relationship between video and direction For each video stream at any time Image frames Preprocessing is required; The image frames of each video stream The target detection is performed, and the set of vehicle target detection results is obtained as follows; in, For the target bounding box, For the target category, To test the confidence level; Multi-target tracking is performed on the target detection results to obtain vehicle trajectories with trajectory markers; in, For the set of tracking results, For vehicle trajectory identification; Based on pre-defined lane areas, stop line areas, and cross-line judgment lines, the trajectories of each vehicle are... Map to the corresponding lane, and determine lane affiliation and crossing status based on whether the target center point enters the corresponding area or the overlap between the target bounding box and the corresponding area; Within a preset time period Δt, under the condition that the same trajectory marker is counted only once, the number of queuing vehicles, waiting time and number of vehicles crossing the line in each lane are counted, and the traffic flow of each lane is obtained based on the number of vehicles crossing the line.

3. The control method according to claim 2, characterized in that, In S2, the construction of the unified intersection state vector is specifically as follows: For any direction ,set up This represents the set of lanes corresponding to that direction, and the direction-level queue length for that direction. Directional waiting time Directional unit time flow and directional flow change trend They respectively satisfy: in, For a moment No. The number of vehicles queuing in each lane; For a moment No. Waiting time statistics for each lane Preset duration Inner Traffic flow in each lane; The unified intersection state vector satisfy: in, This is the number of the current phase. The duration of the current phase. Mark the validity of the video.

4. The control method according to claim 3, characterized in that, In the reinforcement learning control model of S3, during the training process, the single-intersection signal control problem is abstracted into a Markov decision process for signal control, expressed as: in, For state space, For the action space, This is a state transition relation. For the reward function, Discount factor; The states in the state space are represented by a unified intersection state vector. The construction of the unified intersection state vector includes at least a lane-level statistics layer, a direction-level aggregation layer, and a control state layer. The lane-level statistics layer is used to receive lane-level traffic parameters of each lane. The direction-level aggregation layer is used to aggregate the lane-level traffic parameters of each lane according to the approach direction to obtain the direction-level traffic state quantities of each direction. The control state layer is used to combine the direction-level traffic state quantities of each direction with signal phase information to construct the unified intersection state vector. The action space consists of signal phase actions, including at least maintaining the current phase and switching to the next candidate phase; The state transition relationship satisfies: in, This is the traffic environment state transition function at the intersection. For a moment The target phase action, This represents the random disturbance term in traffic flow.

5. The control method according to claim 3, characterized in that, When using dual-lane input modeling, the directional queue length in each direction is obtained by aggregating the statistics of the two lanes in the corresponding direction.

6. The control method according to claim 4, characterized in that, The reinforcement learning control model in S3 adopts a multi-objective weighted reward function during the training process, which includes at least a queue length penalty, a waiting time penalty, a throughput reward, a phase switching cost penalty, a multi-directional flow balancing penalty, and an anomaly penalty. The reward function satisfies: Where α, β, χ, δ, μ, and λ are non-negative weighting coefficients; This is a penalty for queue length. For waiting time penalty items, This is a traffic volume reward item; Indicates time Instant reward value, For the aforementioned abnormal penalty item; The phase switching cost item satisfy: The multi-directional flow balancing item satisfy: 。 7. The control method according to claim 6, characterized in that, The minimum green light duration constraint satisfies: The maximum green light duration constraint satisfies: The conflict phase constraint satisfies: At that time, determine the phase of the candidate target. These are illegal actions, among which, To pre-determine the conflict phase matrix, The current phase; The switching stabilization constraint satisfies: like And the time interval between two adjacent switching events is less than the threshold. If so, maintain the current phase or switch to protection timing; in, For candidate phase actions, For target phase action, For the current phase, For candidate phase action The corresponding candidate target phase, This indicates that the current phase is maintained. This is the maximum green light duration threshold. The minimum green light duration threshold. The phase of the previous control cycle. This is the phase for the next control cycle.

8. The control method according to claim 1, characterized in that, The phase change control in S5 is used to perform phase holding or phase switching on each traffic light group at the intersection according to the target phase action, specifically including: When the target phase corresponding to the target phase action is consistent with the current phase, the green light hold state is maintained; When the target phase corresponding to the target phase action is inconsistent with the current phase, it goes through the yellow light transition state and the all-red clear state in sequence before switching to the target phase effective state, and resets the current phase duration count after the target phase takes effect.

9. A single-intersection traffic signal control system based on video detection results, characterized in that, include: The acquisition and processing module acquires video data and signal phase information from four directions of the intersection, and performs perception mapping on the video data of each direction to obtain lane-level traffic parameters for each lane; wherein, the lane-level traffic parameters include at least the number of queuing vehicles, waiting time, and traffic flow for a preset duration; The state construction module, based on the lane-level traffic parameters of each lane and the correspondence between each lane and its respective approach direction, aggregates them according to the approach direction and combines them with signal phase information to construct a unified intersection state vector. The unified intersection state vector includes at least the direction-level queue length, direction-level waiting time, direction-level flow rate per unit time, direction-level flow rate change trend, and current signal phase for each approach direction. The decision module inputs the unified intersection state vector into a pre-trained reinforcement learning control model to obtain candidate phase actions; Among them, when the video detection confidence of the video data is lower than the confidence threshold The frame drop rate is higher than the frame drop threshold. If the statistical results show a mutation exceeding the mutation threshold, an alternative state is used. Replace the current unified intersection state vector ,in: in, For smoothing coefficients, This is the length of the historical sliding window; When the reinforcement learning model outputs an illegal action, a no-action, or more than K consecutive oscillations, it switches to a preset protection timing scheme, which has a fixed period. Set of phase green light durations: in, These represent the green light duration for each phase in the protection timing scheme. To protect the number of phases in the timing scheme; The constraint processing module applies action constraints to candidate phase actions to obtain target phase actions; wherein, the target phase action is a phase action determined under the action constraints; the action constraints include at least: minimum green light duration constraint, maximum green light duration constraint, conflict phase constraint, and handover anti-shake constraint; The phase execution module performs phase change control according to the target phase action.