A traffic signal control method based on a neuro-symbolic system and reinforcement learning
By integrating neural symbol system and reinforcement learning in the intelligent transportation system and building a traffic light control method, the problems of uninterpretation and poor environmental adaptability of the intelligent transportation system are solved, and higher transparency and real-time adaptability are achieved.
Patent Information
- Application Number
- CN202510495235.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Intelligent transportation systems face problems of inexplicability and poor environmental adaptability, and the prior art has not yet effectively integrated the advantages of neural symbolic systems and reinforcement learning to solve these problems.
The traffic light control method based on neural symbol system and reinforcement learning is adopted, and the video stream is processed through convolutional neural networks to construct traffic flow feature vectors, and the deep learning features are mapped into interpretable traffic rules using neural symbol systems, and combined with reinforcement learning, optimized signal light timing strategy.
It improves the transparency and real-time adaptability of traffic light control, enhances the interpretability and environmental adaptability of the system, and ensures the safety and effectiveness of the strategy.
Smart Images

Figure CN120014842B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation systems, and particularly to a traffic signal control method based on a neuro-symbolic system and reinforcement learning. Background Art
[0002] The main challenges faced by current intelligent transportation systems include the interpretability problem caused by the "black box" characteristics of models, which makes it difficult to provide clear decision-making basis for traffic managers. At the same time, traditional methods show poor environmental adaptability when dealing with sudden traffic events or urban structure changes. Although the neuro-symbolic system can enhance the interpretability of the system by combining deep learning with symbolic logic rules, and reinforcement learning can optimize decision-making strategies through a trial-and-error mechanism, the existing technologies have not effectively integrated the advantages of both to solve the above-mentioned problems.
[0003] Therefore, in view of the above problems, a traffic signal control method based on a neuro-symbolic system and reinforcement learning is provided. Summary of the Invention
[0004] The purpose of the present invention is to provide a traffic signal control method based on a neuro-symbolic system and reinforcement learning to overcome the existing defects. By integrating the perception ability of deep learning, the interpretability of symbolic logic, and the dynamic optimization ability of reinforcement learning, the transparency and real-time adaptability of traffic signal control are improved.
[0005] The technical solution to achieve the above purpose is as follows:
[0006] A traffic signal control method based on a neuro-symbolic system and reinforcement learning includes:
[0007] Step S1, real-time obtain vehicle data through cameras, geomagnetic sensors and GPS, collect weather information through meteorological sensors, obtain weekday, weekend and peak period identifiers through a server, and obtain traffic signal information through a signal controller at the same time;
[0008] Step S2, use a convolutional neural network to process the camera video stream, output the real-time vehicle density and speed distribution of each lane, and construct a traffic flow feature vector using sensor data;
[0009] Step S3, use a neuro-symbolic system to map deep learning features into interpretable traffic rules, and generate constraint conditions in combination with traffic regulations;
[0010] Step S4, construct the state and action spaces of reinforcement learning, design a reward function based on traffic rules and constraint conditions, and use a double deep Q network to optimize the traffic signal timing strategy.
[0011] Preferably, in the step S1, the vehicle data includes but is not limited to traffic flow, vehicle speed, queue length, and weather conditions.
[0012] Preferably, in the step S2, the constructed traffic flow feature vectors include but are not limited to:
[0013]
[0014] where is the set of traffic flow feature vectors of vehicle density, is the th traffic flow feature vector of vehicle density, is the set of traffic flow feature vectors of vehicle speed, is the th traffic flow feature vector of vehicle speed, is the set of traffic flow feature vectors of queue length, is the th traffic flow feature vector of queue length.
[0015] Preferably, in the step S3, mapping the deep learning features to interpretable traffic rules. For vehicle density, that is:
[0016] If the density of the eastbound lane exceeds the threshold, then trigger the extension of the green light. The rule is formally expressed as:
[0017] ;
[0018] In the formula, is the lane density threshold;
[0019] If the average speed of vehicles in a certain direction lane is lower than the preset threshold, then trigger the green light priority strategy to improve the traffic efficiency; otherwise, maintain the current signal timing strategy:
[0020] ;
[0021] In the formula, is the th speed threshold of the lane.
[0022] If the queue length of a certain direction lane exceeds the preset threshold, then trigger the extension or priority strategy of the green light to relieve congestion; otherwise, maintain the current signal timing strategy:
[0023] ;
[0024] In the formula, is the th queue length threshold of the lane;
[0025] Therefore, it is necessary to comprehensively consider multiple factors such as vehicle density, speed, and queue length. The formal representation of the comprehensive rules is as follows:
[0026] 。
[0027] Preferably, in the step S3, the constraint conditions generated in combination with traffic regulations include, but are not limited to, not allowing a certain direction of vehicles to pass through for two consecutive red-light cycles and the north-south direction must be red when the east-west direction is passing.
[0028] Preferably, in the step S4, a state space of reinforcement learning is constructed, and the state includes:
[0029] The current traffic flow characteristics, that is: 、 、 ;
[0030] The historical signal timings, that is: the green-light durations of the last 3 cycles, that is:
[0031] ;
[0032] In the formula, is the green-light duration of the th cycle;
[0033] Weather and time information, where the weather information includes, but is not limited to: weather type, rainfall, visibility, wind speed, road surface condition, that is:
[0034] ;
[0035] In the formula, represents the weather type, represents the rainfall, represents the visibility, represents the wind speed, represents the road surface condition;
[0036] The time information includes, but is not limited to: weekday peak hours and weekend off-peak hours, that is:
[0037] ;
[0038] In the formula, is 1 when it represents a weekday and 0 when it represents a weekend, is 1 when it represents a peak hour and 0 when it represents an off-peak hour;
[0039] The final state vector is expressed as ;
[0040] Construct the action space of reinforcement learning, and the action including but not limited to: green light duration adjustment and lane priority allocation, where
[0041] Green light duration adjustment: Dynamically adjust the green light duration in each direction, that is:
[0042] ;
[0043] In the formula, represents the green light duration adjustment value of the th direction;
[0044] Lane priority allocation: Dynamically allocate the priority of each lane according to real-time traffic demand, affecting the weight of green light duration allocation, that is:
[0045] ;
[0046] Among them, represents the priority of the th lane. When it represents low priority, when
[0047] it represents medium priority,
[0048] Action space ;
[0049] Reward function design, that is:
[0050] Set to minimize the average waiting time, where is the reward for minimizing the average waiting time, is the number of vehicles waiting to pass through the intersection, is the time interval from when the vehicle arrives at the intersection to when it starts to pass through the intersection;
[0051] Set a safety penalty term:
[0052] If the action causes the green lights of vehicles in conflicting directions to turn on simultaneously, impose a penalty where is the reward for action conflict;
[0053] If the queue length exceeds the threshold impose a penalty is the reward for queue length;
[0054] If the average speed is lower than the threshold impose a penalty where is the reward for average speed, is the average speed of the current lane;
[0055] If the weather is rainy and the green light duration is less than 30 seconds, impose a penalty , where is the rainy day reward;
[0056] If the visibility is less than 200 meters and the green light duration is less than 30 seconds, impose a penalty , where is the low visibility reward;
[0057] That is, the comprehensive reward function:
[0058] ;
[0059] In the formula, is the weight coefficient, is the smoothness reward;
[0060] Use a double deep Q-network to mitigate the overestimation bias and introduce prioritized experience replay to accelerate convergence.
[0061] Preferably, in step S4, after each signal light cycle ends, store the current state , action , and reward in the experience replay pool, regularly sample batch data from the pool, and update the parameters of the double deep Q-network through backpropagation.
[0062] The beneficial effects of the present invention are as follows: The present invention explicitly expresses the decision-making logic through symbolic rules, which is convenient for traffic managers to understand the basis of the strategy. Online reinforcement learning enables the system to dynamically adjust the strategy to cope with traffic flow mutations, improving the real-time adaptability. Through symbolic rule constraints, the generation of strategies that violate traffic regulations is avoided, enhancing the safety guarantee. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a flowchart of a traffic signal control method based on a neuro-symbolic system and reinforcement learning according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0064] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0065] The present invention will be further described below in conjunction with the accompanying drawings.
[0066] As Figure 1 shown, a traffic signal control method based on a neuro-symbolic system and reinforcement learning includes:
[0067] Step S1, real-time vehicle data is obtained through cameras, geomagnetic sensors, and GPS, weather information is collected through meteorological sensors, weekday, weekend, and peak-hour identifiers are obtained through a server, and at the same time, signal light information is obtained through a signal light controller.
[0068] In the embodiment, the vehicle data includes, but is not limited to, traffic flow, vehicle speed, queue length, and weather conditions.
[0069] Step S2, a convolutional neural network is used to process the camera video stream, the real-time vehicle density and speed distribution of each lane are output, and a traffic flow feature vector is constructed using the sensor data.
[0070] In the embodiment, the constructed traffic flow feature vector includes, but is not limited to:
[0071]
[0072] Among them, is the set of traffic flow feature vectors of vehicle density, is the th traffic flow feature vector of vehicle density, is the set of traffic flow feature vectors of vehicle speed, is the th traffic flow feature vector of vehicle speed, is the set of traffic flow feature vectors of queue length, is the th traffic flow feature vector of queue length.
[0073] Step S3, a neuro-symbolic system is used to map the deep learning features into interpretable traffic rules, and constraint conditions are generated in combination with traffic regulations.
[0074] In the embodiment, when mapping the deep learning features into interpretable traffic rules, for vehicle density, that is:
[0075] If the density of the eastbound lane exceeds the threshold, then the green light extension is triggered, and the rule is formally expressed as:
[0076] ;
[0077] In the formula, is the lane density threshold;
[0078] If the average speed of vehicles in a lane in a certain direction is lower than the preset threshold, the green light priority strategy is triggered to improve traffic efficiency; otherwise, the current signal timing strategy is maintained:
[0079] ;
[0080] In the formula, is the speed threshold of the th lane.
[0081] If the queue length of a lane in a certain direction exceeds the preset threshold, the green light extension or priority strategy is triggered to relieve congestion; otherwise, the current signal timing strategy is maintained:
[0082] ;
[0083] In the formula, is the queue length threshold of the th lane;
[0084] Therefore, it is necessary to comprehensively consider multiple factors such as vehicle density, speed, and queue length. The formal representation of the comprehensive rule is as follows:
[0085] .
[0086] In the embodiment, the constraint conditions generated in combination with traffic regulations include but are not limited to not allowing a certain direction of vehicles to be released without passing through two consecutive red light cycles and the north-south direction must be red when the east-west direction is passing.
[0087] Step S4, construct the state and action spaces of reinforcement learning, design a reward function based on traffic rules and constraint conditions, and use a double deep Q-network to optimize the signal timing strategy.
[0088] In the embodiment, construct the state space of reinforcement learning, and the state includes:
[0089] The current traffic flow characteristics, that is: , , ;
[0090] The historical signal timing, such as the green light duration of the last 3 cycles, that is:
[0091] ;
[0092] In the formula, is the green light duration of the th cycle;
[0093] Weather and time information, where the weather information includes but is not limited to: weather type, rainfall, visibility, wind speed, road surface condition, that is:
[0094] ;
[0095] In the formula, represents the weather type, represents the rainfall, represents the visibility, represents the wind speed, represents the state of the road surface;
[0096] The time information includes but is not limited to: weekday peak hours and weekend off-peak hours, that is:
[0097] ;
[0098] In the formula, When it is 1, it represents a weekday, and when it is 0, it represents a weekend, When it is 1, it represents peak hours, and when it is 0, it represents off-peak hours;
[0099] The final state vector is expressed as ;
[0100] Construct the action space of reinforcement learning, and the action includes but is not limited to: green light duration adjustment and lane priority allocation, where,
[0101] Green light duration adjustment: Dynamically adjust the green light duration in each direction, that is:
[0102] ;
[0103] In the formula, represents the adjustment value of the green light duration in the th direction;
[0104] Lane priority allocation: Dynamically allocate the priority of each lane according to real-time traffic demand, affecting the weight of green light duration allocation, that is:
[0105] ;
[0106] Among them, represents the priority of the th lane, When it is, it represents low priority, When it is, it represents medium priority, When it is, it represents high priority;
[0107] Then:
[0108] The action space ;
[0109] Design of the reward function, that is:
[0110] Set the minimum average waiting time, , where is the minimum average waiting time reward, is the number of vehicles waiting to pass through the intersection, is the time interval from when the vehicle arrives at the intersection to when it starts to pass through the intersection;
[0111] Set the safety penalty term:
[0112] If the action causes the green lights of conflicting-direction vehicles to turn on simultaneously, impose a penalty , where is the action conflict reward;
[0113] If the queue length exceeds the threshold , impose a penalty , is the queue length reward;
[0114] If the average speed is lower than the threshold , impose a penalty , where is the average speed reward, is the average speed of the current lane;
[0115] If the weather is rainy and the green light duration is less than 30 seconds, impose a penalty , where is the rainy weather reward;
[0116] If the visibility is less than 200 meters and the green light duration is less than 30 seconds, impose a penalty , where is the low visibility reward;
[0117] That is, the comprehensive reward function:
[0118] ;
[0119] In the formula, is the weight coefficient, is the smoothness reward (to avoid frequent signal light switching);
[0120] Use a double deep Q-network to mitigate the overestimation bias and introduce prioritized experience replay to accelerate convergence.
[0121] In the embodiment, after each signal light cycle ends, the current state , action , and reward are stored in the experience replay pool, and a batch of data is sampled from the pool regularly, and the parameters of the double deep Q-network are updated through backpropagation.
[0122] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A traffic light control method based on neural symbolic system and reinforcement learning, characterized in that: include: Step S1, real-time acquisition of vehicle data through a camera, geomagnetic sensor and GPS, weather information through a meteorological sensor, identification of working days, weekends and peak hours through a server, and signal light information through a signal light controller; Step S2, using a convolutional neural network to process the camera video stream, output the real-time vehicle density and speed distribution of each lane, and use the sensor data to construct a traffic flow feature vector; Step S3, using a neural symbolic system to map deep learning features into interpretable traffic rules, and generating constraints in combination with traffic regulations; Step S4, construct the state and action space of reinforcement learning, design the reward function based on traffic rules and constraints, and use a dual deep Q network to optimize the traffic light timing strategy; In step S4, the state space of reinforcement learning is constructed. Include: Current traffic flow characteristics, namely: , , ; is the set of vehicle density traffic flow feature vectors, is the vehicle speed traffic flow feature vector set, is the set of queue length traffic flow feature vectors; Historical traffic light timing, that is, the green light duration of the last three cycles, that is: ; In the formula, For the The green light duration of each cycle; Weather information, including but not limited to: weather type, rainfall, visibility, wind speed, road conditions, namely: ; In the formula, Indicates the weather type. Indicates rainfall, Indicates visibility, Indicates wind speed, Indicates the condition of the road surface; Time information includes, but is not limited to: weekday peak hours and weekend off-peak hours, i.e.: ; In the formula, When it is 1, it indicates a weekday, and when it is 0, it indicates a weekend. When it is 1, it indicates peak hours, and when it is 0, it indicates non-peak hours; The final state vector is expressed as ; Constructing the action space of reinforcement learning, action Including but not limited to: green light duration adjustment and lane priority allocation, among which, Green light duration adjustment: Dynamically adjust the green light duration in each direction, namely: ; In the formula, Indicates Green light duration adjustment value for each direction; Lane priority allocation: Dynamically allocate the priority of each lane according to real-time traffic demand, affecting the weight of green light duration allocation, namely: ; in, Indicates The priority of each lane, , indicating low priority. , indicating medium priority. , indicating high priority; but: Action Space ; Reward function design, namely: Set the minimum average waiting time, ,in, To minimize the average waiting time reward, is the number of vehicles waiting to pass through the intersection, The time interval from when a vehicle arrives at an intersection to when it begins to pass through the intersection; To set security penalties: If the action causes the green lights of vehicles in the conflicting directions to turn on at the same time, the penalty ,in, Reward for action conflict; If the queue length exceeds the threshold ,punish , Rewards for queue length; If the average speed is lower than the threshold ,punish ,in is the average speed reward, is the average speed of the current lane; If it is rainy and the green light duration is less than 30 seconds, the penalty ,in For rainy day rewards; If visibility is less than 200 meters and the green light duration is less than 30 seconds, the penalty ,in Reward for low visibility; That is, the comprehensive reward function: ; In the formula, is the weight coefficient, Reward for smoothness; A dual-depth Q network is used to alleviate over-estimation bias, and priority experience replay is introduced to accelerate convergence.
2. A traffic light control method based on neural symbolic system and reinforcement learning according to claim 1, characterized in that: In step S1, the vehicle data includes but is not limited to vehicle flow, vehicle speed, queue length and weather conditions.
3. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 1, characterized in that: In step S2, the traffic flow feature vector constructed includes but is not limited to: ; ; ; in, For the Vehicle density traffic flow feature vector, For the Vehicle speed traffic flow feature vector, For the The queue length traffic flow feature vector.
4. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 3 is characterized in that: In step S3, the deep learning features are mapped into interpretable traffic rules. For vehicle density, that is: If the density of the eastbound lane exceeds the threshold, the green light extension is triggered. The rule is formalized as follows: ; In the formula, is the lane density threshold; If the average speed of vehicles in a certain direction lane is lower than the preset threshold, the green light priority strategy is triggered to improve traffic efficiency; Otherwise, maintain the current signal light timing strategy: ; In the formula, For the Speed threshold for each lane; If the queue length in a certain direction exceeds the preset threshold, a green light extension or priority strategy is triggered to ease congestion; Otherwise, maintain the current signal light timing strategy: ; In the formula, For the The queue length threshold for each lane; Therefore, it is necessary to comprehensively consider multiple factors such as vehicle density, speed and queue length. The formal expression of the comprehensive rule is as follows: 。 5. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 1, characterized in that: In step S3, the constraint conditions generated in combination with traffic regulations include but are not limited to prohibiting vehicles in a certain direction from being released for two consecutive red light cycles and requiring a red light in the north-south direction when traveling in the east-west direction.
6. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 1, characterized in that: In step S4, after each signal light cycle ends, the current state ,action ,award The experience is stored in the replay pool, batch data is sampled from the pool regularly, and the parameters of the dual-depth Q network are updated through back-propagation.
Citation Information
Patent Citations
Real-time recommendation method for traffic signal control scheme based on deep learning
CN110491146A
Traffic signal timing optimization method based on deep reinforcement learning
CN112700664A