Traffic signal lamp control method based on neural symbol system and reinforcement learning
By integrating neural symbol system and reinforcement learning in the intelligent transportation system, the problems of uninterpretation and poor environmental adaptability of intelligent transportation systems are solved, and the transparency and real-time adaptability of traffic light control are achieved.
Patent Information
- Application Number
- CN202510495235.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Intelligent transportation systems face problems of unexplainability and poor environmental adaptability. The existing technology has not yet effectively integrated neural symbology and reinforcement learning to solve these problems.
The traffic light control method based on neural symbol system and reinforcement learning is adopted, and the video stream is processed through convolutional neural networks. The neural symbol system maps deep learning features as interpretable traffic rules, and combines reinforcement learning to optimize the signal light timing strategy.
It improves the transparency and real-time adaptability of traffic light control, provides clear decision-making basis, and enhances the safety and environmental adaptability of the system.
Smart Images

Figure CN120014842A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation systems, and in particular to a traffic light control method based on a neural symbolic system and reinforcement learning. Background Art
[0002] The main challenges facing the current intelligent transportation system include the unexplainability caused by the "black box" characteristics of the model, which makes it difficult to provide traffic managers with clear decision-making basis. At the same time, traditional methods show poor environmental adaptability when dealing with sudden traffic incidents or changes in urban structure. Although neural symbolic systems can enhance the interpretability of the system by combining deep learning with symbolic logic rules, and reinforcement learning can optimize decision-making strategies through trial and error mechanisms, existing technologies have not yet effectively integrated the advantages of the two to solve the above-mentioned problems.
[0003] Therefore, in response to the above problems, a traffic light control method based on neural symbolic system and reinforcement learning is provided. Summary of the invention
[0004] The purpose of the present invention is to overcome the existing defects and provide a traffic light control method based on neural symbolic system and reinforcement learning, which improves the transparency and real-time adaptability of traffic light control by integrating the perception ability of deep learning, the interpretability of symbolic logic and the dynamic optimization ability of reinforcement learning.
[0005] The technical solution to achieve the above purpose is: A traffic light control method based on neural symbolic system and reinforcement learning, comprising: Step S1, real-time acquisition of vehicle data through a camera, geomagnetic sensor and GPS, weather information through a meteorological sensor, identification of working days, weekends and peak hours through a server, and signal light information through a signal light controller; Step S2, using a convolutional neural network to process the camera video stream, output the real-time vehicle density and speed distribution of each lane, and use the sensor data to construct a traffic flow feature vector; Step S3, using a neural symbolic system to map deep learning features into interpretable traffic rules, and generating constraints in combination with traffic regulations; Step S4, construct the state and action space of reinforcement learning, design the reward function based on traffic rules and constraints, and use the dual deep Q network to optimize the traffic light timing strategy.
[0006] Preferably, in step S1, the vehicle data includes but is not limited to vehicle flow, vehicle speed, queue length and weather conditions.
[0007] Preferably, in step S2, the traffic flow feature vector constructed includes but is not limited to: in, is the set of vehicle density traffic flow feature vectors, For the Vehicle density traffic flow feature vector, is the vehicle speed traffic flow feature vector set, For the Vehicle speed traffic flow feature vector, is the queue length traffic flow feature vector set, For the The queue length traffic flow feature vector.
[0008] Preferably, in step S3, the deep learning features are mapped into interpretable traffic rules. For vehicle density, that is: If the density of the eastbound lane exceeds the threshold, the green light extension is triggered. The rule is formalized as follows: ; In the formula, is the lane density threshold; If the average speed of vehicles in a certain direction is lower than the preset threshold, the green light priority strategy is triggered to improve traffic efficiency; otherwise, the current signal light timing strategy is maintained: ; In the formula, For the The speed threshold for each lane.
[0009] If the queue length of a lane in a certain direction exceeds the preset threshold, the green light extension or priority strategy is triggered to alleviate congestion; otherwise, the current signal light timing strategy is maintained: ; In the formula, For the The queue length threshold for each lane; Therefore, it is necessary to comprehensively consider multiple factors such as vehicle density, speed and queue length. The formal expression of the comprehensive rule is as follows: .
[0010] Preferably, in step S3, the constraint conditions generated in combination with traffic regulations include but are not limited to prohibiting vehicles in a certain direction from being released for two consecutive red light cycles and requiring a red light in the north-south direction when traveling in the east-west direction.
[0011] Preferably, in step S4, a state space of reinforcement learning is constructed, where Include: Current traffic flow characteristics, namely: , , ; Historical traffic light timing, that is, the green light duration of the last three cycles, that is: ; In the formula, For the The green light duration of each cycle; Weather and time information, including but not limited to: weather type, rainfall, visibility, wind speed, road conditions, namely: ; In the formula, Indicates the weather type. Indicates rainfall, Indicates visibility, Indicates wind speed, Indicates the condition of the road surface; Time information includes, but is not limited to: weekday peak hours and weekend off-peak hours, i.e.: ; In the formula, When it is 1, it indicates a weekday, and when it is 0, it indicates a weekend. When it is 1, it indicates peak hours, and when it is 0, it indicates non-peak hours; The final state vector is expressed as ; Constructing the action space of reinforcement learning, action Including but not limited to: green light duration adjustment and lane priority allocation, among which, Green light duration adjustment: Dynamically adjust the green light duration in each direction, namely: ; In the formula, Indicates Green light duration adjustment value for each direction; Lane priority allocation: Dynamically allocate the priority of each lane according to real-time traffic demand, affecting the weight of green light duration allocation, namely: ; in, Indicates The priority of each lane, , indicating low priority. , indicating medium priority. , indicating high priority; but: Action Space ; Reward function design, namely: Set the minimum average waiting time, ,in, To minimize the average waiting time reward, is the number of vehicles waiting to pass through the intersection, It is the time interval from the vehicle arriving at the intersection to the time it starts to pass through the intersection; To set security penalties: If the action causes the green lights of vehicles in the conflicting directions to turn on at the same time, the penalty ,in, Reward for action conflict; If the queue length exceeds the threshold ,punish , Rewards for queue length; If the average speed is lower than the threshold ,punish ,in is the average speed reward, is the average speed of the current lane; If it is rainy and the green light duration is less than 30 seconds, the penalty ,in For rainy day rewards; If visibility is less than 200 meters and the green light duration is less than 30 seconds, the penalty ,in Reward for low visibility; That is, the comprehensive reward function: ; In the formula, is the weight coefficient, Reward for smoothness; A dual-depth Q network is used to alleviate over-estimation bias, and priority experience replay is introduced to accelerate convergence.
[0012] Preferably, in step S4, after each signal light cycle ends, the current state ,action ,award The experience is stored in the replay pool, batch data is sampled from the pool regularly, and the parameters of the dual-depth Q network are updated through back-propagation.
[0013] The beneficial effects of the present invention are as follows: the present invention explicitly expresses the decision logic through symbolic rules, which makes it easier for traffic managers to understand the basis of the strategy. Online reinforcement learning enables the system to dynamically adjust the strategy to cope with sudden changes in traffic volume, thereby improving real-time adaptability. Through symbolic rule constraints, the generation of strategies that violate traffic regulations is avoided, thereby improving safety assurance. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flow chart of a traffic light control method based on a neural symbolic system and reinforcement learning of the present invention. DETAILED DESCRIPTION
[0015] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. In the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.
[0016] The present invention will be further described below in conjunction with the accompanying drawings.
[0017] like Figure 1 As shown, a traffic light control method based on neural symbolic system and reinforcement learning includes: Step S1, real-time acquisition of vehicle data through cameras, geomagnetic sensors and GPS, weather information through meteorological sensors, weekday, weekend and peak hour identification through servers, and signal light information through signal light controllers.
[0018] In an embodiment, the vehicle data includes, but is not limited to, vehicle volume, vehicle speed, queue length, and weather conditions.
[0019] Step S2, using a convolutional neural network to process the camera video stream, output the real-time vehicle density and speed distribution of each lane, and use the sensor data to construct a traffic flow feature vector.
[0020] In the embodiment, the traffic flow feature vector constructed includes but is not limited to: in, is the set of vehicle density traffic flow feature vectors, For the Vehicle density traffic flow feature vector, is the vehicle speed traffic flow feature vector set, For the Vehicle speed traffic flow feature vector, is the queue length traffic flow feature vector set, For the The queue length traffic flow feature vector.
[0021] Step S3, using a neural symbolic system to map deep learning features into interpretable traffic rules, and generating constraints in combination with traffic regulations.
[0022] In the embodiment, the deep learning features are mapped into interpretable traffic rules. For vehicle density, that is: If the density of the eastbound lane exceeds the threshold, the green light extension is triggered. The rule is formalized as follows: ; In the formula, is the lane density threshold; If the average speed of vehicles in a certain direction is lower than the preset threshold, the green light priority strategy is triggered to improve traffic efficiency; otherwise, the current signal light timing strategy is maintained: ; In the formula, For the The speed threshold for each lane.
[0023] If the queue length of a lane in a certain direction exceeds the preset threshold, the green light extension or priority strategy is triggered to alleviate congestion; otherwise, the current signal light timing strategy is maintained: ; In the formula, For the The queue length threshold for each lane; Therefore, it is necessary to comprehensively consider multiple factors such as vehicle density, speed and queue length. The formal expression of the comprehensive rule is as follows: .
[0024] In the embodiment, the constraint conditions generated in combination with traffic regulations include but are not limited to prohibiting vehicles in a certain direction from being released for two consecutive red light cycles and requiring a red light in the north-south direction when traveling in the east-west direction.
[0025] Step S4, construct the state and action space of reinforcement learning, design the reward function based on traffic rules and constraints, and use the dual deep Q network to optimize the traffic light timing strategy.
[0026] In the embodiment, the state space of reinforcement learning is constructed, and the state Include: Current traffic flow characteristics, namely: , , ; Historical traffic light timing, such as the green light duration in the last three cycles, that is: ; In the formula, For the The green light duration of each cycle; Weather and time information, including but not limited to: weather type, rainfall, visibility, wind speed, road conditions, namely: ; In the formula, Indicates the weather type. Indicates rainfall, Indicates visibility, Indicates wind speed, Indicates the condition of the road surface; Time information includes, but is not limited to: weekday peak hours and weekend off-peak hours, i.e.: ; In the formula, When it is 1, it indicates a weekday, and when it is 0, it indicates a weekend. When it is 1, it indicates peak hours, and when it is 0, it indicates non-peak hours; The final state vector is expressed as ; Constructing the action space of reinforcement learning, action Including but not limited to: green light duration adjustment and lane priority allocation, among which, Green light duration adjustment: Dynamically adjust the green light duration in each direction, namely: ; In the formula, Indicates Green light duration adjustment value for each direction; Lane priority allocation: Dynamically allocate the priority of each lane according to real-time traffic demand, affecting the weight of green light duration allocation, namely: ; in, Indicates The priority of each lane, , indicating low priority. , indicating medium priority. , indicating high priority; but: Action Space ; Reward function design, namely: Set the minimum average waiting time, ,in, To minimize the average waiting time reward, is the number of vehicles waiting to pass through the intersection, The time interval from when a vehicle arrives at an intersection to when it begins to pass through the intersection; To set security penalties: If the action causes the green lights of vehicles in the conflicting directions to turn on at the same time, the penalty ,in, Reward for action conflict; If the queue length exceeds the threshold ,punish , Rewards for queue length; If the average speed is lower than the threshold ,punish ,in is the average speed reward, is the average speed of the current lane; If it is rainy and the green light duration is less than 30 seconds, the penalty ,in For rainy day rewards; If visibility is less than 200 meters and the green light duration is less than 30 seconds, the penalty ,in Reward for low visibility; That is, the comprehensive reward function: ; In the formula, is the weight coefficient, Reward for smoothness (avoid frequent switching of lights); A dual-depth Q network is used to alleviate over-estimation bias, and priority experience replay is introduced to accelerate convergence.
[0027] In the embodiment, after each signal light cycle ends, the current state ,action ,award The experience is stored in the replay pool, batch data is sampled from the pool regularly, and the parameters of the dual-depth Q network are updated through back-propagation.
[0028] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A traffic light control method based on neural symbolic system and reinforcement learning, characterized in that: include: Step S1, real-time acquisition of vehicle data through a camera, geomagnetic sensor and GPS, weather information through a meteorological sensor, identification of working days, weekends and peak hours through a server, and signal light information through a signal light controller; Step S2, using a convolutional neural network to process the camera video stream, output the real-time vehicle density and speed distribution of each lane, and use the sensor data to construct a traffic flow feature vector; Step S3, using a neural symbolic system to map deep learning features into interpretable traffic rules, and generating constraints in combination with traffic regulations; Step S4, construct the state and action space of reinforcement learning, design the reward function based on traffic rules and constraints, and use the dual deep Q network to optimize the traffic light timing strategy.
2. A traffic light control method based on neural symbolic system and reinforcement learning according to claim 1, characterized in that: In step S1, the vehicle data includes but is not limited to vehicle flow, vehicle speed, queue length and weather conditions.
3. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 1, characterized in that: In step S2, the traffic flow feature vector constructed includes but is not limited to: in, is the set of vehicle density traffic flow feature vectors, For the Vehicle density traffic flow feature vector, is the vehicle speed traffic flow feature vector set, For the Vehicle speed traffic flow feature vector, is the queue length traffic flow feature vector set, For the The queue length traffic flow feature vector.
4. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 3 is characterized in that: In step S3, the deep learning features are mapped into interpretable traffic rules. For vehicle density, that is: If the density of the eastbound lane exceeds the threshold, the green light extension is triggered. The rule is formalized as follows: ; In the formula, is the lane density threshold; If the average speed of vehicles in a certain direction lane is lower than the preset threshold, the green light priority strategy is triggered to improve traffic efficiency; Otherwise, maintain the current signal light timing strategy: ; In the formula, For the Speed threshold for each lane; If the queue length in a certain direction exceeds the preset threshold, a green light extension or priority strategy is triggered to ease congestion; Otherwise, maintain the current signal light timing strategy: ; In the formula, For the The queue length threshold for each lane; Therefore, it is necessary to comprehensively consider multiple factors such as vehicle density, speed and queue length. The formal expression of the comprehensive rule is as follows: 。 5. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 1, characterized in that: In step S3, the constraint conditions generated in combination with traffic regulations include but are not limited to prohibiting vehicles in a certain direction from being released for two consecutive red light cycles and requiring a red light in the north-south direction when traveling in the east-west direction.
6. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 3 is characterized in that: In step S4, the state space of reinforcement learning is constructed. Include: Current traffic flow characteristics, namely: , , ; Historical traffic light timing, that is, the green light duration of the last three cycles, that is: ; In the formula, For the The green light duration of each cycle; Weather and time information, including but not limited to: weather type, rainfall, visibility, wind speed, road conditions, namely: ; In the formula, Indicates the weather type. Indicates rainfall, Indicates visibility, Indicates wind speed, Indicates the condition of the road surface; Time information includes, but is not limited to: weekday peak hours and weekend off-peak hours, i.e.: ; In the formula, When it is 1, it indicates a weekday, and when it is 0, it indicates a weekend. When it is 1, it indicates peak hours, and when it is 0, it indicates non-peak hours; The final state vector is expressed as ; Constructing the action space of reinforcement learning, action Including but not limited to: green light duration adjustment and lane priority allocation, among which, Green light duration adjustment: Dynamically adjust the green light duration in each direction, namely: ; In the formula, Indicates Green light duration adjustment value for each direction; Lane priority allocation: Dynamically allocate the priority of each lane according to real-time traffic demand, affecting the weight of green light duration allocation, namely: ; in, Indicates The priority of each lane, , indicating low priority. , indicating medium priority. , indicating high priority; but: Action Space ; Reward function design, namely: Set the minimum average waiting time, ,in, To minimize the average waiting time reward, is the number of vehicles waiting to pass through the intersection, The time interval from when a vehicle arrives at an intersection to when it begins to pass through the intersection; To set security penalties: If the action causes the green lights of vehicles in the conflicting directions to turn on at the same time, the penalty ,in, Reward for action conflict; If the queue length exceeds the threshold ,punish , Rewards for queue length; If the average speed is lower than the threshold ,punish ,in is the average speed reward, is the average speed of the current lane; If it is rainy and the green light duration is less than 30 seconds, the penalty ,in For rainy day rewards; If visibility is less than 200 meters and the green light duration is less than 30 seconds, the penalty ,in Reward for low visibility; That is, the comprehensive reward function: ; In the formula, is the weight coefficient, Reward for smoothness; A dual-depth Q network is used to alleviate over-estimation bias, and priority experience replay is introduced to accelerate convergence.
7. The traffic light control method based on neural symbolic system and reinforcement learning according to claim 6, characterized in that: In step S4, after each signal light cycle ends, the current state ,action ,award The experience is stored in the replay pool, batch data is sampled from the pool regularly, and the parameters of the dual-depth Q network are updated through back-propagation.
Citation Information
Patent Citations
Real-time recommendation method for traffic signal control scheme based on deep learning
CN110491146A
Traffic signal timing optimization method based on deep reinforcement learning
CN112700664A
Traffic signal control method and device, electronic equipment and storage medium
CN116189454A
Multi-intersection traffic signal control method based on deep reinforcement learning
CN117409593A
Traffic signal lamp control method suitable for real-time traffic condition
CN117746652A
Cited By
Traffic signal optimization method based on multi-agent deep reinforcement learning
CN120599840A
Intelligent decision-making system construction method for traffic signal control
CN121564997A