An Edge Computing-Based Reinforcement Learning Control Method and System
Through edge computing combined with graph convolution feature coding and reinforcement learning algorithm, the traffic signal phase control is optimized in real time, and the traffic efficiency and pedestrian waiting anxiety problems of traditional methods in complex traffic scenarios are solved, achieving efficient traffic management.
Patent Information
- Application Number
- CN202510559499.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-30
AI Technical Summary
When faced with complex traffic scenarios, existing traffic signal control technology is difficult to respond to changes in traffic flow in real time, resulting in low traffic efficiency and high pedestrian waiting anxiety index. Traditional PID control algorithms and centralized AI control models have limitations in practical applications.
The reinforcement learning control method based on edge computing is adopted, and the traffic information and pedestrian resident hot zone map are collected in real time through edge fusion perception equipment, and the signal phase scheme is determined using graph convolution feature coding and reinforcement learning algorithm, and the phase control of traffic lights is finally realized through digital twin model optimization.
Real-time response to traffic flow changes is achieved, phase control is optimized, traffic efficiency is improved, and pedestrian waiting anxiety is reduced, providing efficient and reliable technical support for smart city traffic governance.
Smart Images

Figure CN120088988B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent traffic control, and particularly to a reinforcement learning control method and system based on edge computing. Background Art
[0002] In modern urban traffic management systems, signal control in complex traffic scenarios such as intersections and highway ramps is a key link to ensure smooth traffic, improve road capacity, and reduce traffic accidents. However, existing signal control technologies face many challenges in practical applications, especially in dealing with sudden changes in short-term traffic flow, the dynamic game between pedestrian crossing demands and motor vehicle flows, and the requirement for real-time response. The limitations of existing technologies are particularly prominent.
[0003] The traditional Proportional-Integral-Differential (PID) control algorithm, as a classic traffic signal control strategy, has been widely used in the field of traffic signal control. However, its linear adjustment mechanism based on fixed parameters is inadequate in the face of complex and variable traffic flows. Especially during the morning and evening rush hours, the traffic flow fluctuates greatly, which makes it difficult for the traditional PID control algorithm to accurately predict and adapt to traffic flow changes. As a result, the signal cycle switching frequency increases significantly, the phase conflict rate rises, and the traffic efficiency decreases, seriously affecting the overall operation efficiency of urban traffic.
[0004] To overcome the deficiencies of the traditional PID control algorithm, rule-based optimization methods such as MAXBAND have been proposed in the industry. These methods attempt to achieve optimized control of traffic signals to a certain extent by presetting phase priorities based on artificial experience. However, they also have significant defects in practical applications: due to their inability to respond in real time to the dynamic game between pedestrian crossing demands and motor vehicle flows, the average waiting anxiety index of pedestrians is relatively high. This not only affects the travel experience of pedestrians but also may exacerbate traffic conflicts and reduce the overall safety of the traffic system.
[0005] With the rapid development of artificial intelligence technology, centralized AI control models have also begun to emerge in the field of traffic signal control. These models achieve data analysis and decision-making through cloud computing and theoretically can more accurately predict and adapt to traffic flow changes. However, in practical applications, centralized AI control models also face challenges. Due to their dependence on cloud computing, there is a certain delay in making decisions by the models, which is difficult to meet the high requirements for real-time response in traffic signal control. In addition, the generalization ability of centralized AI control models is relatively weak. When applied to new traffic intersections, the models need to be retrained to adapt to the new traffic environment. This not only increases the deployment cost but also limits the wide application of the models in complex and variable traffic scenarios.
[0006] In summary, there are many deficiencies in the existing technologies in the field of real-time signal control in complex traffic scenarios such as urban intersections and highway ramps. There is an urgent need for a new signal control technology that can respond to traffic flow changes in real time, optimize phase control, improve traffic efficiency, and reduce pedestrians' waiting anxiety. Summary of the Invention
[0007] To solve the above problems, embodiments of the present invention provide a reinforcement learning control method and system based on edge computing. The embodiments of the present invention adopt the following technical solutions:
[0008] On the one hand, embodiments of the present invention provide a reinforcement learning control method based on edge computing. The method includes: real-time collecting traffic information of the current intersection and a pedestrian residence heat map through an edge fusion perception device; wherein, the traffic information at least includes the average vehicle moving speed and the vehicle queue length;
[0009] Based on a graph convolutional feature encoding model installed in the edge fusion perception device, performing graph convolutional feature encoding on the traffic information and the pedestrian residence heat map to obtain a traffic situation feature vector of the current intersection;
[0010] Based on the traffic situation feature vector and a reinforcement learning algorithm, determining an initial signal phase plan for the traffic signal lights at the current intersection;
[0011] Constructing a digital twin model of the traffic intersection, and optimizing the initial signal phase plan through the digital twin model to obtain the final signal phase plan for each intersection;
[0012] Based on the signal phase plans of each intersection, performing phase control on the corresponding traffic signal lights.
[0013] In a feasible implementation manner, before real-time collecting traffic information of the current intersection and a pedestrian residence heat map through an edge fusion perception device, the method further includes:
[0014] Fusing a millimeter-wave radar, a thermal imaging camera, and a microprocessor to construct a new type of edge fusion perception device;
[0015] Deploying the edge fusion perception device at each intersection in the area to be controlled, and installing a preset graph convolutional feature encoding model and a preset signal phase plan generation algorithm in the edge fusion perception device for edge computing;
[0016] Connecting the edge fusion perception devices deployed at each intersection into an edge fusion perception network for data communication.
[0017] In a feasible implementation, traffic information of the current intersection and a pedestrian staying heat zone map are collected in real time through an edge fusion perception device, specifically including:
[0018] Different target objects at the current intersection are identified through the millimeter-wave radar integrated in the edge fusion perception device; among them, the target objects include static objects and dynamic objects; the static objects at least include lanes and crosswalks; the dynamic objects at least include vehicles and pedestrians;
[0019] Through the millimeter-wave radar, the position information and distance information of the static objects are obtained; and, the average moving speed and queue length of the dynamic objects within a preset time period are obtained; among them, the average moving speed and queue length of the vehicles constitute the traffic information of the current intersection, and the average moving speed and queue length of the pedestrians constitute the pedestrian information of the current intersection;
[0020] A sequence of thermal imaging maps within a preset time period at the current intersection is collected through the thermal imaging camera integrated in the edge fusion perception device;
[0021] Based on the thermal imaging map sequence and the pedestrian information, the pedestrian staying heat zone map is obtained.
[0022] In a feasible implementation, based on the thermal imaging map sequence and the pedestrian information, obtaining the pedestrian staying heat zone map specifically includes:
[0023] In the thermal imaging map sequence, the human target positions in each thermal imaging map are identified, and clustering analysis is performed based on the human target positions to determine the clustering regions where the pedestrian density is higher than the first preset threshold;
[0024] Each thermal imaging map in the thermal imaging map sequence is traversed, and through superposition and comparison, the target clustering regions with an area overlap degree higher than the second preset threshold between the first thermal imaging map and the second thermal imaging map are sequentially determined; then the target clustering regions are compared with the third thermal imaging map, and so on, until after the comparison of the last thermal imaging map is completed, the final target clustering regions are obtained;
[0025] Based on the signal light duration and the crosswalk length at the current intersection, the average moving speed reference value of the pedestrians within the preset time period is calculated;
[0026] In the final clustering regions, the regions where the difference between the average moving speed of the pedestrians and the average moving speed reference value is less than the third preset threshold are screened out and determined as the pedestrian staying heat zones;
[0027] The image at the pedestrian staying heat zone is obtained through the camera installed at the current intersection, and the pedestrian staying heat zone map is obtained.
[0028] In a feasible implementation, before performing graph convolutional feature encoding on the traffic information and the pedestrian residence heat map based on the graph convolutional feature encoding model installed in the edge fusion perception device, the method further includes:
[0029] Abstract the static objects in the current intersection as nodes, and abstract the traffic flow turning relationships between each pair of static objects as edges, obtaining a directed graph G=(V, E) of the current intersection;
[0030] wherein, V={v1,v2,……,v n} represents the nodes, and n is the total number of nodes in the current intersection; the nodes in the intersection at least include each motor vehicle lane and each crosswalk; E is the edge of the directed graph, and the traffic flow turning relationships at least include a straight-ahead relationship, a left-turn relationship, and a right-turn relationship.
[0031] In a feasible implementation, based on the graph convolutional feature encoding model installed in the edge fusion perception device, perform graph convolutional feature encoding on the traffic information and the pedestrian residence heat map to obtain a traffic situation feature vector of the current intersection, specifically including:
[0032] Input the traffic information and the pedestrian residence heat map into the graph convolutional feature encoding model, and output the node feature vectors of each node in the directed graph; wherein, the node feature vectors at least include the vehicle queue length, the traffic flow change rate, the pedestrian density, and the remaining time of the current phase of the traffic signal corresponding to the current node;
[0033] Combine the node feature vectors of each node into a multi-dimensional matrix to obtain a traffic situation feature vector of the current intersection.
[0034] In a feasible implementation, based on the traffic situation feature vector and the reinforcement learning algorithm, determine an initial signal phase plan for the traffic signal at the current intersection, specifically including:
[0035] Based on the reinforcement learning algorithm, construct and train a PPO policy network; wherein, the signal control instruction set of the PPO policy network is: a t ∈{extend the current phase, switch to the next phase, activate the pedestrian priority phase};
[0036] Input the traffic situation feature vector obtained at the current moment into the PPO policy network, and obtain the probability distribution of each control instruction in the signal control instruction set;
[0037] Determine the signal control instruction with the highest probability as the initial signal phase plan.
[0038] In a feasible implementation manner, the digital twin model is used to optimize the initial signal phase plan to obtain the final signal phase plan for each intersection, specifically including:
[0039] Input the initial signal phase plan into the digital twin model to perform simulation control of traffic lights;
[0040] In the digital twin model, obtain the traffic situation feature vector for the next preset duration, and calculate the plan evaluation value corresponding to the initial signal phase plan based on the dual-objective reward function;
[0041] If the plan evaluation value is not lower than the preset standard threshold, determine the initial signal phase plan as the signal phase plan for the current intersection;
[0042] If the plan evaluation value is lower than the preset standard threshold, optimize the policy network parameters and hyperparameters in the PPO policy network, regenerate the initial signal phase plan until the plan evaluation value of the obtained initial signal phase plan is not lower than the preset standard threshold.
[0043] In a feasible implementation manner, based on the dual-objective reward function, calculate the plan evaluation value corresponding to the initial signal phase plan, specifically including:
[0044] According to , calculate the plan evaluation value of the initial signal phase plan;
[0045] Among them, is the average vehicle delay time at the current intersection, is the pedestrian anxiety index at the current intersection, obtained according to the number of people staying and waiting time in each pedestrian stay hot zone; and are weight coefficients, dynamically adjusted through the Pareto front.
[0046] On the other hand, the embodiment of the present invention also provides a reinforcement learning control system based on edge computing, characterized in that the system includes:
[0047] A data acquisition module for real-time collecting traffic information of the current intersection and a pedestrian stay hot zone map through edge fusion perception devices; wherein, the traffic information at least includes the average vehicle moving speed and the vehicle queue length;
[0048] The signal phase plan generation module is used to perform graph convolutional feature encoding on the traffic information and the pedestrian residence heat zone map based on the graph convolutional feature encoding model installed in the edge fusion perception device, so as to obtain the traffic situation feature vector of the current intersection; and determine the initial signal phase plan of the traffic signal lights at the current intersection based on the traffic situation feature vector and the reinforcement learning algorithm.
[0049] The signal phase plan optimization module is used to construct a digital twin model of the traffic intersection, and optimize the initial signal phase plan through the digital twin model to obtain the final signal phase plan for each intersection; and perform phase control on the corresponding traffic signal lights based on the signal phase plans of each intersection.
[0050] Compared with the prior art, the reinforcement learning control method and system based on edge computing provided by the embodiments of the present invention have the following beneficial effects:
[0051] In the present invention, by deploying edge computing devices at traffic intersections, the collected traffic information and pedestrian information are directly processed in the edge devices without being transmitted to the cloud, with a fast response speed, meeting the requirements of real-time response for traffic signal control. And the present invention uses millimeter-wave radars to collect traffic information. Compared with optical seekers such as infrared, laser, and television, the millimeter-wave seeker has strong ability to penetrate fog, smoke, and dust, has the characteristics of obtaining data efficiently all-weather and all-day, and the anti-interference and anti-stealth capabilities of the millimeter-wave seeker are also superior to other microwave seekers, and it is more suitable for collecting traffic road information in the open air.
[0052] Furthermore, the present invention proposes a phase decision method combining a graph convolutional feature encoding model and a reinforcement learning algorithm, taking into account the dynamic game between the pedestrian crossing demand and the motor vehicle flow. And through the digital twin model, the generated phase adjustment plan is simulated, and the generated phase adjustment plan is evaluated and optimized according to the average vehicle delay value and the average pedestrian waiting anxiety index, so as to generate a phase adjustment plan more suitable for the actual scenario. Through the closed-loop design of "perception - decision - optimization", the present invention provides a new signal control technology that can respond to traffic flow changes in real time, optimize phase control, improve traffic efficiency, and reduce pedestrian waiting anxiety, overcoming the problem of control instability of traditional methods in dynamic traffic scenarios, and providing an efficient and reliable technical support for the traffic governance of smart cities. Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0054] Figure 1 Flowchart of a reinforcement learning control method based on edge computing provided by an embodiment of the present invention;
[0055] Figure 2 Structural schematic diagram of a reinforcement learning control system based on edge computing provided by an embodiment of the present invention. Detailed implementation manners
[0056] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] An embodiment of the present invention provides a reinforcement learning control method based on edge computing. As Figure 1 shown, the reinforcement learning control method based on edge computing specifically includes steps S101 - S104:
[0058] S101. Through an edge fusion perception device, the traffic information of the current intersection and the pedestrian stay heat zone map are collected in real time; wherein, the traffic information at least includes the average vehicle moving speed and the vehicle queue length.
[0059] Specifically, a millimeter - wave radar, a thermal imaging camera and a micro - processor are fused to construct a new type of edge fusion perception device.
[0060] As a feasible implementation manner, the millimeter - wave radar, the thermal imaging camera, the micro - processor and some necessary auxiliary devices are encapsulated in a housing and fused to construct a new type of edge fusion perception device, and then the edge fusion perception device is deployed at each intersection in the area to be controlled. The millimeter - wave radar and the thermal imaging camera are used to collect the traffic data at the intersection. A preset graph convolutional feature encoding model and a preset signal phase scheme generation algorithm are installed in the micro - processor of the edge fusion perception device for performing edge computing on the collected traffic data.
[0061] Further, the edge fusion perception devices deployed at each intersection are connected into an edge fusion perception network to conduct data communication between intersections and cooperate in the phase control of traffic signals.
[0062] Further, after the edge fusion perception network is deployed, different target objects at the current intersection are identified by millimeter-wave radar. Among them, the target objects include static objects and dynamic objects; the static objects at least include lanes and crosswalks; the dynamic objects at least include vehicles and pedestrians.
[0063] Further, the position information and distance information of static objects are obtained by millimeter-wave radar; and, the average moving speed and queue length of dynamic objects within a preset time period are obtained. Among them, the average moving speed and queue length of vehicles constitute the traffic information of the current intersection, and the average moving speed and queue length of pedestrians constitute the pedestrian information of the current intersection.
[0064] Further, a sequence of thermal imaging maps within a preset time period at the current intersection is collected by the thermal imaging camera integrated in the edge fusion perception device. Then, based on the sequence of thermal imaging maps and pedestrian information, a pedestrian residence hot zone map is obtained.
[0065] As a feasible implementation manner, based on the sequence of thermal imaging maps and pedestrian information, obtaining a pedestrian residence hot zone map, the specific implementation process is as follows:
[0066] In the sequence of thermal imaging maps, the human target positions in each thermal imaging map are identified, and clustering analysis is performed based on the human target positions to determine the clustering regions where the pedestrian density is higher than the first preset threshold. Each thermal imaging map in the sequence of thermal imaging maps is traversed, and through superposition and comparison, the target clustering regions with an area overlap degree higher than the second preset threshold between the first thermal imaging map and the second thermal imaging map are sequentially determined; then the target clustering regions are compared with the third thermal imaging map, and so on, until after the comparison of the last thermal imaging map is completed, the final target clustering regions are obtained. Based on the signal light duration and crosswalk length of the current intersection, the average moving speed reference value of pedestrians within a preset time period is calculated. In the final clustering regions, the regions where the difference between the average moving speed of pedestrians and the average moving speed reference value is less than the third preset threshold are screened out and determined as pedestrian residence hot zones. The images at the pedestrian residence hot zones are obtained through the cameras installed at the current intersection to obtain the pedestrian residence hot zone map.
[0067] The present invention performs target recognition on the thermal imaging map at the intersection, determines multiple pedestrian gathering areas through clustering analysis technology, and then preliminarily determines the pedestrian staying hot areas according to the coincidence degree of the pedestrian gathering areas in the image sequence, that is, the areas where pedestrians stay for a long time and have a high density. Then, the traffic signal lights are considered in the pedestrian movement speed, and the normal average movement speed that normal pedestrians should have when crossing the road according to the traffic lights is calculated. And as the average movement speed reference value, in the preliminarily determined pedestrian staying hot areas, the situations where crowd gathering may be caused by scenes such as bus stops and store entrances are screened out, so as to determine a more accurate signal light waiting area and provide a more accurate data basis for the subsequent calculation of the pedestrian anxiety index.
[0068] S102. Based on the graph convolutional feature encoding model installed in the edge fusion perception device, perform graph convolutional feature encoding on the traffic information and the pedestrian staying hot area map to obtain the traffic situation feature vector of the current intersection.
[0069] Specifically, the static objects in the current intersection are abstracted as nodes, and the traffic flow turning relationship between each static object is abstracted as an edge to obtain the directed graph G=(V, E) of the current intersection.
[0070] Among them, V={v1, v2, ……, v n} represents nodes, and n is the total number of nodes in the current intersection; the nodes in the intersection at least include each motor vehicle lane and each crosswalk; E is the edge of the directed graph, and the traffic flow turning relationship at least includes a straight relationship, a left-turn relationship, and a right-turn relationship.
[0071] In one embodiment, each different lane at the intersection is a node, and each crosswalk is also a node. Generally, a standard crossroads has 8 motor vehicle lanes in different directions and 4 crosswalks, which can be abstracted as 12 nodes. There are straight, left-turn, and right-turn relationships between each different motor vehicle lane, and left-turn and right-turn relationships between each adjacent crosswalk.
[0072] Further, input the traffic information and the pedestrian staying hot area map into the graph convolutional feature encoding model pre-installed in the edge fusion perception device, and output the node feature vector of each node in the directed graph; among them, the node feature vector at least includes the vehicle queue length, flow rate change rate, pedestrian density, and the remaining time of the current phase of the traffic signal lights corresponding to the current node.
[0073] Finally, combine the node feature vectors of each node into a multi-dimensional matrix to obtain the traffic situation feature vector of the current intersection.
[0074] As a feasible implementation manner, the node feature vector of a single node is: . Among them, q is the vehicle queue length, is the traffic flow change rate within a preset duration, is the pedestrian density, and is the remaining time of the current phase of the traffic signal at the current intersection.
[0075] S103. Based on the traffic situation feature vector and the reinforcement learning algorithm, determine the initial signal phase plan of the traffic signal at the current intersection.
[0076] Specifically, based on the reinforcement learning algorithm, construct and train a PPO policy network, and install it in the edge fusion perception device. Among them, the signal control instruction set of the PPO policy network is: a t ∈ {extend the current phase, switch to the next phase, activate the pedestrian priority phase}.
[0077] Furthermore, input the traffic situation feature vector obtained at the current moment into the PPO policy network to obtain the probability distribution of each control instruction in the signal control instruction set. Determine the signal control instruction with the highest probability as the initial signal phase plan. Then extract the corresponding detailed signal phase control plan from the database.
[0078] In one embodiment, if the probability distribution output by the PPO policy network is P = [p1, p2, p3], and p2 is the largest, then the initial signal phase plan is to switch to the next phase. If p1 is the largest, then the initial signal phase plan is to extend the current phase. Similarly, if p3 is the largest, then the initial signal phase plan is to activate the pedestrian priority phase.
[0079] S104. Construct a digital twin model of the traffic intersection, and optimize the initial signal phase plan through the digital twin model to obtain the final signal phase plan for each intersection; based on the signal phase plans of each intersection, perform phase control on the corresponding traffic signals.
[0080] Specifically, construct an overall digital twin model of the area to be controlled based on digital twin technology. Then input the initial signal phase plans of each intersection into the digital twin model to perform simulation control of the traffic signals.
[0081] In the digital twin model, obtain the traffic situation feature vector of each intersection within the next preset duration, and calculate the evaluation value of the plan corresponding to the initial signal phase plan based on the double-objective reward function.
[0082] If the evaluation value of the plan is not lower than the preset standard threshold, determine the initial signal phase plan as the signal phase plan of the current intersection.
[0083] Further, if the scheme evaluation value is lower than the preset standard threshold, optimize the policy network parameters and hyperparameters in the PPO policy network, and regenerate the initial signal phase scheme until the scheme evaluation value of the obtained initial signal phase scheme is not lower than the preset standard threshold.
[0084] As a feasible implementation manner, based on the dual-objective reward function, calculate the scheme evaluation value corresponding to the initial signal phase scheme, specifically including:
[0085] According to , calculate the scheme evaluation value of the initial signal phase scheme;
[0086] wherein, is the average vehicle delay time at the current intersection, is the pedestrian anxiety index at the current intersection, obtained according to the number of pedestrians staying in the pedestrian stay hot zone and the waiting time; and are weight coefficients, dynamically adjusted through the Pareto front.
[0087] Finally, based on the signal phase schemes of each intersection, perform phase control on the corresponding traffic lights, so as to implement a reinforcement learning control method based on edge computing provided by the present invention.
[0088] In addition, an embodiment of the present invention further provides a reinforcement learning control system based on edge computing, as Figure 2 shown, the reinforcement learning control system 200 based on edge computing specifically includes:
[0089] A data acquisition module 210, configured to collect traffic information of the current intersection and a pedestrian stay hot zone map in real time through an edge fusion perception device; wherein, the traffic information at least includes the average vehicle moving speed and the vehicle queue length;
[0090] A signal phase scheme generation module 220, configured to perform graph convolutional feature encoding on the traffic information and the pedestrian stay hot zone map based on the graph convolutional feature encoding model installed in the edge fusion perception device to obtain a traffic situation feature vector of the current intersection; determine an initial signal phase scheme of the traffic lights at the current intersection based on the traffic situation feature vector and a reinforcement learning algorithm;
[0091] A signal phase scheme optimization module 230, configured to construct a digital twin model of the traffic intersection, optimize the initial signal phase scheme through the digital twin model to obtain the final signal phase schemes of each intersection; perform phase control on the corresponding traffic lights based on the signal phase schemes of each intersection.
[0092] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
[0093] The above describes specific embodiments of the present invention. Additionally, the processes depicted in the drawings do not necessarily require the specific order or consecutive order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0094] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and changes can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the present invention.
Claims
1. A reinforcement learning control method based on edge computing, characterized in that The method includes: Collecting the traffic information of the current intersection and the pedestrian stay heat zone map in real time through an edge fusion perception device; wherein, the traffic information at least includes the average vehicle moving speed and the vehicle queue length; specifically including: Identifying different target objects at the current intersection through the millimeter-wave radar integrated in the edge fusion perception device; wherein, the target objects include static objects and dynamic objects; the static objects at least include lanes and crosswalks; the dynamic objects at least include vehicles and pedestrians; Obtaining the position information and distance information of the static objects through the millimeter-wave radar; and obtaining the average moving speed and queue length of the dynamic objects within a preset time period; wherein, the average moving speed and queue length of the vehicles constitute the traffic information of the current intersection, and the average moving speed and queue length of the pedestrians constitute the pedestrian information of the current intersection; Collecting a sequence of thermal imaging maps of the current intersection within a preset time period through the thermal imaging camera integrated in the edge fusion perception device; based on the sequence of thermal imaging maps and the pedestrian information, obtaining the pedestrian stay heat zone map, specifically including: Identifying the human target positions in each thermal imaging map in the sequence of thermal imaging maps, and performing clustering analysis based on the human target positions to determine the clustering regions where the pedestrian density is higher than the first preset threshold; Traversing each thermal imaging map in the sequence of thermal imaging maps, and successively determining the target clustering regions where the regional coincidence degree between the first thermal imaging map and the second thermal imaging map is higher than the second preset threshold through superposition and comparison; then comparing the target clustering regions with the third thermal imaging map, and so on, until after the comparison of the last thermal imaging map is completed, obtaining the final target clustering regions; Calculating the average moving speed reference value of the pedestrians within the preset time period based on the signal light duration and the crosswalk length at the current intersection; in the final clustering regions, screening out the regions where the difference between the average moving speed of the pedestrians and the average moving speed reference value is less than the third preset threshold, and determining them as the pedestrian stay heat zones; Obtaining the image at the pedestrian stay heat zone through the camera installed at the current intersection, and obtaining the pedestrian stay heat zone map; Performing graph convolution feature encoding on the traffic information and the pedestrian stay heat zone map based on the graph convolution feature encoding model installed in the edge fusion perception device to obtain the traffic situation feature vector of the current intersection; Determining the initial signal phase scheme of the traffic signal at the current intersection based on the traffic situation feature vector and the reinforcement learning algorithm, specifically including: Based on the reinforcement learning algorithm, a PPO policy network is constructed and trained; wherein, the signal control instruction set of the PPO policy network is: a t ∈ {extend the current phase, switch to the next phase, activate the pedestrian priority phase}; input the traffic situation feature vector obtained at the current moment into the PPO policy network to obtain the probability distribution of each control instruction in the signal control instruction set; determine the signal control instruction with the highest probability as the initial signal phase plan; Constructing a digital twin model of the traffic intersection, and optimizing the initial signal phase scheme through the digital twin model to obtain the final signal phase scheme for each intersection, specifically including: Inputting the initial signal phase scheme into the digital twin model to perform simulation control of the traffic signal; Obtaining the traffic situation feature vector of the next preset time period in the digital twin model, and calculating the scheme evaluation value corresponding to the initial signal phase scheme based on the dual-objective reward function; If the evaluation value of the scheme is not lower than the preset standard threshold, determine the initial signal phase scheme as the signal phase scheme of the current intersection; If the evaluation value of the scheme is lower than the preset standard threshold, optimize the policy network parameters and hyperparameters in the PPO policy network, and regenerate the initial signal phase scheme until the evaluation value of the obtained initial signal phase scheme is not lower than the preset standard threshold; Based on the signal phase schemes of each intersection, perform phase control on the corresponding traffic lights.
2. The enhanced learning control method based on edge computing according to claim 1, characterized in that, Before collecting the traffic information of the current intersection and the pedestrian stay heat map in real time through the edge fusion perception device, the method further includes: Fuse the millimeter-wave radar, thermal imaging camera and microprocessor to construct a new type of edge fusion perception device; Deploy the edge fusion perception device at each intersection in the area to be controlled, and install a preset graph convolutional feature encoding model and a preset signal phase scheme generation algorithm in the edge fusion perception device for edge computing; Connect the edge fusion perception devices deployed at each intersection into an edge fusion perception network for data communication.
3. The enhanced learning control method based on edge computing according to claim 1, characterized in that Before performing graph convolutional feature encoding on the traffic information and the pedestrian stay heat map based on the graph convolutional feature encoding model installed in the edge fusion perception device, the method further includes: Abstract the static objects in the current intersection as nodes, and abstract the traffic flow turning relationship between each static object as an edge to obtain a directed graph G=(V, E) of the current intersection; Among them, V = {v1, v2, ……, v n} represents the nodes, where n is the total number of nodes in the current intersection; the nodes in the intersection include at least each motor vehicle lane and each crosswalk; E is the edge of the directed graph, and the traffic flow turning relationship includes at least a straight-ahead relationship, a left-turn relationship, and a right-turn relationship.
4. The reinforcement learning control method based on edge computing according to claim 3, wherein Based on the graph convolutional feature encoding model installed in the edge fusion perception device, perform graph convolutional feature encoding on the traffic information and the pedestrian stay heat map to obtain the traffic situation feature vector of the current intersection, specifically including: Input the traffic information and the pedestrian stay heat map into the graph convolutional feature encoding model, and output the node feature vector of each node in the directed graph; wherein, the node feature vector at least includes the vehicle queue length, flow rate change rate, pedestrian density corresponding to the current node, and the remaining time of the current phase of the traffic lights at the current intersection; Combine the node feature vectors of each node into a multi-dimensional matrix to obtain the traffic situation feature vector of the current intersection.
5. A reinforcement learning control method based on edge computing according to claim 1, characterized in that, Based on the dual-objective reward function, calculate the evaluation value corresponding to the initial signal phase scheme, specifically including: According to , calculate the scheme evaluation value of the initial signal phase scheme; Among them, is the average vehicle delay time at the current intersection, is the pedestrian anxiety index at the current intersection, which is obtained based on the number of pedestrians staying in the heat zone and their waiting time; and are weight coefficients, which are dynamically adjusted through the Pareto front.
6. A reinforcement learning control system based on edge computing, which applies a reinforcement learning control method based on edge computing as described in any one of claims 1-5, characterized in that, The system includes: A data acquisition module for collecting the traffic information of the current intersection and the pedestrian stay heat map in real time through the edge fusion perception device; wherein, the traffic information at least includes the average vehicle moving speed and the vehicle queue length; A signal phase scheme generation module for performing graph convolutional feature encoding on the traffic information and the pedestrian stay heat map based on the graph convolutional feature encoding model installed in the edge fusion perception device to obtain the traffic situation feature vector of the current intersection; and determining the initial signal phase scheme of the traffic lights at the current intersection based on the traffic situation feature vector and the reinforcement learning algorithm; The signal phase scheme optimization module is used to construct a digital twin model of a traffic intersection, and optimize the initial signal phase scheme through the digital twin model to obtain the final signal phase scheme for each intersection; based on the signal phase schemes of each intersection, perform phase control on the corresponding traffic lights.
Citation Information
Patent Citations
Road traffic signal control system and method
CN119649621A