A parameter self-learning signal control method based on cloud edge collaboration
By collecting traffic holographic trajectory data to construct a parameter self-learning model, and combining it with a cloud-edge collaborative control architecture, the dependence of traffic signal control methods on traffic parameter configuration and training samples is solved, and real-time and all-weather, all-scenario adaptability of traffic signal control is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG SUPCON INFORMATION TECH CO LTD
- Filing Date
- 2023-06-21
- Publication Date
- 2026-05-05
AI Technical Summary
Existing traffic signal control methods rely heavily on traffic parameter configuration and training samples, making it difficult to achieve real-time, all-weather, and all-scenario applications. Furthermore, existing technologies are ill-suited to various traffic flow conditions and traffic saturation scenarios.
By collecting holographic traffic trajectory data, a parameter self-learning model is constructed, and combined with a cloud-edge collaborative control architecture, traffic signals can be controlled in real time.
It achieves real-time and all-weather, all-scenario adaptability of traffic signal control, reduces reliance on professional personnel, and improves data accuracy and control effectiveness.
Smart Images

Figure CN116798227B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traffic signal control technology, and in particular to a parameter self-learning signal control method based on cloud-edge collaboration. Background Technology
[0002] In response to the increasingly serious problems of urban traffic congestion and wasted green light time, intelligent signal control methods are becoming increasingly important with the continuous iteration and upgrading of detection and sensing technologies. Current traffic control methods mainly fall into two categories: one is traffic control algorithms based on operations research and control theory, such as fully inductive control, semi-inductive control, and adaptive control. Their basic principle is to input data such as traffic flow and queue length, and then use a control model to generate control parameters such as signal cycle and green light ratio. While these methods have achieved some success at many intersections, they heavily rely on various traffic parameter configurations, require extensive manual experience from professional traffic personnel, and necessitate repeated parameter adjustments for different intersections, making them difficult to generalize. Furthermore, when traffic flow deviates from historical patterns and undergoes sudden changes, fixed parameter settings lead to unreasonable control schemes. The other category is self-generating control strategy algorithms based on machine learning and deep learning. These algorithms extract intersection operational characteristics, set optimization objectives, and directly generate signal schemes. However, they rely too heavily on training samples, and the results achieved on test sets are often difficult to replicate in real-world environments.
[0003] A Chinese patent document, CN113643553A, discloses a "Multi-intersection Intelligent Traffic Signal Control Method and System Based on Federated Reinforcement Learning." This method models real intersections, uses Cityflow traffic simulation software to simulate urban traffic and flow, and applies reinforcement learning algorithms to each reinforcement learning agent. It controls traffic signals in real-time based on intersection traffic flow conditions. Furthermore, it incorporates a cloud-edge collaborative federated reinforcement learning framework, introducing gradient sharing and parameter transfer processes similar to federated learning, achieving good control results in terms of average vehicle travel time. However, this patent relies heavily on traffic simulation software for obtaining actual road condition information, making it difficult to achieve real-time, all-weather, and all-scenario application. Additionally, the reinforcement learning agents primarily calculate pressure on each directional road through observation, without employing more suitable intersection observation methods and traffic parameters to achieve parameter self-learning, thus failing to adapt to various traffic flow conditions and traffic saturation scenarios. Summary of the Invention
[0004] This invention aims to solve the problems of traffic signal algorithm control being heavily reliant on traffic parameter configuration, excessively dependent on training samples, and difficult to achieve real-time, all-weather, and all-scenario application.
[0005] The above technical problems are solved by the following technical solution: a parameter self-learning signal control method based on cloud-edge collaboration, comprising:
[0006] S1. Collect traffic holographic trajectory data, determine the traffic status of each phase, and control the duration of traffic lights based on the relationship between selected traffic parameters;
[0007] S2. Extract holographic sensing control parameters from traffic holographic trajectory data, define the parameter self-learning action space, and use the intersection cell structure construction method to construct the state matrix to define the parameter self-learning state space, and train the parameter self-learning model.
[0008] S3. Establish a cloud-edge collaborative control architecture to conduct data interaction and control traffic signals in real time.
[0009] Detection equipment such as radar and integrated radar-visual systems can collect omnidirectional vehicle trajectories at intersections, including vehicle trajectory positions, speeds, and transit times within the detection zone. Traffic-related parameters such as traffic flow, speed, queue length, and traffic density can be extracted from these trajectories. Traditional inductive control extends phases based on vehicle transit time intervals; this method adds the ability to discriminate traffic conditions at each phase, resulting in more real-time data acquisition and more accurate data analysis. Operations research-based optimization control algorithms can improve the accuracy of the required full dataset; therefore, edge sensing signal control methods based on holographic data can be used to achieve real-time traffic control. The parameter self-learning method obtains a series of training parameters after training all data, and then predicts the values of new samples based on these training parameters. At this point, it no longer relies on previous training data, and the parameter values are fixed. By extracting the parameters required for holographic sensing control from the collected holographic data using a parameter self-learning method, a state space is extracted from the intersection holographic data. Using the holographic sensing control parameters as the action space and queuing balance as the action evaluation index, a parameter self-learning model for holographic sensing control is established. This eliminates reliance on traffic simulation software and reduces the need for numerous professionals to configure and repeatedly adjust parameters based on manual experience. Since the application of the holographic data-based edge sensing control algorithm requires real-time data acquisition and immediate response control, it places significant demands on message transmission latency. Furthermore, the parameter self-learning method requires a large amount of sample data and high computing power, and the self-learning process involves trial and error. Directly learning at actual intersections can easily create traffic hazards. Therefore, a cloud-edge collaborative control architecture can be built to adapt to the method for traffic signal control. This cloud-edge collaborative control system includes a message transmission and data interaction mechanism from the cloud to the edge computing node, a control-driven synchronization mechanism, and a cloud-edge link management and anomaly degradation mechanism. This forms a cloud-edge collaborative traffic signal control system that integrates cloud strategy generation, parameter learning, edge data interaction, and real-time control.
[0010] Preferably, in step S1, the traffic holographic trajectory data includes traffic state information for each phase, and the traffic state information includes traffic density and headway. Controlling the duration of the traffic lights according to the relationship between the selected traffic parameters includes: S101, judging the traffic state of the intersection based on the traffic density of each phase according to traffic flow theory.
[0011] S102. Set the headway threshold and calculate the headway for the current phase.
[0012] S103. Based on the relationship between the headway and the headway threshold, traffic density, and holographic sensing algorithm, determine whether it is necessary to extend the green light time.
[0013] Traffic density is one of the key parameters for determining lane status. After capturing the traffic status of each phase using data acquisition equipment, the traffic density corresponding to each phase can be calculated based on "Traffic Flow Theory" published by People's Transportation Press in September 2002. Related knowledge of traffic density includes theoretical congestion density and optimal density. Headway is a core element for conditional judgment in the holographic sensing algorithm. Headway represents the time interval from the current moment to the arrival of the previous vehicle. If no vehicle has arrived in the current phase, its initial value is the time interval from the current moment to the start of the phase. A headway threshold is preset for comparison. A smaller headway indicates a higher frequency of vehicle arrivals, requiring a corresponding extension of the phase's green light time. If the headway is greater than the preset threshold, it indicates sparse vehicle arrivals, making it unnecessary to extend the phase's green light time; the next phase can be switched after the minimum green light time. The calculated traffic density of each phase and the headway of the current phase are then substituted into the holographic sensing algorithm flow for traffic light duration control.
[0014] As a preferred approach, the holographic sensing algorithm determines the following: When the traffic density of other phases is less than the theoretical congestion density, if the headway of the current phase exceeds the headway threshold, the green light time is not extended; if the headway of the current phase does not exceed the headway threshold, the green light time is extended. When the traffic density of other phases equals the theoretical congestion density, if the traffic density of the current phase does not exceed the optimal density, a phase switch is performed; if the traffic density of the current phase exceeds the optimal density, the green light time of each phase is extended sequentially. When the traffic density of other phases has not reached the theoretical congestion density, it indicates that the actual road conditions are relatively smooth, and it is only necessary to determine whether to extend the green light time based on the relationship between the headway of the current phase and the headway threshold. Generally, when other phases are close to the theoretical congestion density and the traffic density of the current phase is low, it is necessary to switch phases in a timely manner to allow the congested phase to proceed, and the headway of the current phase is not considered at this time. When the traffic density of the current phase exceeds the optimal density, it indicates that multiple phases at the intersection are congested, and it is necessary to allow each phase to proceed sequentially to disperse the vehicles. Edge sensing control uses holographic sensing algorithms to scientifically and flexibly control different traffic conditions at intersections, and promptly alleviate potential pressure on various directional roads.
[0015] Preferably, in step S2, the parameter self-learning action space and state space are defined, including:
[0016] S201. Based on the holographic trajectory data of traffic at the intersection, a vehicle state matrix is constructed using the intersection cellular structure construction method.
[0017] S202. Construct an intersection signal status matrix based on the current actual traffic signal situation at the intersection;
[0018] S203. The convolutional neural network algorithm is used to extract the feature vector of the current time step from the result of the superposition of two matrices. The feature vector of each time step is serialized as the signal periodic feature. Then, the Transformer algorithm is used to integrate the signal periodic features into a feature vector, which is defined as the state space.
[0019] S204. Under traffic saturation scenarios, the green light limit extension time that ensures the queuing balance of each phase is learned, and the sequence of green light limit extension times of each phase is defined as the action space.
[0020] S205. The DDPG algorithm is used to generate learning parameters, and the parametric model is trained based on the defined state space and action space.
[0021] The parameter self-learning model is established by extracting the state space from the holographic data of the intersection, taking the holographic sensing control parameters as the action space, and the queuing balance as the action evaluation index. The intersection lane information is converted into a vehicle state matrix by the designed intersection cell structure construction method, and the intersection signal state matrix is constructed by combining the actual situation of the current intersection traffic lights. Convolutional neural network is a type of feedforward neural network that includes convolution calculation and has a deep structure. It is one of the representative algorithms of deep learning. The two two-dimensional state matrices of the same size are superimposed to form a 2*m*N three-dimensional feature matrix. Then, CNN (convolutional neural network) is used to extract the features of the three-dimensional matrix and flatten it into a one-dimensional feature vector to represent the features at the current moment. The CNN algorithm can be applied as described in the literature [1] Wang Xiaopu, Huo Jianqing, Liu Tonghuai. Recognition method of handwritten digits by neural network for extracting feature information using related convolution operation [J]. Acta Automatica Sinica, 1996(01):123-125.DOI:10.16383 / j.aas.1996.01.020. In order to integrate the signal cycle features into a feature vector, the Transformer algorithm can be used to introduce an attention mechanism, which can discover the data of key sequence nodes such as the end of the red light and the end of the green light, and can better extract the feature vector of traffic operation of the entire signal cycle. When applying the Transformer algorithm, the method described in the literature [1] Zhang Yingjun, Bai Xiaohui, Xie Binhong. CNN-Transformer feature fusion multi-target tracking algorithm [J / OL]. Computer Engineering and Applications: 1-14 [2023-06-16]. http: / / kns.cnki.net / kcms / detail / 11.2127.TP.20230411.1035.002.html can be used. In the parameter extraction process, the convolutional neural network algorithm is first used to extract the feature vectors of the two state matrices containing real-time information, and then the Transformer algorithm integrates the signal cycle features into a feature vector. The resulting state space can well represent the actual traffic conditions of the intersection, and provide a guarantee for the accuracy of subsequent edge control. As described in the preceding steps, single-point holographic sensing control can effectively reduce the problem of idle traffic at intersections. However, in saturated scenarios, the green light extension time is extended to the maximum. Under congested conditions, it's impossible to guarantee load balance across phases. The goal of parameter learning is to ensure that the maximum green light extension time guarantees queuing balance across phases in saturated scenarios. A sequence of maximum green light extension times for each phase is used as the parameter self-learning action space to achieve balanced resource allocation. Finally, the DDPG algorithm is used to generate the learning parameters. DDPG is a deep deterministic policy gradient algorithm that mainly introduces experience replay and a dual-network approach. It can learn based on a continuous action space and achieves better convergence than Actor-Critic.It mainly includes four neural networks: Actor current network, Actor target network, Critic current network and Critic target network. The DDPG algorithm can be implemented in the manner described in reference [1] Zhai Jianwei. Research on Algorithm and Model Based on Deep Q Network [D]. Suzhou University, 2017. The defined action space and state space will be trained according to this learning network. The trained parameter model will be applied to the actual environment. The limit extension time will be generated according to the current traffic state as the control parameter of the holographic sensing edge control algorithm in S1. The test training model can achieve better results in the real environment, thereby achieving a stable and excellent traffic control effect.
[0022] Preferably, in step S201, the vehicle state matrix is constructed using the intersection cell structure construction method, including: S20101, dividing the maximum detection range of a lane in the intersection into several cells of fixed length, wherein the cells contain the vehicle's position information and speed value;
[0023] S20102. Set the cell state of each lane to row and the lanes at each intersection to column;
[0024] S20103. Normalize the vehicle speed value in each cell to the maximum speed limit and use it as the value at that position. Depending on whether there is a parking speed, use different fixed values to represent it, thereby constructing the vehicle state matrix.
[0025] The maximum detection range of a lane at the intersection is divided into a finite number of cells, each 7m long. The vehicle position and speed information contained within each cell describes the real-time state of the current lane. The cell states of each lane are then arranged as rows, and all lanes at the intersection as columns, forming a matrix. For ease of analysis, the vehicle speed value of each cell is normalized to the maximum speed limit and used as the value for that position. Since the results are the same for no vehicles and a speed of 0 in actual calculations, a -1 is used to distinguish if a vehicle is stopped and its speed is 0. The intersection signal state matrix is constructed based on the actual traffic light conditions at the intersection. The feature vectors extracted from the processed dual matrix containing real-time traffic information provide an important foundation for real-time traffic control.
[0026] Preferably, in step S202, constructing the intersection signal state matrix involves representing the different states of traffic lights for each lane with corresponding fixed values, thus constructing an intersection signal state matrix of the same size as the vehicle state matrix. Setting the intersection signal state matrix to be the same size as the traffic state matrix facilitates subsequent overlay processing and feature vector extraction.
[0027] Preferably, in step S204, the queuing balance index is the variance of queuing lengths in each direction between intersections. The intersection queuing balance index D(q) is the queuing lengths {q1, q2, ..., q} in each direction between intersections. n The variance between intersections is used to achieve a balanced allocation of resources in all directions.
[0028] Preferably, in step S3, a cloud-edge collaborative control architecture is established, including:
[0029] S301. Data interaction is carried out through the message transmission channel from the cloud to the edge computing node. The data exchanged through the message transmission channel includes traffic status information and parameter models trained in the cloud.
[0030] S302. Establish a control drive synchronization mechanism. The control drive includes a program that converts the signal scheme into a program that can interact with the control device and send control commands through a communication protocol. Synchronization is the determination of the control method and priority of the edge generation scheme and the original central light control command after the edge computing node is established.
[0031] S303, Application Cloud-Edge Link Management and Abnormal Degradation Control, Cloud Center Unified Management Information includes Edge Device Information, Abnormal Degradation Situations include Edge Device Abnormalities and Intersection Communication Abnormalities with Cloud Center.
[0032] Cloud-edge collaborative control includes message transmission mechanisms, data interaction mechanisms, control-driven synchronization mechanisms, cloud-edge link management, and anomaly degradation mechanisms from the cloud to edge computing nodes. The message transmission and data interaction mechanisms ensure normal interaction of various data types. Control is achieved by converting signal schemes into programs that can interact with control devices and send control commands via communication protocols, synchronously determining the control methods and priorities of edge-generated schemes and central control commands. Cloud-edge link management and anomaly degradation control are applied, with the central terminal uniformly managing all edge computing nodes, including device IDs, IPs, and port information. Control commands are sent through the links, and the central terminal can directly communicate with the devices when they malfunction. Even when communication between the edge computing node and the central terminal is abnormal at intersections, the edge computing nodes can autonomously control the signal equipment, ensuring intersection communication efficiency and improving the stability of the traffic signal control system. By establishing cloud-based strategy generation, parameter learning, and edge data interaction, a real-time cloud-edge collaborative traffic signal control system is formed.
[0033] Preferably, in step S3, the traffic state information undergoes preliminary preprocessing to ensure a consistent data format. The initial state input for the parameter model trained in the cloud includes this traffic state information. Edge nodes access multi-source data, including radar trajectory data, real-time vehicle passage data, queue length data, traffic density data, etc. This data undergoes preliminary preprocessing to convert it into a unified format, facilitating transmission through the message transmission channel and ensuring communication efficiency. This data is then uploaded to the cloud as the initial state input for model training. The model trained in the cloud is downloaded to the edge computing nodes via the message transmission channel, enabling data interaction.
[0034] The beneficial effects of this invention are as follows: Based on traditional inductive control that extends the phase according to the time interval between passing vehicles, this invention uses holographic traffic trajectory data to determine the traffic state of each phase, prioritizing the passage of congested phases, resulting in strong real-time control capabilities. The holographic inductive control, based on operations research optimization methods and combined with the intersection cell structure construction method designed in this invention, employs a parameter self-learning method, enabling it to adapt to various traffic flow conditions and saturation scenarios, achieving all-weather, all-scenario application of traffic control. A cloud-edge collaborative control system is established, fully utilizing the characteristics of cloud-edge computing to ensure the real-time performance of the holographic inductive algorithm and avoid the computational power and trial-and-error costs required for parameter learning. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of traffic feature extraction on the lane according to the present invention.
[0036] Figure 2 The intersection periodic traffic parameter feature extraction map of the present invention.
[0037] Figure 3 Cloud-edge control flowchart of the present invention Detailed Implementation
[0038] Example: This example provides a parameter self-learning signal control method based on cloud-edge collaboration, including:
[0039] S1. Collect traffic holographic trajectory data and perform edge sensing signal control;
[0040] S2. Extract holographic sensing control parameters from holographic data using a parameter self-learning method;
[0041] S3. Establish a cloud-edge collaborative control architecture, with parameter self-learning and data interaction to control traffic signals in real time.
[0042] In S1, traffic holographic trajectory data can be obtained using detection and acquisition equipment such as radar, integrated radar-visual equipment, and UAV 3D capture. State parameters such as vehicle flow, vehicle speed, queue length, and traffic density are extracted from the traffic state information for edge sensing signal control. This invention adds traffic state discrimination for each phase to the traditional sensing control that extends the phase based on the time interval of passing vehicles. For control algorithms based on operations research optimization, the accuracy of the required full data can be improved; therefore, the following edge sensing control method based on holographic data can be selected, including...
[0043] S101. Based on traffic flow theory, calculate the traffic density of each phase;
[0044] S102. Set the headway threshold and calculate the headway for the current phase.
[0045] S103. Based on the relationship between the headway and the headway threshold, traffic density, and holographic sensing algorithm, determine whether it is necessary to extend the green light time.
[0046] In S101, during actual calculations, the traffic density of each phase can be obtained by referring to "Traffic Flow Theory" published by People's Transportation Press in September 2002. Traffic density-related knowledge includes theoretical congestion density and optimal density. In S102, a headway threshold is set. During the calculation, the headway represents the time interval between the current moment and the arrival of the previous vehicle. If no vehicle has arrived in the current phase, its initial value is the time interval between the current moment and the start of the phase. In practice, to ensure safety, the shortest headway is generally taken as a travel distance of about 2 seconds. In S103, when the traffic density of other phases has not reached the theoretical congestion density, the green light time is extended based on the headway of the current phase. When other phases are close to the theoretical congestion density, and the current phase density is low, the phase needs to be switched in time to allow the congested phase to proceed; at this time, the headway of the current phase is not considered. When the current density exceeds the optimal density, it indicates that multiple phases at the intersection are congested, and the phases need to be released one by one to disperse the vehicles. The actual holographic sensing algorithm flow is as follows:
[0047] At the beginning of each cycle, S10301 initializes the green light time for each phase;
[0048] S10302 has reached the initial green light time and is beginning to assess whether the green light time needs to be increased.
[0049] S10303 determines whether the traffic density of the current phase is greater than the optimal density, then proceeds to S10304; if the threshold is met, it determines whether there are other phases with a traffic density greater than the congestion density, and if so, switches the phase; if not, it proceeds to S10304.
[0050] S10304 calculates the current headway t.h That is, the current time minus the arrival time of the previous vehicle. If no vehicle has arrived in the current phase, its initial value is the current time minus the start time of the phase, which is then compared with the headway threshold t. veh Size:
[0051] If t h ≥t veh Determine whether the green light time has reached the total green light time. If not, wait for the next moment and repeat S10303; if it has reached the total green light time, proceed to S10306.
[0052] If t is not satisfied h ≥t veh If the current green light time of the phase is increased by one unit of extended green light time, it is determined whether the value exceeds the maximum green light time: if it does not exceed the maximum green light time, the total green light time is increased by one extended time, and S10303 is repeated at the next moment; if it exceeds the maximum green light time, it goes to S10305.
[0053] S10305: Determine whether the current green light time of the phase has reached the maximum green light time: If it has, update the total green light time, wait for the next moment, and return to S10304; if the maximum green light time has been reached, go to S10306.
[0054] S10306: The phase ends after passing through yellow flash and all red. Determine whether all phases within the cycle have been traversed. If not, proceed to the next phase and return to Step 1; if they have been traversed, the cycle ends.
[0055] In S2, Figure 1 and Figure 2 An exemplary schematic diagram of the feature extraction method for the holographic sensing control parameters of the present invention is provided. The process of extracting holographic sensing control parameters includes:
[0056] S201. Based on the holographic data of the intersection, construct the vehicle state matrix using the intersection cellular structure construction method;
[0057] S202. Construct an intersection signal status matrix based on the current actual traffic signal situation at the intersection;
[0058] S203. The convolutional neural network algorithm is used to extract the feature vector of the current time step from the result of the superposition of two matrices. The feature vector of each time step is serialized as the signal periodic feature. Then, the Transformer algorithm is used to integrate the signal periodic features into a feature vector, which is defined as the state space.
[0059] S204. Under traffic saturation scenarios, the green light limit extension time that ensures the queuing balance of each phase is learned, and the sequence of green light limit extension times of each phase is defined as the action space.
[0060] S205. The DDPG algorithm is used to generate learning parameters, and the parametric model is trained based on the defined state space and action space.
[0061] In S201, to obtain the parameter self-learning state space, this invention uses an intersection cell structure construction method to construct the vehicle state matrix. The intersection at a certain moment is constructed as a cell structure, and the maximum detection range of a lane in the intersection is divided into a finite number of cells with a length of 7m. The vehicle position and speed information contained in each cell can describe the real-time state of the current lane. For example... Figure 1 As shown, the cell state of each lane is represented as a row, and all lanes at the intersection are represented as columns. The vehicle speed value of each cell is normalized to the maximum speed limit and used as the value at that position, constructing the intersection vehicle state matrix. Assuming the number of lanes is m, the matrix size is m*N. In S202, traffic operation is affected not only by vehicle states, but also by the current signal scheme, which significantly restricts the intersection's traffic efficiency. Therefore, an intersection signal state matrix needs to be constructed. Referring to the vehicle state matrix constructed in S201, the intersection signal state matrix can be constructed based on the traffic status of each lane. For example, a green light value of 1 and a red light value of 0 can be defined, and the matrix size is also m*N. In S203, the state matrices constructed in S201 and S202 can describe the current operating state of the intersection, from which features can be extracted for subsequent parameter learning tasks. Figure 2The intersection state feature extraction process of the present invention is given by way of example. Two two-dimensional matrices of the same size are superimposed to form a three-dimensional feature matrix of 2*m*N. The features are extracted by CNN (convolutional neural network) and flattened into a one-dimensional feature vector to represent the features at the current moment. Since the operation of the intersection is continuous, the features at each moment have a time series correlation. Therefore, all the states from the first second to the end of a signal cycle are serialized as a complete feature of a signal cycle. When extracting the feature vector of traffic operation of a whole signal cycle, Transformer can be used to introduce an attention mechanism and define the feature vector of traffic operation of a whole signal cycle as the state space S. The CNN algorithm mentioned above can be applied as described in the literature [1] Wang Xiaopu, Huo Jianqing, Liu Tonghuai. A neural network method for recognizing handwritten digits by extracting feature information using correlation convolution operation [J]. Acta Automatica Sinica, 1996(01):123-125. DOI:10.16383 / j.aas.1996.01.020. The Transformer algorithm can adopt the content described in reference [1] Zhang Yingjun, Bai Xiaohui, Xie Binhong. CNN-Transformer Feature Fusion Multi-Target Tracking Algorithm [J / OL]. Computer Engineering and Applications: 1-14 [2023-06-16]. http: / / kns.cnki.net / kcms / detail / 11.2127.TP.20230411.1035.002.html. In S204, the action space is defined as:
[0062] A={t max,1 , t max,12 , ..., t max,n}
[0063] Among them, t max,i Let i be the limit extension time for phase i, where i = 1, 2, ..., n, and n is the number of phases. We need to learn the limit extension time for the green light for each phase. Single-point holographic sensing control can effectively reduce the problem of idle traffic at intersections. In saturated scenarios, the green light extension time will be extended to the limit, which cannot guarantee load balance among phases under congested conditions. The goal of parameter learning is to ensure that the limit extension time for the green light in saturated scenarios guarantees queue balance among phases. The intersection balance index D(q) represents the queue lengths {q1, q2, ..., q...} in each direction between intersections. n The variance between nodes. In S105, learning parameters are generated based on DDPG. DDPG mainly introduces experience replay and a dual network, enabling learning based on a continuous action space and achieving better convergence than Actor-Critic. It mainly includes four neural networks:
[0064] The current Actor network is responsible for the iterative update of the policy network parameters θ, and for selecting the current action A based on the current state S. It is used to interact with the environment to generate S′, R.
[0065] The Actor target network is responsible for selecting the optimal next action A′ based on the next state S′ sampled from the experience replay pool. Network parameters θ′ are periodically copied from θ.
[0066] The Critic current network is responsible for iteratively updating the value network parameters w and calculating the current Q-value Q(S,A,w). The target Q-value yi = R + γQ′(S′,A′,w′);
[0067] The Critic target network is responsible for calculating the Q′(S′,A′,w′) part of the target Q-value. The network parameter w′ is periodically copied from w.
[0068] A portion of the target Q-value is calculated based on the empirical replay pool and S′, A′ provided by the target Actor network. Since this part is only for evaluation, it is still completed in the Critic target network. After the Critic target network calculates a portion of the target Q-value, the current Critic network calculates the target Q-value, updates its network parameters, and periodically copies the network parameters to the Critic target network. Furthermore, the current Actor network also updates its network parameters based on the target Q-value calculated by the current Critic network and periodically copies the network parameters to the Actor target network. Regarding the determination of the loss function for iterative training of deep learning models:
[0069] For the current Critic network, the mean square error is calculated as follows:
[0070]
[0071] For the current network of the Actor:
[0072]
[0073] The DDPG algorithm involved in the actual calculation can be adopted as described in reference [1] Zhai Jianwei. Research on Deep Q-Network Algorithm and Model [D]. Suzhou University, 2017. Based on the action space and state space defined in the aforementioned steps, the parameter model is trained, and the trained parameter model is applied to the actual environment. The limit extension time is generated according to the current traffic state as the control parameter for holographic sensing edge control.
[0074] In S3, cloud-edge collaborative control includes message transmission mechanism from cloud to edge computing node, data interaction mechanism, control-driven synchronization mechanism, cloud-edge link management, and anomaly degradation mechanism. Figure 3This invention provides a cloud-edge collaborative control system that establishes a message transmission channel and data interaction mechanism from the cloud to edge computing nodes. Edge nodes uniformly access multi-source data, such as radar trajectories, real-time vehicle passage, queue lengths, and traffic density, perform preliminary preprocessing, convert the data into a unified format, and upload it to the cloud as the initial input for model training. The cloud then downloads the trained model to the edge computing nodes through the message transmission channel. Figure 3 As shown, the parameter self-learning method is deployed on the central platform, and a simulated intersection with the same channelization as the actual intersection is established using a traffic simulator for model training. The trained model is downloaded to the edge computing nodes, which then generate parameters for the control algorithm based on the actual traffic environment of the intersection. A control-driven synchronization mechanism is established, whereby the control driver is a program that converts the signal scheme into a program that can interact with the control equipment through a specific communication protocol and send control commands. After the edge computing nodes are established, the central control commands for traffic lights also need to be retained; therefore, it is necessary to clarify the control methods and priorities of the edge-generated schemes and the central control commands. A cloud-edge link management and anomaly degradation mechanism is established, with the central end uniformly managing all edge computing nodes, including the management of device IDs, IPs, ports, and other information. Control commands are sent to the link, and when an edge device malfunctions, the central end can directly communicate with the device; when communication between the intersection and the central end is abnormal, the edge computing nodes can also autonomously control the signal equipment.
Claims
1. A parameter self-learning signal control method based on cloud-edge collaboration, characterized in that, include: S1. Collect traffic holographic trajectory data, determine the traffic status of each phase, and control the duration of traffic lights based on the relationship between selected traffic parameters; When the traffic density of other phases equals the theoretical congestion density, if the traffic density of the current phase is less than the optimal density, a phase switch is performed; otherwise, the green light time of each phase is extended one by one. S2. Extract holographic sensing control parameters from traffic holographic trajectory data, define a parameter self-learning action space, and construct a state matrix using an intersection cell structure construction method to define a parameter self-learning state space, then train the parameter self-learning model. Based on the intersection traffic holographic trajectory data, construct a vehicle state matrix using an intersection cell structure. According to the actual situation of the current intersection traffic lights, represent different states of the traffic lights in each lane with different values, constructing an intersection signal state matrix of the same size as the state matrix. Use a CNN algorithm to extract the current-time feature vector from the superimposed matrix results, taking the feature vector sequence at each time moment as periodic features, and integrate them into the feature vector of the state space using the Transformer algorithm. Learn the balanced green light limit time for each phase under saturation scenarios, and use the sequence of limit times as the action space. Generate learning parameters and train the model. S3. Establish a cloud-edge collaborative control architecture to conduct data interaction and control traffic signals in real time.
2. The parameter self-learning signal control method based on cloud-edge collaboration according to claim 1, characterized in that, In step S1, the traffic holographic trajectory data includes traffic state information for each phase, including traffic density and headway. Controlling the signal light duration based on the relationship between selected traffic parameters includes: S101. Based on traffic flow theory, determine the traffic state of the intersection according to the traffic density of each phase; S102. Set the headway threshold and calculate the headway for the current phase. S103. Based on the relationship between the headway and the headway threshold, traffic density, and holographic sensing algorithm, determine whether it is necessary to extend the green light time.
3. The parameter self-learning signal control method based on cloud-edge collaboration according to claim 2, characterized in that, In step S1, the determination based on the holographic sensing algorithm is as follows: when the traffic density of other phases is less than the theoretical congestion density, if the headway of the current phase exceeds the headway threshold, the green light time is not extended; if the headway of the current phase does not exceed the headway threshold, the green light time is extended.
4. The parameter self-learning signal control method based on cloud-edge collaboration according to claim 1, characterized in that, In step S2, the definition of the parameter self-learning action space and state space includes: S201. Based on the holographic trajectory data of traffic at the intersection, a vehicle state matrix is constructed using the intersection cellular structure construction method. S202. Construct an intersection signal status matrix based on the current actual traffic signal situation at the intersection; S203. The convolutional neural network algorithm is used to extract the feature vector of the current time step from the result of the superposition of two matrices. The feature vector of each time step is serialized as the signal periodic feature. Then, the Transformer algorithm is used to integrate the signal periodic features into a feature vector, which is defined as the state space. S204. Under traffic saturation scenarios, the green light limit extension time that ensures the queuing balance of each phase is learned, and the sequence of green light limit extension times of each phase is defined as the action space. S205. The DDPG algorithm is used to generate learning parameters, and the parametric model is trained based on the defined state space and action space.
5. A parameter self-learning signal control method based on cloud-edge collaboration according to claim 1 or 4, characterized in that, In step S2, the construction of the vehicle state matrix using the intersection cell structure construction method includes: S20101. Divide the maximum detection range of a lane in the intersection into several cells of fixed length, and the cells contain the vehicle's position information and speed value. S20102. Set the cell state of each lane to row and the lanes at each intersection to column; S20103. Normalize the vehicle speed value in each cell to the maximum speed limit and use it as the value at that position. Depending on whether there is a parking speed, use different fixed values to represent it, thereby constructing the vehicle state matrix.
6. The parameter self-learning signal control method based on cloud-edge collaboration according to claim 4, characterized in that, In step S2, the index for queue balance is the variance of queue lengths in each direction between intersections.
7. The parameter self-learning signal control method based on cloud-edge collaboration according to claim 1, characterized in that, In step S3, establishing the cloud-edge collaborative control architecture includes: S301. Data interaction is carried out through the message transmission channel from the cloud to the edge computing node. The data exchanged through the message transmission channel includes traffic status information and parameter models trained in the cloud. S302. Establish a control drive synchronization mechanism. The control drive includes a program that converts the signal scheme into a program that can interact with the control device and send control commands through a communication protocol. Synchronization is the determination of the control method and priority of the edge generation scheme and the original central light control command after the edge computing node is established. S303, Application Cloud-Edge Link Management and Abnormal Degradation Control, Cloud Center Unified Management Information includes Edge Device Information, Abnormal Degradation Situations include Edge Device Abnormalities and Intersection Communication Abnormalities with Cloud Center.
8. The parameter self-learning signal control method based on cloud-edge collaboration according to claim 7, characterized in that, In step S3, the traffic state information is preprocessed with the same data format, and the initial state input of the parameter model trained in the cloud includes the traffic state information.
Citation Information
Patent Citations
Multi-intersection intelligent traffic signal lamp control method and system based on federal reinforcement learning
CN113643553A
Holographic intersection signal control autonomous optimization method
CN113658439A
Single-point traffic signal control method based on intersection holographic data
CN115691167A
Traffic signal phase timing control method and system based on cloud side-end cooperation
CN116246474A