A method and software for intelligent traffic signal control optimization based on Markov decision process

By using an intelligent traffic signal control method based on the Markov decision process and constructing models of state space and action space, the problem of poor timeliness of traditional traffic signal control is solved, and refined dynamic adjustment and high adaptability to complex traffic flows are achieved.

CN115547050BActive Publication Date: 2025-10-17YUNKONG ZHIXING (SHANGHAI) AUTOMOTIVE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211244345.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-10-17
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

Traditional traffic signal control methods have poor timeliness, cannot adapt to complex and changing traffic flow conditions, and lack the ability to learn independently and perform dynamic optimization.

Method used

An intelligent traffic signal control method based on the Markov decision process is adopted. By acquiring real-time traffic flow data, a traffic signal control model in state space and action space is constructed to predict the traffic flow conditions at the next moment and dynamically adjust the control strategy.

Benefits of technology

It achieves higher adaptability and timeliness, provides more refined signal control strategies, and can adapt to complex and changing traffic flow conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115547050B_ABST
    Figure CN115547050B_ABST
Patent Text Reader

Abstract

The application provides a Markov decision process-based intelligent traffic signal control optimization method and software, and the method comprises the following steps: acquiring real-time traffic flow data of a target road intersection; predicting a traffic flow condition of the target road intersection at a next time point through a traffic signal control model constructed based on a Markov decision process according to the real-time traffic flow data, and obtaining a prediction result; and executing a control strategy on the traffic signal according to the prediction result; wherein a construction factor of the traffic signal control model comprises a state space and an action space; the state space is used for representing a state of vehicle flow of each period of the target road intersection; and the action space is used for representing a signal control strategy of each period of the target road intersection under different states. At least the technical problem that the existing traffic signal control method has poor timeliness and cannot adapt to the current complex and changeable traffic flow condition can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent transportation, and particularly relates to an intelligent traffic signal control optimization method and software based on a Markov decision process. BACKGROUND

[0002] Traffic signal control is to allocate road right to traffic flow in different directions on the basis of intersection channelization. Traffic flow is separated in time, and traffic signal control parameters such as green ratio are automatically optimized and adjusted through advanced traffic models and algorithms, so that traffic signals of a group of intersections or intersections in a region are realized optimal coordinated control, and finally the purpose of safely and effectively organizing traffic flow through intersections is achieved. Traffic signal control technology has experienced four major development stages:

[0003] The first stage is mechanical traffic signal control technology;

[0004] The second stage is fixed timing traffic signal control technology. The signal cycle and green ratio of a single signal machine are mainly determined by experience and historical traffic data, and automatic cycle control and multi-period control are realized by a computer;

[0005] The third stage is inductive traffic signal control technology. The control mode of signal display time of a single signal machine is mainly adjusted according to traffic flow data measured by a vehicle detector, and is divided into semi-inductive control (only part of the phases of the intersection have inductive requests) and full-inductive control (all phases of the intersection have inductive requests);

[0006] The fourth stage is line control technology (coordinated control of traffic signals of multiple adjacent intersections on a road) and surface control technology (coordinated control of all traffic signals in a region). There are three kinds of fixed timing coordinated control system, real-time scheme selection coordinated control system and real-time adaptive coordinated control system.

[0007] The mature line control system at present mainly includes PASSER-II and MAXBAND of the United States. PASSER-II is a line control system coordination software that combines the mutual influence method of Blox and the "unequal width optimization model" of Liddell. The optimal ratio of traffic demand-capacity of each intersection is determined, and then the green ratio of each signal is determined, and then the optimal signal timing scheme of the widest through band is determined by changing the trial cycle length, phase and time difference. MAXBAND is to optimize the signal time difference according to the "mixed integer programming model" of Liddell under the conditions of given cycle length, green ratio, intersection distance and continuous traffic speed, so as to determine the optimal bandwidth according to different traffic conditions.

[0008] The earliest area control system is TRANSYT, and then more representative ones are SCOOT, SCATS, ACTRA, UTCS, etc. Most of these area control schemes acquire and analyze traffic information through detector timing, generate the best timing scheme with the cooperation of traffic models and optimization programs, and finally send it to the intersection signal machine for implementation. The optimization program uses a small step-by-step asymptotic optimization method to continuously adjust the green ratio, cycle and phase difference three parameters in real time, which not only reduces the calculation amount but also easily tracks and grasps real-time traffic trends.

[0009] However, the inventors have found that at least the following technical problems exist in the related art:

[0010] In the control method of traffic signals provided in the conventional line control technology and area control technology, the control rules are fixed, the timeliness is poor, and it is unable to adapt to the complex and variable situation of the current traffic flow. In addition, although a batch of representative traffic signal control systems such as NATS, HiCon, SMOOTH, etc. have appeared after the HT-UTCS urban traffic signal control system in China, they have the adaptability and real-time performance to different cities, different regions and different traffic flow characteristics, and realize the scheme of coordinated optimization of traffic efficiency, safety and order, which to some extent meets the real-time requirement, but the effect is not very ideal. SUMMARY

[0011] An object of the present application is to provide an intelligent traffic signal control optimization method and software based on Markov decision process, at least to solve the technical problem that the existing control method of traffic signals has poor timeliness and is unable to adapt to the complex and variable situation of the current traffic flow.

[0012] To achieve the above object, some embodiments of the present application provide a traffic signal control method, which comprises: acquiring real-time traffic flow data of a target road intersection; predicting the traffic flow condition of the target road intersection at the next time according to the real-time traffic flow data through a traffic signal control model constructed based on Markov decision process, to obtain a prediction result; and executing a control strategy on the traffic signal according to the prediction result; wherein the construction factors of the traffic signal control model include a state space and an action space; the state space is used to represent the state of vehicle flow of each period at the target road intersection; and the action space is used to represent the signal control strategy of each period at the target road intersection under different states.

[0013] Some embodiments of the present application also provide a traffic signal control device, which comprises: one or more processors; and a memory storing computer program instructions which, when executed, cause the processor to execute the method as described above.

[0014] Some embodiments of the present application also provide a computer readable medium having stored thereon computer program instructions executable by a processor to implement the traffic signal control method.

[0015] Compared with the prior art, in the traffic signal control scheme provided by the embodiments of the present application, real-time traffic flow data of a target road intersection is obtained, and then a traffic signal control model constructed based on a Markov decision process is used to predict traffic flow conditions of the target road intersection at the next time according to the real-time traffic flow data, to obtain a prediction result. Finally, a control strategy is executed on the traffic signal according to the prediction result. The construction factors of the traffic signal control model include a state space and an action space. The state space is used to represent the state of vehicle flow of each period at the target road intersection, and the action space is used to represent the signal control strategy of each period at the target road intersection in different states. Since the definitions of the state space and the action space are added to the traffic signal control model constructed based on the Markov decision process, on the one hand, each state of the state of vehicle flow of each period at the target road intersection can be exhausted, and meanwhile, in the description of the traffic flow state, the correlation between the selected feature information and other feature information can be added on the basis of the selected feature information, so that the data dimension is higher and the state description is more detailed. On the other hand, since the action space is added, the dynamic adjustment of the control strategy can be realized while the traffic flow conditions are predicted by the traffic signal control model constructed based on the Markov decision process. It can be seen that the scheme provided by the embodiments of the present application has higher adaptability and better timeliness to the actual traffic flow conditions, and is beneficial to providing more detailed signal control strategies. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A flowchart of a traffic signal control method provided by the embodiments of the present application;

[0017] Figure 2 A flowchart of another traffic signal control method provided by the embodiments of the present application;

[0018] Figure 3 A flowchart of another traffic signal control method provided by the embodiments of the present application;

[0019] Figure 4 An example schematic diagram of a traffic signal control method provided by the embodiments of the present application;

[0020] Figure 5 A schematic diagram of a traffic signal control model constructed based on a Markov decision process provided by the embodiments of the present application;

[0021] Figure 6 A schematic diagram of training a traffic signal control model based on a Markov decision process provided in an embodiment of the present application;

[0022] Figure 7 A schematic diagram of the structure of a device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] The following terms are used in this document.

[0025] Markov Decision Process: The full English name is Markov Decision Process, referred to as "MDP", which is used to simulate the stochastic strategies and rewards that can be achieved by intelligent agents in an environment where the system state has Markov properties.

[0026] Deep RL: Deep reinforcement learning refers to combining the perception capabilities of deep learning with the decision-making capabilities of reinforcement learning. It can directly control based on the input image and is an artificial intelligence method that is closer to human thinking.

[0027] Green-to-light ratio: refers to the proportion of time available for vehicle passage within a traffic light cycle. That is, the ratio of the effective green light time in a phase to the cycle length.

[0028] Road network: refers to a road system within a certain area that is composed of various roads that are interconnected and interwoven into a network.

[0029] In the related art, the method of timing control and signal control of vehicle driving in the traditional line control model and surface control model has fixed control rules, which are not suitable for the current complex and variable traffic flow (traffic flow "abnormalities" caused by different sudden traffic events at complex intersections, different time traffic flow changes, such as holidays, morning and evening rush hours, major events, etc.), and the higher the complexity of the intersection, the lower the fitting degree of the relatively fixed signal control rules and the actual physical world traffic flow model. In a high-dimensional state space, the traditional RL algorithm cannot effectively calculate the value function and policy function for each state. Although some linear function approximation methods are proposed in RL to solve the state space problem, their ability is still limited. In high-dimensional and complex systems, traditional RL methods cannot learn the feature information of the environment for efficient function approximation. The actual traffic flow is complex and changes quickly, and the types of traffic flow feature information are more, and the traditional algorithm is limited in the description of the traffic flow state space.

[0030] The timing scheme obtained by the previous intelligent traffic signal control system is mostly based on the assumed state and the relatively fixed configuration mode relying on historical experience. For example, based on a specific state, assuming that the phase sequence or cycle length is unchanged, only adjusting the green ratio, etc., the scheme is single and cannot be flexibly adjusted according to the real-time traffic situation. At the same time, there is a lack of long-term data monitoring, effect evaluation feedback of the scheme execution effect, and no ability of autonomous learning and dynamic optimization.

[0031] The embodiment of the present application provides a traffic signal control method, which comprises the following steps: acquiring real-time traffic flow data of a target road intersection; then, according to the real-time traffic flow data, a traffic signal control model constructed based on a Markov decision process is used to predict the traffic flow condition of the target road intersection at the next moment, and a prediction result is obtained; finally, a control strategy is executed on the traffic signal according to the prediction result; wherein the construction factors of the traffic signal control model include a state space and an action space; the state space is used to represent the state of vehicle flow of each period at the target road intersection; and the action space is used to represent the signal control strategy of each period at the target road intersection under different states.

[0032] In the embodiments of the present application, the definition of state space and action space is added in the traffic signal control model constructed based on Markov decision process. Therefore, on the one hand, each state of the state of vehicle flow of each period of the target road intersection can be exhausted, and meanwhile, in the description of the traffic flow state, the correlation description of the selected feature information and other feature information can be added based on the selected feature information, so that the data dimension is higher and the state description is more detailed; on the other hand, since the action space is added, the dynamic adjustment of the control strategy can be realized while the traffic flow condition is predicted by the traffic signal control model constructed based on Markov decision process. In summary, the scheme provided by the embodiments of the present application has higher adaptability and better timeliness to the actual traffic flow condition, and is beneficial to provide more detailed signal control strategy.

[0033] As shown in Figure 1 The traffic signal control method provided by the embodiments of the present application can include the following steps:

[0034] In step S101, real-time traffic flow data of a target road intersection is acquired.

[0035] In step S102, according to the real-time traffic flow data, the traffic flow condition of the target road intersection at the next time is predicted by a traffic signal control model constructed based on Markov decision process, and a prediction result is obtained.

[0036] In step S103, according to the prediction result, a control strategy is executed on the traffic signal.

[0037] The construction factors of the traffic signal control model include state space and action space; the state space is used to represent the state of vehicle flow of each period of the target road intersection; and the action space is used to represent the signal control strategy of the target road intersection in different states of each period.

[0038] For step S101, specifically, the real-time traffic flow data of the target road intersection of the city can be acquired by a cloud control basic platform. Here, the real-time traffic flow data of the target road intersection can be lane-level real-time traffic flow data of the target road intersection.

[0039] For step S102, specifically, when the traffic signal control model constructed based on Markov decision process is used to predict the traffic flow condition, the construction factors of the traffic signal control model can include state space and action space; the state space is used to represent the state of vehicle flow of each period of the target road intersection; and the action space is used to represent the signal control strategy of the target road intersection in different states of each period.

[0040] The state can be understood as information reflecting the state change of the vehicle flow at the target road intersection in each period, such as the speed of the vehicle at the target road intersection, the density of the vehicle at the target road intersection, and the like. Of course, the information used to represent the state can be selected and flexibly set according to actual needs, and is not limited specifically herein. The action can be all possible signal control timing schemes corresponding to the state that the vehicle in the target road intersection can take when the vehicle is in a certain state.

[0041] In some examples, feature information reflecting the features of the traffic flow data of the road intersection can be extracted, and then the state of the vehicle flow at the target road intersection in each period is described according to the feature information.

[0042] For step S103, specifically, a control strategy is performed on the traffic signal according to the prediction result output by the traffic signal control model constructed based on the Markov decision process.

[0043] It can be found that, compared with the related art, in the embodiments of the present application, the definition of the state space and the action space is added to the traffic signal control model constructed based on the Markov decision process. Therefore, on the one hand, each state of the state of the vehicle flow at the target road intersection in each period can be exhausted, and meanwhile, in the description of the traffic flow state, the correlation description of the selected feature information and other feature information can be added on the basis of the selected feature information, so that the data dimension can be higher and the description of the state can be more detailed; on the other hand, since the action space is added, the dynamic adjustment of the control strategy can be realized while the traffic flow condition is predicted by the traffic signal control model constructed based on the Markov decision process. In summary, the scheme provided in the embodiments of the present application has higher adaptability and better timeliness to the actual traffic flow condition, and is beneficial to providing more detailed signal control strategies.

[0044] In some embodiments of the present application, the state space is determined according to the speed feature of the vehicle at the target road intersection and the density feature of the vehicle; and the action space is determined according to the phase sequence of the traffic signal at the target road intersection, and the cycle length and green ratio of the corresponding signal lamp under different phase sequences.

[0045] Specifically, the correlation of the speed feature and the density feature can be constructed by a Markov chain to calibrate the vehicle flow at the target road intersection, and this is used as the state space.

[0046] The phase sequence of the traffic signal refers to the sequence of traffic flow in different directions. Thus, in some examples, the signal control timing scheme output by the traffic signal control model based on the Markov decision process can include a variable phase sequence, a cycle length, and a green ratio three-dimensional array.

[0047] It can be found that, compared with the related art, the traffic signal control method provided in the present application calibrates the vehicle flow in the target road intersection by constructing the correlation between the speed feature and the density feature, so that the description of the state in the state space in the present embodiment is more fine; the action space is determined according to the phase sequence of the traffic signal of the target road intersection and the cycle length and the green ratio of the corresponding signal lamp under different phase sequences, so that the timing scheme of the intersection signal can be dynamically adjusted, and the optimized timing scheme obtained is more flexible and better fits the traffic condition.

[0048] In some embodiments of the present application, the determination method of the state space can include determining the speed feature and the density feature according to the real-time traffic flow data; determining the flow feature of the vehicle according to the speed feature and the density feature; and determining the state space according to the flow feature of the vehicle.

[0049] Specifically, referring to Figure 2 The method of the present embodiment can include the following steps:

[0050] Step S201: determining the speed feature and the density feature according to the real-time traffic flow data;

[0051] Step S202: determining the flow feature of the vehicle according to the speed feature and the density feature;

[0052] Step S203: determining the state space according to the flow feature of the vehicle.

[0053] For step S201, the speed feature and the density feature of the vehicle can be extracted by analyzing the real-time traffic flow data obtained by the cloud control basic platform.

[0054] For step S202, the correlation between the speed feature of the vehicle, the density feature of the vehicle, and the flow feature of the vehicle can be established by a large amount of real-time traffic flow data.

[0055] In some examples, the correlation can be established by the following formula:

[0056] Q=KV

[0057] Wherein, K represents the density feature of the vehicle, V represents the speed feature of the vehicle, and Q represents the flow feature of the vehicle.

[0058] For step S203, the flow feature determined according to step S202 can be recorded to obtain the state space.

[0059] It can be found that, compared with the related art, in the scheme provided by the embodiments of the application, the correlation description of the speed feature information of the vehicle and the density feature information of the vehicle is added on the basis of the flow feature information of the vehicle in the process of determining the state space, so that the data dimension can be higher, and the description of the state can be more accurate. That is, in the description of the state of the traffic flow, the correlation description of the selected feature information and other feature information is added on the basis of the selected feature information, so that the data dimension is increased, and the purpose of more accurate description of the state is achieved.

[0060] In some embodiments of the application, the determining the state space according to the flow feature of the vehicle can include: dividing the flow feature of the vehicle according to the state of the vehicle flow of the target road intersection in each period; and obtaining the state space according to each divided flow feature of the vehicle.

[0061] Referring to Figure 3 The method of the embodiments of the application can include the following steps:

[0062] Step S301, dividing the flow feature of the vehicle according to the state of the vehicle flow of the target road intersection in each period.

[0063] Step S302, obtaining the state space according to each divided flow feature of the vehicle.

[0064] Specifically, the flow feature of the vehicle can be divided into several small intervals, and each flow interval represents one of the states of the state of the vehicle flow of the target road intersection in each period.

[0065] In some examples, the following sequence can be established:

[0066]

[0067] Wherein, k is the intersection number, and t is the continuous time. Each sequence can be divided into n states, which is an n-order multi-state Markov chain. For example, if the sequence includes 10 states, the flow is divided into 10 equal parts from 0 to X, and the value of n is 10. Here, n is an integer greater than or equal to 1, and the value of n can be determined according to the actual data.

[0068] The sequence established in the embodiments of the application can be the form of the state space.

[0069] It can be found that, compared with the related art, the embodiment of the present application provides a specific implementation manner for determining the state space, which is advantageous to improve the processing efficiency of data by changing the complex and unordered state values into a relatively ordered array through a sequence.

[0070] In some embodiments of the present application, the control strategy performed on the traffic signal according to the prediction result can include: receiving feedback information issued by the traffic signal control model after the control strategy performed on the traffic signal according to the prediction result; and adjusting the control strategy performed on the traffic signal according to the feedback information.

[0071] Specifically, in the embodiment of the present application, the traffic signal control model constructed based on the Markov decision process is tested, and the test effect reaches the correct rate required for actual use before being put into use in the intelligent traffic signal control system. After being put into use, running data in the actual dynamic environment can be received, feedback information issued by the traffic signal control model can be calculated according to the feedback information, and training and testing can be performed, so that the traffic signal control model can perform the optimal control strategy under the real-time changing traffic condition.

[0072] Compared with the related art, in the scheme provided by the embodiment of the present application, after the control strategy performed on the traffic signal according to the prediction result, the feedback information issued by the traffic signal control model is received, and the control strategy performed on the traffic signal is adjusted according to the feedback information, which is advantageous to better adapt to the complex and changeable traffic condition.

[0073] In some embodiments of the present application, the construction factors of the traffic signal control model can further include: a mapping relationship between a service level of an intersection and an average delay time of a vehicle; the average delay time is used to represent the time lost by the vehicle waiting for a red light at the intersection; and the receiving of the feedback information issued by the traffic signal control model can include: after the traffic signal control model determines the average delay time of the vehicle according to the mapping relationship between the service level of the intersection and the average delay time of the vehicle, receiving the feedback information issued by the traffic signal control model according to the average delay time of the vehicle.

[0074] In some examples, the mapping relationship between the average delay time and the service level of the intersection can be established, as shown in Table 1:

[0075]

[0076] Table 1

[0077] In this example, the service level of the intersection is divided into six levels.

[0078] Compared with the related art, the method provided by the embodiments of the present application considers the average delay duration, evaluates the control result of the control strategy after the traffic signal executes the control strategy, and further facilitates optimization of the control strategy.

[0079] In some embodiments of the present application, the adjusting the control strategy executed on the traffic signal according to the feedback information can include: adjusting the control strategy executed on the traffic signal according to the feedback information by a deep reinforcement learning algorithm.

[0080] Specifically, the optimal solution of the control method of the traffic signal provided by the embodiments can be obtained by using a neural network model of deep reinforcement learning as a nonlinear function approximator. Meanwhile, the data set and the test set can be obtained by combining the feedback information with actual sampling by the deep reinforcement learning algorithm, which is beneficial to autonomous optimization of the data set and the test set, and thus beneficial to iterative optimization of the control strategy.

[0081] Deep RL is one of the most successful artificial intelligence models and the machine learning paradigm closest to human learning mode. It combines deep neural networks and reinforcement learning, making function approximation more effective and stable, especially for high-dimensional and infinite state problems. Specifically, for high-dimensional state space, the deep RL method is superior to the traditional RL method, and by training a deep neural network to learn the optimal strategy or value function, the value function and the strategy function can be effectively calculated for each state. In terms of action space, the policy-based deep RL method is more suitable for continuous action space than the value-based deep RL method. For discrete action space, the controller usually uses DQN and its variants because their structure is simpler compared with the policy-based method. In a large state space, different neural network structures, such as convolutional neural network (CNN) and recurrent neural network (RNN), can be used to train the reinforcement learning algorithm.

[0082] Compared with the related art, the embodiments of the present application provide a specific implementation manner of adjusting the control strategy executed on the traffic signal according to the feedback information, which is beneficial to flexible and diverse implementation of the manner of adjusting the control strategy of the embodiments of the present application.

[0083] In some embodiments of the present application, the adjusting the control strategy executed on the traffic signal according to the feedback information by the deep reinforcement learning algorithm can include: evaluating the traffic state at a future time by the deep reinforcement learning algorithm to obtain an evaluation result; and adjusting the control strategy executed on the traffic signal in combination with the evaluation result and the feedback information.

[0084] Specifically, the feedback information can include an immediate reward. Wherein, based on the control strategy output by the traffic signal control model, the feedback information output after performing an action is the immediate reward; the feedback information output on the future impact of vehicle flow, that is, the evaluation of the traffic state at the future time, obtains an evaluation result. The evaluation result can also be understood as an additional reward. Then, the control strategy for the traffic signal is adjusted in combination with the evaluation result and the feedback information.

[0085] Wherein, the additional reward can also be referred to as a future reward.

[0086] Compared with the related art, in the method provided by the embodiments of the present application, the current immediate reward and the additional reward affecting the future are considered before adjusting the control strategy for the traffic signal, so as to facilitate the adjusted control strategy to further adapt to the complex and changeable traffic conditions.

[0087] In summary, the control method of the traffic signal provided by the embodiments of the present application can obtain the real-time traffic flow data of the target road intersection, and then can predict the traffic flow condition of the target road intersection at the next time according to the real-time traffic flow data through the traffic signal control model constructed based on the Markov decision process, to obtain a prediction result. Finally, the control strategy for the traffic signal is executed according to the prediction result; wherein, the construction factors of the traffic signal control model include a state space and an action space; the state space is used to represent the state of the vehicle flow at each period of the target road intersection; and the action space is used to represent the signal control strategy of the target road intersection at each period under different states. Since the definition of the state space and the action space is added in the traffic signal control model constructed based on the Markov decision process, on the one hand, each state of the state of the vehicle flow at each period of the target road intersection can be exhausted, and meanwhile, in the description of the traffic flow state, the correlation description of the selected feature information and other feature information can be added on the basis of the selected feature information, so as to make the data dimension higher and the state description more fine; on the other hand, since the action space is added, the dynamic adjustment of the control strategy can be realized while the traffic signal control model constructed based on the Markov decision process predicts the traffic flow condition. It can be seen that the scheme provided by the embodiments of the present application has higher adaptability and better timeliness to the actual traffic flow condition, and is beneficial to provide more fine signal control strategy.

[0088] Briefly, refer to Figure 4As shown, in the embodiments of the present application, the intelligent traffic signal control of urban road intersection based on Markov decision process (①) is performed, the traffic flow characteristic information is extracted through the cloud-based platform, the flow, speed and density are defined as the model state space, the variable phase sequence, cycle length and green ratio are defined as the action space, the state transition relationship is defined based on the traffic flow big data to predict the traffic condition at the next time (②). The mapping relationship between the intersection service level and the average delay time is established to obtain the immediate reward of the actual environment feedback after the execution of the strategy, and the additional reward of the strategy based on the prediction of the traffic condition at the future time is evaluated (②), and the reward function is established. The neural network model of deep reinforcement learning, i.e., the traffic signal control model based on Markov decision process (④), is used to optimize the strategy based on the reward value in the process of continuous interaction between the system and the environment (⑤), and the optimal solution of the model is tested by collecting the long-term data set, and the optimization effect of the scheme is verified.

[0089] In addition, in order to facilitate everyone to understand the scheme, an example of the traffic signal control model based on Markov decision process is provided, and the traffic signal control model is described in detail as follows.

[0090] In some examples, the traffic signal control model based on Markov decision process can be represented as a five-tuple, such as:

[0091] MDP = <S, A, R, P, γ>

[0092] Wherein, S represents the state space, which is used to represent the non-empty finite set of all possible states of vehicles in the target road intersection;

[0093] A represents the action space, which is used to represent the non-empty finite behavior set of actions that can be performed at time t under the state s ∈ S;

[0094] R represents the reward function, which is used to represent the reward obtained when the state S t is executed under the action a t , the state S t is transferred to S t+1 ;

[0095] P represents the transition function, which is used to represent the transition probability of the state S i to S j when the action a is executed under the state S t+1 ; The state-action mapping at a certain time is defined as the distribution matrix P(s t+1 |s t , a t ) of the next state S t+1 , which is the state transition function;

[0096] γ represents a discount factor, used to characterize the importance of immediate rewards and additional rewards.

[0097] The following five aspects involved in the quintuple are described respectively:

[0098] I. State space S

[0099] The state space S represents a non-empty finite set of all possible states of vehicles in the target road intersection. By analyzing the characteristics of the traffic flow data of the cloud control basic platform, the relevant information of the speed characteristics and the density characteristics can be extracted, and the correlation between the vehicle speed, the vehicle density and the vehicle flow can be established. The flow characteristics of the intersection are recorded in the form of time series data, which is defined as the state space of the model.

[0100] Specifically, the correlation between speed, density and flow can be established based on a large amount of traffic flow data of the cloud control basic platform.

[0101] Q = KV

[0102] In the formula, K is the density characteristic information of the vehicle, V is the speed characteristic information of the vehicle, and Q is the flow characteristic information of the vehicle.

[0103] Further, the flow can be divided into small intervals, and each flow interval represents a state, and the following sequence is established:

[0104]

[0105] Wherein, k is the intersection number in the road network, and t is the continuous time. Each sequence can be divided into n states, which is an n-order multi-state Markov chain. For example, assuming that the flow is divided into 10 equal parts from 0 to X, the sequence includes 10 states, and the value of n is 10; here, n is an integer greater than or equal to 1, and the value of n can be determined according to the actual data.

[0106] II. Action space A

[0107] The action space A represents a non-empty finite set of actions that can be performed when the state s ∈ S at time t. That is, based on the Markov model, the control strategy adopted for multiple intersections in the road network, i.e. the signal control timing scheme, is defined as the action set. Wherein:

[0108] Assuming that k represents the intersection number in the road network, agent k represents the traffic signal controller of the kth intersection. In a road network with n intersections, the set of signal controllers of the intersections is:

[0109] K = {agent0, agent1, … agent k, …agent n}

[0110] Generally, the traffic signal controller of each intersection can only execute a set of signal control timing schemes at the same time. Assuming A k is the signal control timing scheme of the kth intersection signal controller agent k , a k is the action executed by the kth intersection traffic signal controller agent k , then a k ∈A k .

[0111] Based on the traffic signal control model constructed based on the Markov decision process in the embodiments of the present application, the phase sequence of the traffic signal and the corresponding green light phase length under different phase sequences are defined as actions, that is, a high-dimensional array including the green light length under different phase sequence strategies, which can be expressed as:

[0112] A={b k (T m +c k ), m=1, 2, 3, 4; k=1, 2……}

[0113] Wherein, m is the phase sequence of the traffic signal, T is the cycle length of the signal light, b k is the green ratio, and c k is the length parameter of the increase and decrease.

[0114] In this example, the traffic signal control model constructed based on the Markov decision process mainly considers the following four green light phases:

[0115] North-South Green (NSG);

[0116] East-West Green (EWG);

[0117] North-South Advance Left Green (NSLG);

[0118] East-West Advance Left Green (EWLG).

[0119] Wherein, the specific values of the green ratio and the length parameter of the increase and decrease can be set according to the traffic signal control experience value.

[0120] In practical applications, A should satisfy the following conditions:

[0121] 1) The cycle length of the signal light and the length of the increase and decrease under different phase sequences are all integers. It can be understood that the cycle length of the general signal light and the green light length do not have decimals.

[0122] 2) The value of the cycle length of the signal lamp should be within a preset range. For example, according to an empirical value setting, the value of the cycle length of the signal lamp can be between 60 s and 180 s.

[0123] 3) The increase / decrease length should be less than a preset value. For example, according to an empirical value setting, the value of the increase / decrease length should be less than or equal to 120 s.

[0124] III. Transfer function P

[0125] The transfer function P represents the transition probability from state S i to state S j when the action a is performed at time t. The distribution matrix P(s t+1 |s t+1 , a t , t) of the state-action mapping at a certain time to the next state S t is the state transition function.

[0126] Specifically, there is a certain time and space correlation between the traffic flow of different intersections at different times of the road network. The traffic flow at a certain intersection at the current time is affected by multiple factors in the previous time step, including the upstream intersection flow situation and the traffic signal control mediation situation at the previous time. In this model, a high-order multivariate Markov chain is applied to determine the state transition relationship and construct the state transition matrix.

[0127] The traffic flow sequence of the intersection can be represented as Here, the vehicle flow change at each intersection and each period is regarded as a state, and the state probability distribution of the jth sequence at time r+1 depends on the probability distribution of all sequences at times r, r-1, …, r-n+1, which can be represented as:

[0128]

[0129] wherein and is the h-step transition probability matrix from the state of the jth sequence at time r-h+1 to the state of the rth sequence at time r+1.

[0130] Let wherein The high-order multivariate Markov chain can be represented by the following matrix:

[0131]

[0132] wherein

[0133]

[0134]

[0135] Next, estimate parameters Q should satisfy X = XQ, and a method for minimizing ||X-XQ|| is needed to solve Consider the optimization problem:

[0136]

[0137] Since where, Then the first prediction vector of the vector for each k is Then the predicted value of the state at time t+1 is: where,

[0138] Next, define the parameters P ij and the state transition matrix:

[0139]

[0140] In the formula, is the flow of intersection k in a certain period within the n-order partition.

[0141] Based on the directed connection graph of the real road network, define the weighted transition probability K ij , which represents the probability of the state transitioning from s to s' when the strategy is executed, which is equal to the sum of the probability of executing all actions in the state and the probability of the corresponding action enabling the state to transition from s to s':

[0142] K ij = P ij · C ij · α ij

[0143] where, α ij represents the weight of the flow change between the i intersection and the j intersection.

[0144] Those skilled in the art can understand that in the Markov model, the closer the historical state is to the present moment, the greater the influence on the decision of the next moment state. That is, the position access point closer to the flow prediction moment has a higher weight, and vice versa, the corresponding weight is lower. Therefore, in some examples, it can be assumed that α i is the weight of intersection i, and α j is the weight of intersection j. Then when i < j, α i > α j . Therefore, for α i > 0, 1 ≤ i ≤ k, α i is a non-decreasing function.​

[0145] The data analysis and statistics can be based on a cloud-based platform, and the traffic time series between i intersection and j intersection is selected as an empirical value. Then, in the range of a road network containing m intersections, the transition matrix X m×n The matrix is:

[0146]

[0147] Through the state transition matrix, the state S t at the next time after the decision is executed and the action a is taken t+1 . That is, through the intersection state S kt at time t, the decision a kt is executed, and the sequence matrix of the intersection state in the road network at time t+1 is obtained.

[0148] Four, reward function R

[0149] Specifically, the reward function R(S t , a t , S t+1 ) represents the reward obtained when the system moves to the state S t after taking the action a t in the state S t+1 .

[0150] For example, the intersection service level can be divided into 6 levels as shown in Table 1.

[0151] Assume that the average delay time is d, d=d1+d2, where:

[0152]

[0153]

[0154] d1—uniform delay time, i.e., the delay time generated by uniform vehicle arrival, unit: s / pcu;

[0155] d2—random additional delay, i.e., the additional delay time generated by random vehicle arrival and causing oversaturation period, unit: s / pcu;

[0156] C—period length;

[0157] λ—green ratio of the calculated lane;

[0158] x—saturation of the calculated lane;

[0159] CAP—throughput of the calculated lane (pcu / h);

[0160] T—analysis period duration. In some examples, it can be taken as 0.25h;

[0161] e—single intersection signal control type correction coefficient. In some examples, the timing control is taken as 0.5; the inductive control e varies with the saturation and green light extension time, and the value range is preferably 0.04-0.5.

[0162] Wherein each data in the above formula can be provided by the cloud platform in real time.

[0163] In some examples, the reward after performing an action can include an immediate reward and an additional reward that has an impact on the future. The immediate reward R t is obtained after performing an action according to the strategy for k intersections in the road network at time t.

[0164] R t = R(S t , π(S t ))

[0165] 1) After regulation, the service level rises, and the return reward R t = 1 is returned.

[0166] 2) After regulation, the service level falls, and the return reward R t = -1 is returned.

[0167] 3) After regulation, the service level basically remains unchanged, and the return reward R t = 0 is returned.

[0168] 4) In other cases, it is indicated that the regulation effect is not obvious, and it is not enough to judge the advantages and disadvantages, and the return reward R t = 0 is returned.

[0169] Based on the above rules, the cumulative reward R t can be represented by the following formula:

[0170]

[0171] The goal of MDP is to find the best strategy π* to maximize the cumulative reward expectation E(R t |s, π), where the cumulative reward R t is:

[0172]

[0173] Five, discount factor γ,

[0174] The discount factor γ controls the importance of immediate rewards and future rewards, and the value can be between 0 and 1, i.e. γ ∈ (0, 1). Choosing a small γ means that the agent's action pays more attention to real-time rewards.

[0175] After that, the neural network model of deep reinforcement learning can be used as a nonlinear function approximator to solve the optimal solution of intelligent traffic signal control.

[0176] Specifically, in reinforcement learning, the goal of the agent is to learn an action selection policy π to guide the action selection of the agent to maximize the expectation, that is, to select a series of actions to obtain the most average reward, which means that at this time the agent set executed by the system for all intersections of the road network is the best decision sequence. The system starts from time t, and according to the state S t , the decision execution a t is made, and the next time state S t+1 and the next decision execution action a t+1 are obtained, and finally the actions of all decision times of each intersection of the road network are traversed to obtain a decision sequence, and each decision sequence can be regarded as a round of MDP.

[0177] The possible strategy π(s,a) of executing action a in state s is represented by the system:

[0178] π(sa) = P[S t = sA t = a]

[0179] The action-state value function Q π (s,a) is defined to evaluate the expected reward of the strategy, which represents the mathematical expectation of the reward obtained by the system under the initial condition of state s according to the policy function π sequence decision, that is, expressed as:

[0180]

[0181] According to the Bellman equation, the action-state value function of the tth decision is only related to the action-state value function of the t-1th decision, so the action-state value function can be simplified as:

[0182]

[0183] The optimal solution is found by maximizing the action-state value function q(s,a) through the greedy strategy, that is, the decision behavior of the system starting from any state s can satisfy that the action-state value function Q π (s,a) reaches the maximum value:

[0184]

[0185] In some examples, a reinforcement learning algorithm based on collaborative Q-learning can be used to obtain the optimal policy π. The influence of adjacent intersections is considered by integrating the Q-value transfer strategy into deep learning, and an MLP evaluation network is constructed to automatically extract features from the original state and approximate the optimal Q-value.

[0186] In some cases, the target network can be introduced into the model to assist the intersection evaluation network and calculate it according to the following formula:

[0187]

[0188] The target Q value can be defined as:

[0189] The action of the agent depends not only on its own Q value, but also on the Q value of the adjacent intersection.

[0190] After transferring the Q values ​​of the adjacent intersection agents, the Q value of each intersection i can be updated according to the following formula:

[0191]

[0192] Among them, θ i and are the parameters of the evaluation network and the target network respectively, N is the number of adjacent intersections of intersection i, ω i,j is the weight of the Q value from intersection j. In practical applications, different weights can be set according to the influence of adjacent intersection j on intersection i.

[0193] In some examples, ω can be determined by the following formula i,j Value:

[0194]

[0195] Among them, c1 and c2 are proportional coefficients, d ij represents the distance from the i-th intersection to the j-th intersection, T ij represents the traffic flow from the i-th intersection to the j-th intersection. Specifically, the closer to the adjacent intersection, the greater the traffic flow, and the greater the impact.

[0196] In addition, the loss function of each agent can be determined according to the following formula:

[0197]

[0198] Where m is the batch size, Status The best target Q value for all actions, is the output of the evaluation network.

[0199] Referring to Figure 5 , Figure 5 Fig. 1 shows a schematic diagram of the above process:

[0200] Specifically, at each time step t, the state s observed by the agent (i.e. the traffic sequence at the current moment) is input into the evaluation network. The agent selects an action a to be performed (i.e. the timing scheme of the intersection signal control, which can be determined according to the phase sequence, cycle length and green light length selected according to the intersection traffic at the current moment) using the Q value output by the evaluation network according to the greedy strategy, and the agent obtains a reward r and enters the next state s';

[0201] The information {s, a, r, s'} obtained by the current intersection agent interacting with the environment at each time step is stored in the experience pool M. During the training process, a certain batch size of samples is randomly selected from M each time, and the samples are trained by the double Q network. The two networks have the same results but different parameters, and the action selection and policy evaluation are separated.

[0202] When calculating the reward function of the current Q value and the target Q value, the corresponding experience is sampled from the experience pool of the upstream intersection agent, and the optimal Q value of the downstream intersection agent is calculated using the evaluation network of the upstream intersection agent. The Q value is transferred to the current network to calculate the loss function, and the gradient descent algorithm is used to update the parameters of the timing scheme.

[0203] Then, the training of the evaluation model can be performed on the data collected by the cloud platform at the roadside for a certain period. The historical traffic data collected in a month is used to construct the training set, the validation set and the test set, and the iteration is performed month by month.

[0204] Similarly, in order to facilitate everyone to understand the present scheme, a process example for training a traffic signal control model based on a Markov decision process is also provided.

[0205] Specifically, as shown in Figure 6 For any intersection, assume that the adjacent intersections of the current intersection agent are four, n1, n2, n3 and n4. Then, at time t, the current intersection agent has all its historical data, which can be represented as:

[0206] (s1,a1;s2,a2;…st,at)

[0207] The state-action data set of the four adjacent intersections that need to be obtained at time t is:

[0208] (sn1,an1;sn2,an2;…snt,ant)

[0209] The observation S of the current agent at time t is represented as the union of the above two sets:

[0210] S = (s1, a1; s2, a2…st, at; sn1, an1; sn2, an2; sn3, an3; sn4, an4;)

[0211] For any multi-intersection traffic network, the training steps of the deep reinforcement learning algorithm based on Markov decision process are as follows:

[0212] Step 1: Initialize the road network intersection agent i The state matrix, evaluation network parameter θ i and target network parameter Discount factor γ, experience pool max_size and min_size, target network update step C, initialize the instantaneous reward value r of each agent, and the upper limit of the number of iterations Iter max .

[0213] Step 2: input the observed real-time data into the evaluation network, agent i According to the output value of the evaluation network, a phase action is selected by using the greedy strategy Based on the calculation of the average delay time of the intersection in the real-time data, the service level change is obtained, and the reward is obtained And enter the next state After that, according to t = t + 1, the assignment is made.

[0214] Step 3: store the experience In the experience pool Mi, if the experience pool overflows, delete the old experience data, when the number of experience pools is greater than min_size, start training, enter step 4, otherwise go to step 2.

[0215] Step 4: input the sampled data in the current experience pool Mi into the current evaluation network and target network, calculate the current value function and target value function; sample the corresponding historical traffic data from the experience pool of the adjacent intersection, input the evaluation network of the adjacent intersection, and obtain the transition Q value of the adjacent intersection. Calculate the loss function according to the formula.

[0216] Step 5: update the network weight θ i and of the current intersection, and repeat the above calculation for each intersection agent.

[0217] Step 6: if t < Iter max And s t ≠ terminal (terminal state), go to step 2.

[0218] Afterwards, a test set can be constructed according to the parameters of the trained model, and the test effect can be obtained.

[0219] Further, if the training test effect reaches the correct rate required by actual use, the model algorithm is integrated into the intelligent traffic signal control system, and the intelligent signal control system is used to perform calculation and feedback training in the actual dynamic environment, so that the system can obtain the sequence decision with the maximum cumulative return based on the real-time changing traffic conditions.

[0220] Compared with the related art, in the embodiment of the present application, the state and action value are stored in the deep neural network indexed by s and a by combining deep reinforcement learning, the neural network is updated by constantly interacting with the environment and obtaining the reward function feedback, and finally the state and action value stored in the neural network can correctly guide the intelligent agent to perform the sequence decision with the highest reward value in the environment. The intelligent traffic signal control model constructed at the same time can dynamically adjust the phase sequence and the length of each phase of the intersection signal timing according to the real-time traffic flow information, and has strong adaptability.

[0221] In addition, the embodiment of the present application also provides a device, the structure of the device is as shown in Figure 7 The device includes a memory 11 for storing computer readable instructions and a processor 12 for executing computer readable instructions, wherein when the computer readable instructions are executed by the processor, the processor is triggered to execute the control method of the traffic signal.

[0222] In some examples, the device can be an automatic driving controller.

[0223] The method and / or embodiment in the embodiment of the present application can be implemented as a computer software program. For example, the embodiment of the present disclosure includes a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above-mentioned functions defined in the method of the present application are performed.

[0224] Note that the computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable medium can be, for example but not limited to, a computer-readable storage medium such as a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic disk, an optical fiber, a compact disk (CD), a digital versatile disk (DVD), a Blu-ray disk, a memory stick, a floppy disk, a random access memory (RAM), a read-only memory (ROM), a programmable ROM (EPROM), an erasable and programmable ROM (EPROM), an electrically erasable and programmable ROM (EEPROM), a flash ROM, a flash memory, a solid state disk, a hard disk, a computer-readable storage medium, or any appropriate combination thereof. In the present application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by an instruction execution system, apparatus, or device to execute the program.

[0225] In the present application, the computer-readable signal medium can include a data signal that propagates in a baseband or as part of a carrier wave by any medium, including but not limited to wire, wireline, optical fiber, cable, RF, etc., or any appropriate combination thereof. The computer-readable signal medium can also be any computer-readable medium that can send, propagate, or transport programming code to be used by an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wireline, optical fiber, cable, RF, etc., or any appropriate combination thereof.

[0226] The computer program code for carrying out operations of the present application can be written in one or more programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as C or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0227] The computer readable medium can include a non-transitory computer readable medium (e.g., tangible computer readable storage medium). The above described device or apparatuses implement the methods and / or techniques according to the embodiments of the present application. As another aspect, the embodiments of the present application also provide a computer readable medium, which can be contained in the above described device or apparatus, or exist separately without being assembled into the device or apparatus. The above described computer readable medium carries one or more computer readable instructions, which can be executed by a processor to implement the steps of the above described methods and / or techniques according to the embodiments of the present application.

[0228] As another aspect, the embodiments of the present application also provide a computer readable medium, which can be contained in the above described device or apparatus, or exist separately without being assembled into the device or apparatus. The above described computer readable medium carries one or more computer readable instructions, which can be executed by a processor to implement the steps of the above described methods and / or techniques according to the embodiments of the present application.

[0229] In a typical configuration of the present application, the devices of the terminal and the service network each include one or more processors (CPU), input / output interface, network interface and memory.

[0230] The memory can include non-persistent memory and / or volatile memory, e.g., random access memory (RAM) requiring power to maintain state where data is stored, and / or non-volatile memory, e.g., read-only memory (ROM), flash memory, etc. The memory is an example of computer readable media.

[0231] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented using any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices.

[0232] In addition, the embodiment of the present application further provides a computer program, which is stored in a computer device, so that the computer device executes the method performed by the control code.

[0233] It should be noted that the present application can be implemented in software and / or a combination of software and hardware, for example, can be implemented by using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer readable recording medium, for example, a RAM memory, a magnetic or optical drive or a soft disk and similar devices. In addition, some steps or functions of the present application can be implemented by using hardware, for example, as a circuit cooperating with the processor to execute the respective steps or functions.

[0234] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, but can be implemented in other concrete forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all aspects as exemplary and non-restrictive, and the scope of the present application is defined by the appended claims rather than the above description, and all changes falling within the meaning and range of the equivalent elements of the claims are intended to be embraced in the present application. Any reference signs in the claims should not be considered as limiting the involved claims. In addition, it is obvious that the word "comprise" does not exclude other units or steps, and the singular does not exclude the plural. The plurality of units or devices stated in the device claims can also be implemented by one unit or device by software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

Claims

1. A traffic signal control method, characterized in that: The method comprises: Obtain real-time traffic flow data at the target road intersection; According to the real-time traffic flow data, a traffic signal control model constructed based on a Markov decision process is used to predict the traffic flow condition at the target road intersection at the next moment to obtain a prediction result; executing a control strategy for the traffic signal according to the prediction result; The construction factors of the traffic signal control model include state space and action space; the state space is used to represent the state of vehicle flow at the target road intersection in each time period; the action space is used to represent the signal control strategy of the target road intersection in different states at each time period; The state space is determined based on the speed characteristics and density characteristics of vehicles at the target road intersection; The action space is determined according to the phase sequence of the traffic signals at the target road intersection, and the cycle duration and green-to-signal ratio of the traffic lights corresponding to different phase sequences; The method for determining the state space includes: Determining the speed characteristics and the density characteristics based on the real-time traffic flow data; determining the vehicle flow characteristics based on the speed characteristics and the density characteristics; determining the state space based on the vehicle flow characteristics; The determining of the state space according to the traffic characteristics of the vehicle includes: Classifying the vehicle flow characteristics according to the state of vehicle flow at each time period at the target road intersection; The state space is obtained according to the flow characteristics of each divided vehicle.

2. The method according to claim 1, characterized in that The executing a control strategy for the traffic signal according to the prediction result includes: After executing the control strategy for the traffic signal according to the prediction result, receiving feedback information from the traffic signal control model; Adjusting a control strategy executed on the traffic signal according to the feedback information.

3. The method according to claim 2, characterized in that The construction factors of the traffic signal control model also include: a mapping relationship between the service level of the intersection and the average delay time of the vehicle; the average delay time is used to represent the time lost by the vehicle waiting for the red light at the intersection; The receiving feedback information sent by the traffic signal control model includes: After the traffic signal control model determines the average delay time of the vehicle based on the mapping relationship between the intersection service level and the average delay time of the vehicle, feedback information sent by the traffic signal control model based on the average delay time of the vehicle is received.

4. The method according to claim 3, characterized in that The adjusting the control strategy executed on the traffic signal according to the feedback information includes: The control strategy executed on the traffic signal is adjusted according to the feedback information through a deep reinforcement learning algorithm.

5. The method according to claim 4, characterized in that The adjusting the control strategy executed on the traffic signal according to the feedback information by using a deep reinforcement learning algorithm includes: Using the deep reinforcement learning algorithm, the traffic state at a future moment is evaluated to obtain an evaluation result; The control strategy executed on the traffic signal is adjusted in combination with the evaluation result and the feedback information.

6. A traffic signal control device, characterized in that: The device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the method according to any one of claims 1 to 5.

7. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for evaluating service level of plane signal intersection under mixed traffic environment

    CN101604479A

  • Traffic-light control system and method based on Markovian decision

    CN108597239A

  • Road intersection signal lamp split control method, device and equipment

    CN113963553A