An urban traffic control optimization method based on adaptive dynamic programming
By using an adaptive dynamic programming method, combined with data-driven and neural network iterative solutions, the problem of obtaining model parameters in urban transportation systems is solved, achieving efficient traffic flow optimization control and significantly reducing the total vehicle travel time.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2026-03-13
Smart Images

Figure CN117456734B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traffic control technology, and in particular to an urban traffic control optimization method based on adaptive dynamic programming. Background Technology
[0002] In recent years, rapid urban modernization has led to an exponential increase in transportation demand, resulting in severe traffic congestion in metropolitan areas. Traffic congestion can cause long queues of vehicles on roads, and even more seriously, it can severely reduce the utilization rate of available infrastructure. Therefore, alleviating traffic congestion through reasonable traffic control methods has attracted considerable attention from scholars.
[0003] Modeling urban transportation systems is a complex task. Traffic modeling based on detailed traffic conditions (such as flow, speed, and density) at each road segment and intersection within the urban transportation network is a micro-level modeling approach. This method is often difficult to implement in practice due to the difficulty in obtaining accurate and comprehensive data. To avoid the problem of data scarcity, in recent decades, increasing attention has shifted to research on macro-level holistic modeling of urban transportation networks. This includes the boundary control theory based on a macro-level fundamental graphical model proposed in "Daganzo C F. Urban gridlock: Macroscopic modeling and mitigation approaches[J]. Transportation Research Part B: Methodological, 2007, 41(1): 49-62," which can effectively improve traffic operation in urban networks. It has brought convenience to macro-level urban traffic road network modeling and has aroused great interest among scholars. The macro-basic map is a typical macro model of urban traffic systems and an inherent characteristic of urban traffic road networks. The theory of macro-basic maps brings convenience to the modeling of macro-urban traffic road networks because the macro-basic map model uses the number of vehicles in each region as a bridge to model the macro-traffic area as a whole, and the vehicles in each region are better measurement data.
[0004] Currently, most macro-control research in traffic control uses boundary control methods, treating the urban traffic system as the controlled object and employing various model-based methods, such as model predictive control, linear quadratic integral control, and adaptive boundary control, to optimize traffic flow transfer between traffic areas. However, since urban traffic systems are multivariable, strongly coupled, and nonlinear systems, their model parameters are not easily obtained. Therefore, some model-free control methods have been proposed, such as model-free adaptive control. Model-free adaptive methods discretize the traffic model and solve for the control input step by step. However, the actual traffic system is a continuous system. Therefore, in this invention, we use an adaptive dynamic programming method to treat the traffic system as a continuous system and consider minimizing traffic time and control costs as the objective, thus achieving optimized control of urban macro-traffic flow.
[0005] If the system is linear and the cost function exhibits a quadratic form for the state and control inputs, then according to optimal control theory, the optimal control strategy can be achieved through state feedback, typically by solving the standard Riccati equations. In this case, the optimal control problem is relatively easy to handle. However, when the system is nonlinear or the cost function does not satisfy the quadratic form for the state and control inputs, obtaining the optimal control strategy becomes much more complex. In this case, it is necessary to solve the Hamilton-Jacobi-Bellman equations, which are partial differential equations, and solving them is usually quite challenging.
[0006] However, with the continuous development of reinforcement learning, researchers have begun to explore methods for applying it to optimal control problems. One approach is to use an evaluation-execution structure, namely adaptive dynamic programming, to approximate the Hamilton-Jacobi-Bellman equations. Compared to traditional methods, reinforcement learning has unique advantages in handling such complex optimal control problems. Reinforcement learning allows the system to find the optimal policy through learning and adaptation without explicitly solving the Hamilton-Jacobi-Bellman equations. This method is particularly suitable for systems with nonlinear dynamics or cost functions because it can more flexibly adapt to complex environments and problem settings. Therefore, reinforcement learning provides an innovative approach to addressing various optimal control problems, bringing new possibilities for solving these problems.
[0007] Adaptive dynamic programming (ADP) algorithms, with their unique algorithmic and structural characteristics, offer significant advantages over existing methods in the field of optimal control. One of its greatest breakthroughs lies in addressing optimal control problems where traditional variational theory cannot handle closed-set constraints. Unlike some other methods that only apply to open-set constraints, ADP's uniqueness, compared to the maximum principle, lies in the amount of information it provides. The maximum principle only provides necessary conditions for optimal control problems, while traditional dynamic programming methods and ADP provide sufficient conditions, facilitating more accurate design and implementation of optimal control systems.
[0008] Although dynamic programming methods can become complex in some cases due to the difficulty of solving the Hamilton-Jacobi-Bellman equations and the "curse of dimensionality," adaptive dynamic programming, as an approximate solution to dynamic programming, successfully overcomes these limitations. This makes adaptive dynamic programming excellent for applications involving highly coupled, highly nonlinear, and complex systems, such as intelligent transportation systems, power systems, navigation systems, aircraft control, and communication systems. Adaptive dynamic programming is rapidly developing and expanding, not only making breakthroughs in theoretical research but also demonstrating broad potential in various practical applications. Summary of the Invention
[0009] To address the difficulty in obtaining parameters from urban macro-level traffic network models in practical control, this invention proposes an urban traffic control optimization method based on adaptive dynamic programming. This method can calculate the optimized boundary control law without needing to know the parameters of the macro-level basic traffic model, and it can ensure the continuous stability of the control input.
[0010] The technical solution of this invention: A method for optimizing urban traffic control based on adaptive dynamic programming, specifically including the following steps:
[0011] Step S1: Establish a macro-level urban transportation network model;
[0012] Based on the theory of macroscopic fundamental graphs, a macroscopic urban transportation network model is established using the classical parameters of existing urban macroscopic transportation networks to obtain the number of vehicles n in each region. ij It is worth noting that the method of this invention is based on data-driven model-free boundary control, and the macroscopic basic graph model is only used to obtain the number of vehicles n in each region. ij The calculation of the control law does not use the parameters of the macroscopic basic graphical model.
[0013] The urban macro-transport network model is rewritten as an affine nonlinear system;
[0014]
[0015] where u(t)=[u 12 u 21 ] T , representing the control input, x(t) = [n 11 n 12 n 21 n 22 ] T f(x(t))=[-M 11 +q 11 q 12 q 21 -M 22 +q 22 ] T ;
[0016] Where q ij (t), i, j = 1, 2, represents the traffic demand from region i to destination region j at time t, where q 11 (t), q 22 (t) represents the traffic demand flow within each region, q 12 (t), q 21 (t) represent traffic demand flow outside each region, u 12 (t), u 21 (t) represent the boundary control release ratio from region i to destination region j at time t, and 0≤u ij u ji ≤1; M ij (t) represents the transfer flow from region i to destination region j at time t, M ii (t), M jj The meaning of (t) is the same;
[0017]
[0018] Among them G i (n i (t) is the macro basic map of region i, representing the completed travel flow within region i. The completed travel flow refers to the sum of the flow of vehicles arriving at their destination within region i and thus ending their travel, and the flow of vehicles leaving region i.
[0019] S2: Design the objective function; considering two aspects: system state and control input. For the system state, this invention aims to minimize the total time to travel (TTS) index within the road network of the macroscopic traffic system. For the control input, since it is constrained (between 0 and 1), special processing is needed for the control input portion of the objective function. To address the problem of constrained input control quantity, this invention proposes using a saturator to solve this problem, making the control input u's range unrestricted.
[0020]
[0021] in Represents the unrestricted control input, where A and c are the upper and lower boundaries that guarantee the control input, i.e., the control input v satisfies -A+c≤υ≤A+c; a and b represent the offset and coordinate translation, respectively.
[0022] The part of the objective function concerning the control input is as follows:
[0023]
[0024] Φ(υ)=[φ(υ1), φ(υ2)] T , Φ -1 (u)=[φ -1 (υ1), φ -1 (υ2)] T . It is a positive definite diagonal matrix, r i These are elements in R = diag{r1, r2}, representing the weights for different control outputs; according to optimal control theory, the objective function is defined as (5).
[0025]
[0026] Where Q(x) = Px(t), It is a column vector in which all elements are 1;
[0027] S3: Online iteration method;
[0028] Based on the urban macro-transportation network model and objective function, the optimal control law is designed using optimal control theory, and V*(x(t)) is defined as:
[0029]
[0030] The Hamiltonian equation is derived from the above formula.
[0031]
[0032] By solving To obtain the optimal control law;
[0033]
[0034] However, this calculation method will cause the amount of computation and storage to increase sharply with the increase of the dimension of the state and control, which is the so-called "curse of dimensionality" problem. Therefore, we consider using the strategy iteration method to iterate the calculation to avoid the "curse of dimensionality" problem; the strategy iteration method is used to iterate the calculation of equations (9) and (10);
[0035] S3.1: Strategy Evaluation;
[0036] Within a fixed time period N, for i = 0, 1, 2, ..., solve for V. i (x),
[0037]
[0038] in V represents the value at the i-th iteration. * The partial derivative of (x) with respect to x, μ i (x) represents the control law obtained in the previous iteration;
[0039] S3.2: Strategy evaluation initialization;
[0040] For i = 0, 1, 2, ..., update the control law according to the following formula.
[0041]
[0042] When ||μ i+1 (x(t))-μ i If (x(t))||≤∈, then exit the current iteration loop; ∈ is a set threshold, ∈>0, which serves as the condition for exiting the loop, and will be used as... μ i+1 (x(t)) is the control input for the current time period. It is input into the urban macro-traffic network model, i.e., equation (1), to obtain the system state.
[0043] S4: Introduce a neural network to solve the problem;
[0044] For urban macro-transportation network models Where f(x) and g(x) are unknown, for the total vehicle travel time index within the road network of the urban macro-transportation network model, the value function V(x(t)) is chosen as...
[0045]
[0046] Since the control input is input to the system in the form of state feedback, u(t) ≡ μ(x), which can be obtained according to the definition of definite integral;
[0047]
[0048] Combining equation (12) with system equation (1) and removing f(x), we get
[0049]
[0050] V was fitted using three neural networks: an evaluation network, a class model network, and an execution network. i (x) and μ i (x);
[0051]
[0052]
[0053]
[0054] in and To evaluate the weights of the three neural networks: the network itself, the model network, and the execution network, and The basis functions of the evaluation network, the model network, and the execution network; and Representing the approximation error, the following equation is obtained.
[0055]
[0056] Get V(x) and How the weights are updated;
[0057]
[0058] in
[0059]
[0060] Get μ i (x) How the weights are updated.
[0061]
[0062] The beneficial effects of this invention are:
[0063] This invention employs a data-driven approach for boundary control of macroscopic traffic flow. Three neural networks are introduced to iteratively solve for the optimized control input, eliminating the need for specific parameters of the urban macroscopic traffic network and fully utilizing information from the input and output data of the macroscopic basic traffic system. Furthermore, since the boundary control quantities of macroscopic traffic flow are constrained, a saturator is added to the value function to ensure that the control input obtained through iterative solution is bounded. Compared to traditional nonlinear model predictive control methods, this invention significantly improves the total vehicle travel time index. Attached Figure Description
[0064] Figure 1 Here is a flowchart of the strategy iteration algorithm;
[0065] Figure 2 This is a schematic diagram of the traffic model for the two regions' macro-basic maps.
[0066] Figure 3 Create a map for traffic demand;
[0067] Figure 4 To execute the network weight change graph;
[0068] Figure 5 A graph showing the changes in control inputs for adaptive dynamic programming.
[0069] Figure 6 To obtain the state change diagram of the traffic system using the adaptive dynamic programming method;
[0070] Figure 7 This is a graph showing the change in control input for nonlinear model predictive control.
[0071] Figure 8 The state change diagram of the traffic system is obtained by the nonlinear model predictive control method. Detailed Implementation
[0072] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions. A model-free adaptive dynamic programming algorithm is provided as a strategy for solving the boundary control of macroscopic traffic networks.
[0073] The steps are as follows:
[0074] S1: Establish a macro-level urban transportation network model
[0075] In urban macro-transport network models, a key parameter is the completed travel flow G(n(t)). Based on the properties of the macro-basic graph, in a uniform road network or region, the completed travel flow exhibits a unimodal, low-dispersion relationship with the number of vehicles within that region. This relationship can generally be approximated using piecewise linear functions, quadratic functions, and cubic functions, with the cubic function generally providing the highest fit. To maintain generality and increase accuracy, this invention uses a cubic function to approximate the macro-basic graph model for each region, namely:
[0076] G i (n i (t))=A i ·n i (t) 3 +B i ·n i (t) 2 +C i ·n i (t), A i B i C i These are coefficients related to region i. In this invention, the parameter is selected as...
[0077]
[0078] Where i = 1, 2. are parameters, and the G(n(t)) parameter settings are the same for both regions. It is worth noting that this parameter setting is only for observing the effect of the control input calculated by the adaptive dynamic programming algorithm. The information of this model parameter is not used when the adaptive dynamic programming algorithm calculates the control input.
[0079] like Figure 2 The diagram shows a city's macro-transport system comprising two traffic zones. Based on the principle of traffic flow conservation in each zone, we can derive the dynamic equilibrium equations for the two-zone traffic model as follows:
[0080]
[0081] Where q ij (t), i, j = 1, 2, represents the traffic demand from region i to destination region j at time t, where q 11 (t), q 22 (t) represents the traffic demand flow within each region, q 12 (t), q 21 (t) represent the external traffic demand flow of each region, u 12 (t), u 21 (t) represent the boundary control release ratio from region i to region j at time t, and 0≤u ij uji ≤1. M ij (t) represents the transfer flow from region i to region j at time t, M ii (t), M jj The meaning of (t) is the same, and the calculation method is as follows:
[0082]
[0083] S2: Design the objective function
[0084] Because the control input range in this invention is [0, 1], the system of equations -A+c = 0 and A+c = 1 is used to obtain A = c = 0.5, which satisfies the condition. Furthermore, to ensure that the unrestricted control input maintains high accuracy after passing through the saturator by adjusting the weighting coefficient a and the bias coefficient b, a = 0.5 and b = 0 are chosen as simulation parameters in this invention, and the identity matrix is selected as R.
[0085]
[0086] Simultaneously choose Q(x) = 10 -4 ·Px(t), where 104 is the weight value, the purpose of which is to make the Hamilton-Jacobi-Bellman equation operate on the same order of magnitude and reduce the error.
[0087] S3: Initialize the control law μ0(·)
[0088]
[0089] The strategy iteration requires an initial admissible control law. Only with an initial admissible control law can the system iteratively approach the optimal solution. From a mathematical perspective, the initial admissible control law can be viewed as selecting a suboptimal solution, which itself requires solving a nonlinear partial differential equation. To date, there is still no good method to solve this problem. Therefore, we determine the parameters of the initial admissible control law through multiple experiments and observations.
[0090] S4: Introducing a neural network for solving
[0091] According to the higher-order Weierstrass approximation theorem, a continuous function can be represented by a set of infinite-dimensional linearly independent basis functions. To implement the proposed iterative algorithm, we introduce three neural networks: an evaluation network to approximate the value function V. i (x), an execution network to approximate μ i (x), a type model network approximation The basis function structure is chosen as follows:
[0092]
[0093] in
[0094] X1 = n 11 X2 = n 12 X3 = n 21 X4 = n 22 As a basis function of the neural network, the initial state is designed as x0 = [2600, 2700, 1900, 2000]. T Within a time step N=8s, according to the process Figure 1 The solution is obtained by iterative steps, with the total time step set to 120.
[0095] S5: Simulation Results and Analysis
[0096] To demonstrate the superior control performance of our proposed method, we conducted comparative experiments with nonlinear model predictive control (MMDC) to illustrate its effectiveness. The parameter settings for the MMDC method were as follows: minimizing the total vehicle travel time within the road network was the optimization objective; the control step size and single-step solution step size were set to 15 and 80, respectively. To facilitate comparison of control performance, the total number of solution steps was set to 120, the same as the parameter settings for the adaptive dynamic programming method.
[0097] The total vehicle travel time obtained using the linear model predictive control method is 1.17 * 10^6. 9 The total vehicle travel time obtained using this method is 3.46 * 10^6 seconds. 7 Compared to the classic nonlinear model predictive control method, the total vehicle travel time is reduced by approximately 98%, demonstrating the effectiveness of this method.
[0098] Figure 3 The invention presents the traffic demand settings for two algorithms. The traffic demand variation is set as a piecewise function to represent the different traffic demand in different areas during different time periods. The traffic demand values are set to be smaller at the beginning and end of the time periods, while the traffic demand values are set to be larger in the middle time periods, which can reasonably simulate the traffic demand variation within a traffic cycle.
[0099] Figure 4 The changes of the executor function in each iteration of the adaptive dynamic algorithm are presented. It can be seen that after a period of iteration, the weights of the neural network gradually converge to a stable range. That is, the control input gradually converges to the optimal value.
[0100] Figure 5 and Figure 7The changes in boundary control input obtained using the adaptive dynamic programming algorithm and the nonlinear model prediction method are presented respectively. The simulation results show that the boundary control variable obtained by the adaptive dynamic programming algorithm changes more gradually and eventually converges to a stable interval. However, the boundary control law obtained by the nonlinear model prediction method reaches its maximum value in the third time step, and the change is more drastic. At the same time, it remains unchanged in subsequent time steps, which does not well reflect that the control input should change with the state feedback.
[0101] Figure 6 and Figure 8 The simulation results show the changes in the number of vehicles (i.e., the system state) of the macroscopic traffic system obtained from the boundary control input using both adaptive dynamic programming and nonlinear model prediction methods. The simulation results indicate that the adaptive dynamic programming method results in a smoother change in the number of vehicles, with the number of vehicles in each region eventually stabilizing at around 1000, which is a reasonable value. However, the classical nonlinear model prediction method results in a less smooth change in the number of vehicles, ultimately converging to around 100. Compared to the actual number of vehicles in the macroscopic traffic area, this stable number of 100 is clearly too low and not a reasonable figure.
[0102] The above embodiments are merely illustrative of the implementation methods of the present invention, but should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the protection scope of the present invention.
Claims
1. A method for optimizing urban traffic control based on adaptive dynamic programming, characterized in that, The specific steps are as follows: Step S1: Establish a macro-level urban transportation network model; Based on the theory of macroscopic fundamental graphs, a macroscopic urban transportation network model is established using the classic parameters of existing urban macroscopic transportation networks to obtain the number of vehicles in each region. ; The urban macro-transport network model is rewritten as an affine nonlinear system; (1) ; in Represents control input, ; ; ; in represent t From the area i To the destination area j The traffic demand, of which This represents the traffic demand flow within each region. These represent traffic demand flow outside each region. Represent t From the area i To the destination area j The boundary control release ratio, and ; represent t From the area i Heading to the destination area j The transfer flow, The meaning is the same; (2) ; in That is, the region i A macro-level basic map representing the region i Completed travel traffic within a region; completed travel traffic refers to the total travel within a region. i The flow of people arriving at their destination within the region ends the travel flow and exits the area. i The sum of traffic flow; S2: Design the objective function; Minimize the total vehicle travel time index within the city's macro-transportation network; use a saturator to adjust the control input. u The scope is not limited; (3) ; in This represents an unrestricted control input. A , c It ensures the upper and lower boundaries of the control input, i.e., the control input. satisfy ; a , b These represent the offset and coordinate translation, respectively. The part of the objective function concerning the control input is as follows: (4) ; , It is a positive definite diagonal matrix. yes The elements in the equation represent the weights for different control outputs; according to optimal control theory, the objective function is defined as (5). (5) ; in ,in It is a column vector in which all elements are 1; S3: Online iteration method; Based on the urban macro-transport network model and objective function, the optimal control law is designed using optimal control theory, and defined as follows: for; (6) ; The Hamiltonian equation is derived from the above formula. (7) ; By solving Thus, the optimal control law is obtained; (8) ; The strategy iteration method is used to iteratively calculate the equations (9) and (10) to avoid the "curse of dimensionality" problem; the strategy iteration method is used to iteratively calculate the equations (9) and (10); S3.1: Strategy Evaluation; Within a fixed time period N Inside, for Solve (9) ; in Representing the i In the next iteration right The partial derivatives, The control law obtained from the previous iteration; S3.2: Strategy evaluation initialization; for Update the control law according to the following formula. (10) ; when Then exit the current iteration loop; It is a set threshold. This will serve as the condition for exiting the loop, and will be used as... The control input for the current time period is input into the urban macro-transport network model, i.e., equation (1), to obtain the system state; S4: Introduce a neural network to solve the problem; For urban macro-transportation network models Among them, for the total vehicle travel time index within the road network of the urban macro-transportation network model, a value function is selected. for, (11) ; Since the control input is input to the system in the form of state feedback, According to the definition of a definite integral, we get: (12) ; Equation (12) is combined with system equation (1) and deleted. ,have to (13) ; The evaluation network, class model network, and execution network were fitted respectively. , and ; (14) ; (15) ; (16) ; in , and To evaluate the weights of the three neural networks: the network itself, the model network, and the execution network, ; and The basis functions of the evaluation network, the model network, and the execution network; and Represents approximation error; The following equation is obtained. ; get and How the weights are updated; (17) ; in , , ; get The method of updating weights, (18)。
Citation Information
Patent Citations
Coordination control method for multiple MFD sub-area boundaries based on random distributed control algorithm
CN109559510A
Urban traffic area iterative learning boundary control method considering disturbance
CN113538897A