A mixed traffic flow cooperative control method based on H-infinity direct strategy search
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWEST NORMAL UNIVERSITY
- Filing Date
- 2026-04-29
- Publication Date
- 2026-08-04
AI Technical Summary
但在真实混合交通中,HDV行为高度非线性、非光滑、随机性强,且AV自身存在执行延迟与传感器噪声
[0026] 1. Construct a linear state-space model for mixed traffic, explicitly coupling disturbance inputs and performance outputs. This scheme establishes a linear model with the motion deviations of each vehicle relative to the target steady state as state variables. In this model, the disturbance input w(t) explicitly represents external uncertainties, such as non-smooth disturbances like sudden braking of HDV and lane changes, while the control input u(t) is executed by AV. Simultaneously, the performance output z(t) is defined as a combination of weighted state and control variables, flexibly characterizing the trade-off between smoothness and control energy consumption through positive definite weighting matrices Q and R. Compared to traditional LQR models that implicitly assume Gaussian white noise and only minimize quadratic costs, this model models disturbances as bounded energy signals, better reflecting the sudden, non-Gaussian disturbance characteristics in real traffic. Furthermore, the performance output design supports explicit constraints on traffic wave suppression capabilities (such as speed fluctuation amplitude), avoiding empirical parameter tuning.
Smart Images

Figure CN122511085A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent traffic management and control, and particularly relates to a hybrid traffic flow cooperative control method based on H∞ direct policy search. Background Technology
[0002] Over the past few decades, the ever-increasing demand for travel has continuously exacerbated the operational pressure on transportation systems, and traffic congestion has become a key bottleneck restricting sustainable urban development. The stop-and-go traffic waves caused by frequent acceleration and deceleration of human-driven vehicles (HDVs) are the core mechanism for the formation and propagation of congestion; while human factors such as lane changes and sudden braking further amplify the instability of traffic flow, causing local disturbances to rapidly amplify along the traffic flow and form cascading congestion.
[0003] To address these challenges, traditional traffic control strategies primarily rely on fixed roadside infrastructure, such as traffic lights and variable speed limit signs. While effective in specific scenarios, these methods struggle to achieve dynamic, precise, and proactive traffic flow control. In recent years, the rapid development of autonomous vehicles (AVs) has provided a new path for the paradigm shift in traffic control. AVs possess integrated perception, decision-making, and execution capabilities, and through coordinated formations, such as ACC / CACC, they can significantly improve traffic efficiency and safety. More importantly, in mixed traffic scenarios where AVs and HDVs coexist, a small number of AVs can act as mobile actuators, utilizing their high-precision perception and controllable dynamic characteristics to proactively intervene in disturbance propagation from within the traffic flow—a concept known as Lagrangian control. Cui et al. theoretically demonstrated that a single AV can effectively dissipate traffic oscillations in a ring road; real-world experiments have also verified the feasibility of the "slow in, fast out" strategy (i.e., delayed acceleration, early deceleration) in smoothing traffic fluctuations, and this has been extended to more complex scenarios such as head-to-tail sequence stability analysis.
[0004] However, existing AV collaborative control methods still face three core bottlenecks:
[0005] First, the model is highly dependent and lacks robustness. Mainstream methods are mostly based on linearized dynamic models or the LQR framework, implicitly assuming that the system is continuously differentiable and that the uncertainty structure is known. However, in real mixed traffic, HDV behavior is highly nonlinear, non-smooth, and highly stochastic, and the AV itself has execution delays and sensor noise. Although traditional robust control introduces disturbance compensation terms, it still requires parameterized modeling of uncertainties. When disturbances exhibit strong non-convex, time-varying, or abrupt characteristics, such as cascading braking triggered by accident avoidance, its performance degrades sharply or even becomes unstable.
[0006] Second, strategy optimization lacks a unified theoretical guarantee. Existing collaborative strategies mostly rely on empirical rule design, such as fixed following distance, threshold-triggered control, or local parameter tuning, such as manual tuning of the LQR weight matrix, which is difficult to adapt to changing traffic conditions. Although data-driven methods such as reinforcement learning have shown some adaptability, their training process is computationally expensive, their policy generalization ability is weak, and they lack strict stability and performance boundary guarantees. In safety-critical traffic control, the uninterpretability and potential failure risk of black-box strategies pose significant hidden dangers.
[0007] Third, existing optimization frameworks struggle to directly address non-smooth policy search problems under the H∞ performance metric. H∞ control aims to minimize the worst-case energy gain, naturally aligning with the robustness requirements of traffic systems for disturbance suppression. However, traditional H∞ controller design relies on accurate models and solving the Riccati equations, making it unsuitable for direct policy space optimization. While direct policy search can optimize control laws without model identification, its objective function is typically non-convex and non-smooth, with missing or unstable gradient information, making standard gradient descent algorithms prone to getting trapped in local minima, compromising convergence and optimality.
[0008] Therefore, how to achieve robust suppression of non-smooth, time-varying disturbances in mixed traffic flow and ensure that the control strategy has theoretically provable stability and performance boundaries has become an urgent problem to be solved. Summary of the Invention
[0009] To address the aforementioned shortcomings of existing technologies, the present invention aims to provide a cooperative control method for mixed traffic flow based on H∞ direct policy search, which can achieve robust suppression of non-smooth, time-varying disturbances in mixed traffic flow and ensure that the control strategy has theoretically provable stability and performance boundaries.
[0010] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0011] A hybrid traffic flow cooperative control method based on H∞ direct policy search is applied to a hybrid traffic system on a ring road, the system including at least one AV and at least one HDV; the method includes the following steps:
[0012] S1. Establish a linear state-space model of the hybrid transportation system:
[0013] ;
[0014] Where x(t) is the state vector, representing the motion deviation of each vehicle relative to the target steady state; u(t) is the control input vector of the AV; w(t) is the disturbance input vector; A, B, and H are system matrices;
[0015] S2, Set the target equilibrium speed v* for the mixed traffic system;
[0016] Each HDV is assumed to have the same dynamic characteristics, and its equilibrium spacing s* at v* is determined based on these dynamic characteristics;
[0017] Set the desired vehicle spacing for each AV at v* And each It satisfies the accessibility conditions determined by the geometric constraints of the ring road;
[0018] Set the control law of AV as u(t) = -K x(t), where K is the feedback gain matrix to be optimized;
[0019] Define the performance output vector z(t):
[0020] ;
[0021] Where Q and R are positive definite weighting matrices;
[0022] To minimize the H∞ norm of the closed-loop transfer function from w(t) to z(t) To optimize the objective, we construct a controller optimization problem with K as the optimization variable;
[0023] S3. Using the direct policy search method, solve the controller optimization problem in S2 to obtain the result. Minimize the optimal gain matrix K*;
[0024] S4. Configure K* in each AV; during system operation, each AV acquires or estimates x(t) in real time, and calculates and executes acceleration control commands according to the control law u(t) = -K* x(t).
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] 1. Construct a linear state-space model for mixed traffic, explicitly coupling disturbance inputs and performance outputs. This scheme establishes a linear model with the motion deviations of each vehicle relative to the target steady state as state variables. In this model, the disturbance input w(t) explicitly represents external uncertainties, such as non-smooth disturbances like sudden braking of HDV and lane changes, while the control input u(t) is executed by AV. Simultaneously, the performance output z(t) is defined as a combination of weighted state and control variables, flexibly characterizing the trade-off between smoothness and control energy consumption through positive definite weighting matrices Q and R. Compared to traditional LQR models that implicitly assume Gaussian white noise and only minimize quadratic costs, this model models disturbances as bounded energy signals, better reflecting the sudden, non-Gaussian disturbance characteristics in real traffic. Furthermore, the performance output design supports explicit constraints on traffic wave suppression capabilities (such as speed fluctuation amplitude), avoiding empirical parameter tuning.
[0027] 2. The optimization objective is to minimize the H∞ norm, ensuring disturbance suppression capability in the worst-case scenario. This scheme transforms the controller design into minimizing the H∞ norm of the closed-loop transfer function from the disturbance w(t) to the performance output z(t). This approach aims to minimize the upper bound of the output energy gain under any bounded perturbation. Unlike adaptive control, which relies on uncertainty boundary estimation, or reinforcement learning, which relies on fitting the statistical properties of perturbations with a large number of samples, this method does not presuppose the distribution or structure of the perturbation. It only requires that the energy of the perturbation be finite, which provides a provable worst-case performance guarantee. Even when faced with unmodeled strong nonlinear perturbations, such as a series of sudden braking, the system can still maintain the output fluctuation within a controllable range, significantly improving the upper bound of robustness.
[0028] 3. A direct policy search method is used to solve for the feedback gain matrix K, avoiding the dependence on model accuracy. In step S3, the traditional path of solving the Riccati equation for the H∞ controller is abandoned, and instead a direct policy search is adopted, such as gradient-free optimization, Bayesian optimization, or evolutionary algorithms, to directly search in the policy space for... The minimum feedback gain matrix K*. This process only needs to evaluate the closed-loop system's response to disturbances, without needing to identify system matrices such as A, B, and H. Compared to model-driven methods, such as the LQR / H∞ standard design, which have stringent requirements for model accuracy, this method is naturally immune to dynamic modeling errors; at the same time, it avoids the problems of unstable policy network training, slow convergence, and difficulty in verification in deep reinforcement learning, realizing a model-free but verifiable optimization path, improving engineering feasibility while ensuring theoretical rigor.
[0029] 4. Implement real-time distributed control command generation, supporting multi-AV collaboration and dynamic adaptation. Step S4 clarifies that each AV acquires or estimates x(t) in real time, and executes u(t) = -K*x(t) based on the uniformly optimized K*. That is, each AV independently calculates acceleration commands without the need for central coordination or communication synchronization. This design aligns with the advantages of the distributed architecture of AVs. Unlike collaborative strategies that rely on centralized scheduling or frequent vehicle-to-vehicle communication and are susceptible to communication delays and packet loss, this method achieves collaborative control with low communication overhead and high fault tolerance. Furthermore, since K* has already undergone global robust optimization in the offline stage, it does not require online re-optimization at runtime, ensuring real-time performance and deterministic response, making it particularly suitable for rapid intervention requirements under sudden disturbances.
[0030] In summary, this method can achieve robust suppression of non-smooth, time-varying disturbances in mixed traffic flows and ensure that the control strategy has theoretically provable stability and performance boundaries.
[0031] Preferably, in step S1, based on the optimal speed model, a linearized car-following dynamics model of HDV containing external disturbance terms is established; a linearized dynamics model with acceleration as the control input is established for AV; the vehicle spacing deviation state and speed deviation state of all vehicles are aggregated into a system state vector x(t), the control input of all AVs is aggregated into a control input vector u(t), and the external disturbances experienced by all vehicles are aggregated into a disturbance input vector w(t);
[0032] In step S2, the equilibrium distance s* of HDV at v* is determined according to the linearized car-following dynamics model of HDV.
[0033] This approach addresses several issues: 1) Traditional methods often directly linearize the AV-HDV coupled system as a black box or ignore HDV dynamics, leading to model distortion. Our solution, however, linearizes the system based on the widely validated car-following model OVM and explicitly separates the disturbance term. This ensures the resulting linear model reflects the inherent dynamic characteristics of the HDV, such as response delay and asymmetric acceleration / deceleration, while also meeting the stringent requirements of H∞ optimization on system structure. Compared to purely data-driven modeling, such as neural network fitting, this method requires no large amount of training data, and the model parameters have clear physical meaning, significantly improving the reliability and portability of the controller design.
[0034] 2. By using the HDV linearization model to inversely deduce its equilibrium distance s under v, we can ensure that the desired vehicle distance set by the AV does not violate the HDV's own stable operating conditions. For example, this avoids the risk of rear-end collisions caused by excessively small vehicle distances, thus preventing closed-loop instability caused by unreasonable vehicle distance settings at the source. This differs from the existing strategies that often use fixed vehicle distances or set vehicle distances solely based on the AV's own performance. This gives the cooperative control strategy intrinsic stability assurance and provides an effective feasible region for subsequent H∞ optimization.
[0035] Preferably, in step S1, a linearized car-following dynamics model of the HDV is constructed by applying the car-following function of the optimal velocity model. At the equilibrium point The result is obtained by performing a first-order Taylor expansion at that point.
[0036] In this setup, the OVM, as a classic micro-carriage-following model, can effectively characterize the expected-to-actual speed deviation driving mechanism and response lag characteristics of the HDV. However, its strong nonlinearity makes it difficult to directly apply to optimization frameworks based on linear system theory, such as H∞. This approach, by performing a first-order expansion at a physically clear equilibrium point, preserves the main dynamic characteristics of the HDV in normal traffic flow, such as negative feedback stability and damping effects, while obtaining a linear approximation model that meets the requirements of the LTI system. Compared to global linearization or simplification that ignores HDV dynamics, this method significantly improves the model's fidelity near the operating point and avoids the complexity and parameter inflation caused by higher-order expansions, laying a reliable modeling foundation for constructing a cooperative controller with provable performance.
[0037] Preferably, the linearized car-following dynamics model takes the form of:
[0038] ;
[0039] ;
[0040] in, , Let represent the distance and speed of vehicle i at time t, respectively; Let α1, α2, and α3 represent the distance deviation and speed deviation of vehicle i at time t, respectively; α1, α2, and α3 are constants greater than zero, representing the optimal speed model at the equilibrium point. linearization coefficient at; .
[0041] This setup, by defining the deviation variable The original nonlinear car-following system is transformed into a system of homogeneous linear differential equations with the zero equilibrium point as the origin, allowing the entire hybrid traffic system to naturally merge into the standard state-space framework. This treatment not only eliminates the interference of steady-state offset on controller design and avoids integral saturation or static error, but also ensures the zero-mean characteristic of the performance output z(t) in subsequent H∞ optimization. When the traffic flow stabilizes at the target state, x(t)→0 and z(t)→0, thus making the minimization of the H∞ norm truly correspond to the improvement of disturbance suppression capability, rather than compensating for steady-state deviation.
[0042] Preferably, the optimal velocity model is:
[0043] ;
[0044] Among them, s i (t) represents the distance between vehicle i and the preceding vehicle i-1; v represents the relative speed between vehicle i and the vehicle in front i-1; i (t) represents the speed of vehicle i at time t; A parameter indicating the driver's sensitivity to the speed difference between the current vehicle speed and the desired vehicle speed; A parameter indicating the driver's sensitivity to the relative speed between their vehicle and the vehicle in front; This represents the driver's desired speed function;
[0045] The linearization coefficients α1, α2, α3 are: .
[0046] Traditional modeling often simplifies HDV to a pure integrator or an empirical second-order system, losing sight of the core logic of driver decision-making. This approach, however, utilizes the two-factor structure of OVM (Optical Dynamics Model), with the α term driving speed tracking and the β term suppressing relative motion, accurately characterizing the feedforward + feedback hybrid strategy of human driving. More importantly, the linearization coefficients α1, α2, and α3 are analytically derived from the original mechanistic parameters α and β and the derivative of the desired velocity function. This allows controller design to move beyond black-box identification and proactively configure the model based on driver characteristics, such as a larger β for conservative drivers, significantly improving the interpretability and adjustability of the strategy.
[0047] 2. As shown by α3=β, the speed disturbance of the preceding vehicle is directly injected into the acceleration dynamics of the current vehicle through the coefficient β, clearly revealing the propagation path of the traffic disturbance. α2=α+β indicates that the total damping is composed of both speed tracking stiffness and relative motion damping. Compared to coarse-grained modeling that treats HDV disturbances as white noise or unknown bounded inputs, this scheme achieves physical-level refinement of the disturbance channel, significantly improving the generalization robustness of the robust controller in real mixed traffic scenarios.
[0048] Preferably, the mathematical expression for the driver's desired speed function is:
[0049] ;
[0050] in, Indicates stationary distance and driving distance; 's' represents the maximum distance between vehicles; 's' represents the distance between vehicles.
[0051] This setup ensures that the desired velocity function not only conforms to human factors engineering principles but also provides a solid and reliable underlying support for linear modeling, equilibrium point analysis, and robust controller integration of mixed traffic systems, based on mathematical smoothness and monotonicity.
[0052] Preferably, in step S2, the target equilibrium velocity v* and the HDV equilibrium spacing s* are related through a desired velocity function.
[0053] v∗=V(s*).
[0054] This setup eliminates the dynamic mismatch between AV and HDV at steady-state operating points, ensuring a natural balance of undisturbed coexistence of mixed traffic flows under target conditions.
[0055] Preferably, in step S2, the accessibility condition determined by the geometric constraints of the ring road is:
[0056] ;
[0057] Among them, S AV Let L represent the set of autonomous vehicles, L be the total length of the ring road, n be the total number of vehicles, and k be the number of autonomous vehicles (AVs).
[0058] Preferably, in step S3, the direct policy search reduces the overall state feedback gain matrix K by iteratively adjusting the total state feedback gain matrix K. This continues until the preset convergence condition is met.
[0059] Preferably, in step S4, when acquiring or estimating x(t) in real time, the vehicle spacing and speed information of each vehicle are acquired through the vehicle-mounted sensing device and / or the roadside sensing device. Attached Figure Description
[0060] To make the objectives, technical solutions, and advantages of the invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:
[0061] Figure 1 This is a flowchart of the method;
[0062] Figure 2 This is a diagram showing the equilibrium state of following behavior in Example 1;
[0063] Figure 3 This is a schematic diagram of the response curve of the ring road traffic system in Example 1 to the pulse disturbance;
[0064] Figure 4 This is a schematic diagram illustrating a scenario where an autonomous vehicle in Example 1 increases traffic speed.
[0065] Figure 5 This is a schematic diagram illustrating the stabilization of traffic flow and improvement of traffic speed in Example 2;
[0066] Figure 6 The numerical experimental results show the strong disturbance scenario of the 6th vehicle under different control strategies in Example 2;
[0067] Figure 7These are simulation results for different system scales in Example 2. Detailed Implementation
[0068] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0070] Example 1
[0071] like Figure 1 As shown, this invention provides a hybrid traffic flow cooperative control method based on H∞ direct policy search, applied to a hybrid traffic system on a ring road, the system including at least one AV and at least one HDV; the method includes the following steps:
[0072] S1. Establish a linear state-space model of the hybrid transportation system:
[0073] ;
[0074] Where x(t) is the state vector, representing the motion deviation of each vehicle relative to the target steady state; u(t) is the control input vector of the AV; w(t) is the disturbance input vector; and A, B, and H are system matrices.
[0075] In practice, based on the optimal speed model, a linearized car-following dynamics model of HDV including external disturbance terms is established; a linearized dynamics model with acceleration as the control input is established for AV; the vehicle spacing deviation state and speed deviation state of all vehicles are gathered into a system state vector x(t), the control input of all AVs is gathered into a control input vector u(t), and the external disturbances experienced by all vehicles are gathered into a disturbance input vector w(t).
[0076] Construct a linearized car-following dynamics model for HDV by applying the car-following function of the optimal velocity model. At the equilibrium point The result is obtained by performing a first-order Taylor expansion at that point.
[0077] The linearized car-following dynamics model takes the following form:
[0078] ;
[0079] ;
[0080] in, , Let represent the distance and speed of vehicle i at time t, respectively; Let α1, α2, and α3 represent the distance deviation and speed deviation of vehicle i at time t, respectively; α1, α2, and α3 are constants greater than zero, representing the optimal speed model at the equilibrium point. linearization coefficient at; .
[0081] The optimal velocity model is:
[0082] ;
[0083] Among them, s i (t) represents the distance between vehicle i and the preceding vehicle i-1; v represents the relative speed between vehicle i and the vehicle in front i-1; i (t) represents the speed of vehicle i at time t; A parameter indicating the driver's sensitivity to the speed difference between the current vehicle speed and the desired vehicle speed; A parameter indicating the driver's sensitivity to the relative speed between their vehicle and the vehicle in front; This represents the driver's desired speed function;
[0084] The linearization coefficients α1, α2, α3 are: .
[0085] The mathematical expression for the driver's desired speed function is:
[0086] ;
[0087] in, Indicates stationary distance and driving distance; 's' represents the maximum distance between vehicles; 's' represents the distance between vehicles.
[0088] S2. Set the target equilibrium speed v* for the mixed traffic system; assume that each HDV has the same dynamic characteristics, and determine its equilibrium spacing s* at v* based on these dynamic characteristics; set the desired vehicle spacing at v* for each AV. And each The accessibility conditions determined by the geometric constraints of the ring road must be satisfied:
[0089] ;
[0090] Among them, S AV Let L represent the set of autonomous vehicles, L be the total length of the ring road, n be the total number of vehicles, and k be the number of autonomous vehicles (AVs).
[0091] Set the control law of AV as u(t) = -K x(t), where K is the feedback gain matrix to be optimized; define the performance output vector z(t): Where Q and R are positive definite weighting matrices;
[0092] To minimize the H∞ norm of the closed-loop transfer function from w(t) to z(t) To optimize the objective, we construct a controller optimization problem with K as the optimization variable.
[0093] In practice, ;
[0094] in, Represents the maximum singular value; It is a transfer function In the frequency domain, j is the imaginary unit and ω is the angular frequency. This is the upper bound.
[0095] In practice, based on the linearized car-following dynamics model of HDV, the equilibrium distance s* of HDV at v* is determined.
[0096] The target balancing velocity v* is related to the HDV balancing distance s* through the desired velocity function, v∗=V(s*).
[0097] To facilitate a better understanding of the technical content of S1 and S2 by those skilled in the art, the following explanation is provided.
[0098] In existing research, the Optimal Speed Model (OVM) and the Intelligent Driving Model (IDM) are two typical models describing the car-following behavior of manually driven vehicles. Both OVM and IDM can be represented in the following general form:
[0099] (1)
[0100] in, This represents the acceleration of vehicle i at time t. This represents the relative speed between vehicle i and the vehicle in front of it (i-1). The following function model represents the acceleration of vehicle i, which is determined by the following distance, relative speed, and its own speed. Equation (1) physically means that the driver typically considers the state of the vehicle in front and their own vehicle when making a following decision. This method considers the heterogeneous HDV following behavior under normal circumstances, i.e., the acceleration functions of different HDVs. The expressions may differ, and there may be differences between different driver behaviors.
[0101] When the traffic flow is in equilibrium, each vehicle travels at the same equilibrium speed. Exercise, that is , At the same time, each vehicle has a corresponding balanced following distance. According to formula (1), each HDV should satisfy the following in a balanced state: (2)
[0102] As can be seen from formula (2), each vehicle may have different equilibrium speeds at different equilibrium speeds. Figure 2 Showcasing the possibilities of different cars and Relationship: Generally speaking, along with It grows with the increase; for the same Balanced vehicle spacing The required distance may vary depending on the driver. It is important to emphasize that the balance distance for an HDV is determined by the driver and is often a positional parameter in practical applications. In contrast, the balance distance for an AV at balanced speed can generally be designed and specified by the controller.
[0103] Assuming each car is in equilibrium There is a small disturbance nearby. The error state between the actual state and the equilibrium state of vehicle i is defined as:
[0104] (3)
[0105] in, Let $\mathbf$ and $\mathbf$ represent the distance deviation and speed deviation of vehicle $i$ at time $t, respectively. Based on the equilibrium state equation (2), a first-order Taylor expansion is applied to the HDV following dynamics equation (1) to obtain the linearized car-following model for each HDV:
[0106] (4)
[0107] Among them, the coefficients of each term have a carousel function. In equilibrium state Taking the partial derivative, we get These three coefficients reflect the driver's sensitivity to distance error, their own speed error, and the speed of the vehicle in front, respectively. To reflect real driving behavior, it is generally assumed that acceleration should increase when the distance to the vehicle increases, the driver's own speed decreases, or the speed of the vehicle in front increases. Therefore, it is assumed that... ,and .
[0108] In the subsequent analysis, the Optimal Velocity Model (OVM) is used to derive the specific forms of formulas (2) and (4). The OVM model is defined as:
[0109] (5)
[0110] in, A parameter representing the driver's sensitivity to the speed difference between the current vehicle speed and the desired vehicle speed. This parameter represents the driver's sensitivity to the relative speed between their vehicle and the vehicle in front. This represents the driver's desired speed function, which is related to the current distance to the vehicle and is usually expressed in a continuous piecewise form:
[0111] (6)
[0112] Among them, parameters Indicates stationary distance and driving distance. This indicates the maximum distance between vehicles.
[0113] For the general form of the OVM model (5), the equilibrium state can be directly obtained according to formula (2):
[0114] (7)
[0115] Furthermore, the parameter values in the linearized model (4) can be calculated as follows:
[0116] (8)
[0117] in Represents the desired velocity function In balance spacing The derivative of with respect to s.
[0118] For the autonomous vehicle numbered i = 1, its acceleration signal is directly used as the control input. Therefore, its longitudinal dynamic model can be expressed as:
[0119] (9)
[0120] in and For autonomous vehicles at speed The target vehicle distance is adjustable. It should be noted that... yes The design parameters of autonomous vehicles, and This represents the desired traffic flow speed, which is the speed that the autonomous vehicles need to stabilize the hybrid traffic system at. To obtain the lumped state-space equations of the hybrid traffic system, the deviation states of all vehicles are reduced to a single lumped state vector:
[0121] (10)
[0122] From the perspective of the global system, the following standard linear dynamic equations are obtained: (11)
[0123] The coefficient matrix is expressed as follows:
[0124] (12)
[0125] The expression for each sub-block is:
[0126] (13)
[0127] In equation (11), the state evolution of each vehicle is determined solely by its own state and the state of the vehicle directly preceding it. A theoretical analysis of the potential of autonomous vehicles in smoothing mixed traffic flows will be conducted, and control inputs for autonomous vehicles will be designed. .
[0128] Multiple real-world road experiments have shown that in a ring road system composed of HDVs (High-Distance Vehicles), traffic is highly prone to generating and causing severe congestion. Theoretical studies have also confirmed that a linearized traffic system will remain stable when the following conditions are met:
[0129] (14)
[0130] Otherwise, the traffic system may become unstable, and even small disturbances can trigger stop-and-go waves. Condition (14) was initially derived from frequency analysis in the literature, which is another proof method based on eigenvalue analysis. For the OVM model (5), condition (14) can be simplified to: (15)
[0131] This indicates that, to ensure traffic stability, human drivers must be more sensitive to speed deviations than to the optimal speed function; otherwise, stop-and-go traffic may occur.
[0132] Rather than attempting to alter human driving behavior, this approach aims to stabilize hybrid transportation systems by controlling individual autonomous vehicles. Specifically, this method discusses two fundamental system theory concepts of hybrid transportation systems (11): controllability and stability.
[0133] The controllability of a dynamic system reflects its ability to guide the system's behavior to a desired state through appropriate control inputs. If all uncontrollable modes of a system (11) are stable, then the system is stabilizable. Essentially, controllability states that a controlled system can be adjusted from any initial position to any final position, while weakened stabilization states that a controlled system can be stabilized from any initial position to an equilibrium position (i.e., the origin). Therefore, three practical lemmas are given first.
[0134] Lemma 1 (Kalman controllability criterion): Linear systems The necessary and sufficient condition for controllability in (11) is .
[0135] Lemma 1 is the well-known Kalman controllability criterion, which provides a necessary and sufficient mathematical condition for controllability. However, directly calculating the rank of a matrix requires knowing... The specific numerical expression of the system is not readily available, and the calculation of the rank of a large-scale system is numerically unreliable. Therefore, by applying a specific linear transformation and re-expressing the linear system using different basis vectors, system decoupling can be achieved. Specifically, given a non-singular matrix T, a new state is defined. It satisfies:
[0136] (16)
[0137] Then an equivalent linear system can be obtained. For the system before and after the linear transformation, the controllability invariance described in the following lemma holds.
[0138] Lemma 2 (Invariance of Linear Transformation): Linear Systems Controllable if and only if It is controllable for any non-singular matrix T.
[0139] If it can be passed For system matrix Diagonalization is then performed. The controllability will be easier to derive. For the system matrix in the original ring-shaped mixed traffic system model (11) Direct diagonalization is very difficult; however, specific state feedback can be applied to simplify the system dynamics. Consider a feedback control law. You can then obtain:
[0140] (17)
[0141] Lemma 3 (Linear Feedback Invariance): Linear Systems Controllable if and only if compatible with all dimensions ,system It is also controllable.
[0142] Now, the following are the results showing the controllability of the mixed transportation system (11).
[0143] Theorem 1: Consider a mixed traffic system on a circular road, consisting of one AV and n-1 HDVs, given by (11). The following conclusions are given.
[0144] 1) The system is not fully controllable.
[0145] 2) There exists an uncontrollable mode corresponding to zero eigenvalue, and this uncontrollable mode is stable.
[0146] The main idea is to utilize the invariance of controllability under linear transformations and state feedback. Through a series of state feedbacks and linear transformations, the system is diagonalized, thus obtaining analytical conclusions regarding the controllability of equation (11). The main idea is to first utilize state feedback to... Transform into a system in block cyclic matrix form Then, through linear transformation, Diagonal decoupling The specific steps are as follows:
[0147]
[0148] First, introduce virtual input. system Transform into The virtual input is defined as follows:
[0149] (18)
[0150] Its physical meaning is the difference between the actual control value of AV and the acceleration value of HDV. Therefore, the transformed system state-space equation is:
[0151] (19)
[0152] The coefficient matrix is expressed as follows:
[0153] (20)
[0154] Essentially, formula The text describes a circular traffic system where all vehicles are HDVs (High-Density Vehicles), with each HDV having a virtual control input. .Depend on The controlled AV is controlled by a vehicle with virtual input. It was replaced by HDV.
[0155] According to Lemma 3, the system and the original system The controllability between them remains unchanged. It is important to note that... It is a block cyclic matrix, which can be obtained through Fourier matrix. Achieve diagonalization.
[0156] Specifically, a transformation matrix can be used. Will Transformation For, here express The conjugate transpose of gives the new system matrix:
[0157] (twenty one)
[0158] Where represents the Kroll product, and the expression for each parameter matrix is:
[0159] (twenty two)
[0160] and Indicates a It is a block diagonal matrix with diagonal elements. Because... It is a unitary matrix, representing the state variables of the new system after the transformation. The expression is:
[0161] (twenty three)
[0162] Control coefficient matrix of the new system for:
[0163] (twenty four)
[0164] therefore, The dynamic equation is:
[0165] (25)
[0166] but It can be decomposed into n independent subsystems:
[0167] (26)
[0168] It is easy to see from this This means It is an uncontrollable mode, but it remains unchanged during dynamic evolution. According to... Therefore, we can conclude that:
[0169] (27)
[0170] Due to the system The linear transformation is equivalent to the system .at the same time , , They all have the same controllability, therefore, the original system It is not fully controllable, and has at least one uncontrollable mode (27) that remains constant throughout the system's dynamic evolution and corresponds to a zero eigenvalue. Since the zero eigenvalue only appears in sub-blocks of the system after diagonal decoupling... In the middle, so it is The algebraic multiplicity of the expression is 1. Therefore, this uncontrollable mode is itself stable.
[0171] In summary, uncontrollable modes are identified. The existence of , as can be seen from equation (27), means that due to the ring road structure of the mixed traffic system, the sum of the distances between each vehicle should remain constant during the dynamic evolution of the system. Some literature argues that the ring mixed traffic system is completely controllable, while Theorem 1 utilizes the block loop characteristics of the mixed traffic system to clearly point out the existence of uncontrollable modes, thus demonstrating that the system is not completely controllable.
[0172] Having revealed the uncontrollable components (27), the next step is to demonstrate that the mixed traffic system is stabilizable. For this, the PBH test is used for stabilization analysis.
[0173] Lemma 4 (PBH controllability criterion): Linear systems It is controllable if and only if it is Each eigenvalue They all .also, If it exists in Make:
[0174] (28)
[0175] in, yes correspond The left eigenvector, then This corresponds to an uncontrollable mode.
[0176] Theorem 2 (Stabilizability of a Mixed Traffic System with One AV): Consider a mixed traffic system on a ring road with one AV and n-1 HDVs, as given by Equation (11). The system is stabilized, and the controllability matrix is given. The rank satisfies:
[0177] (29)
[0178] As explained in the proof of Theorem 1, the diagonalized system given by Equation (25) has the same controllability characteristics as the original system. Based on this, the proof of this theorem directly examines the stabilization. The idea is to find all uncontrollable modes and prove that they are all stable, which shows that the system is stabilizable. The following discussion will cover two cases.
[0179] Scenario 1: .because It is a block diagonal matrix because it satisfies:
[0180] (30)
[0181] in, express The eigenvalues. Substituting equation (22) into equation (30) (i=1,2,…,n), we get:
[0182] (31)
[0183] Since equation (31) is a second-order complex equation, it is difficult to directly obtain its analytical solution. Therefore, this equation will be used to analyze the properties of the eigenvalues, which will be carried out in two steps.
[0184] Step 1: Prove by contradiction and They have no common eigenvalues. Assume there exists λ such that... , Then we have:
[0185] (32)
[0186] because ,available and Therefore:
[0187] (33)
[0188] This is This contradicts the assumption. Therefore and They do not share common eigenvalues.
[0189] Step 2: Prove that all system modes corresponding to non-zero eigenvalues are controllable. Let... for Let the non-zero eigenvalues be denoted as . Let it be its corresponding left eigenvalue. According to Lemma 4, it is necessary to prove... Only then can it be said that the corresponding system mode is controllable.
[0190] remember ,in, .Depend on We can obtain:
[0191] (34)
[0192] because no eigenvalues Therefore, .therefore, Assuming, Substituting (22) into (34) yields:
[0193] (35)
[0194] The unique solution to equation (35) is This indicates that the left eigenvector This contradicts the condition that the eigenvector is not zero. Therefore, The assumption is false, therefore we get This means The corresponding modes are controllable, thus indicating that the system modes corresponding to non-zero eigenvalues are all controllable. Furthermore, due to the unique zero eigenvalue... Appeared Furthermore, the corresponding mode is uncontrollable, leading to the conclusion that: if ,system There is A controllable mode, namely .
[0195] Scenario 2: Substituting this condition into equation (31) yields:
[0196] (36)
[0197] Therefore, we obtain The eigenvalues are as follows:
[0198] (37)
[0199] The following proof consists of two steps.
[0200] Step 1: Proof This eigenvalue corresponds to An uncontrollable mode. Clearly, It is each block The eigenvalues of , that is, The algebraic multiplicity is Consider its left eigenvector. Similar to equation (34), we can obtain:
[0201] (38)
[0202] Expanding the equation, we get:
[0203] (39)
[0204] Therefore, we obtain Therefore, one can choose... corresponding There are n linearly independent left eigenvectors, namely:
[0205] (40)
[0206] Based on these left eigenvectors, it is easy to know and This means for This eigenvalue exists. An uncontrollable mode.
[0207] Step 2: Consider the remaining eigenvalues, i.e.:
[0208] (41)
[0209] Among them, zero eigenvalue This still corresponds to an uncontrollable mode, as shown in (27). Next, we will prove that... , The corresponding mode is controllable. For , Let its left eigenvector be . ,in, .because Not other sub-blocks , eigenvalues, thus having .for ,have This holds true, similar to the argument in equation (35). Therefore, This means , All corresponding modes are controllable.
[0210] In summary, eigenvalues Corresponding to one controllable mode and An uncontrollable mode, due to These uncontrollable modes are all stable; eigenvalues Corresponding Each mode is controllable, and the zero eigenvalue corresponds to an uncontrollable mode, but that mode is stable. Therefore, the system... There are a total of n controllable modes. Because... and Controllable suppression, therefore ultimately leading to the system It can calm you down.
[0211] Theorem 2 shows that a single AV in a ring-shaped mixed traffic system (11) always has one uncontrollable mode, which corresponds to a zero eigenvalue, while the other modes are either controllable or stable.
[0212] To characterize the impact of external disturbances in traffic flow, it is assumed that each vehicle's acceleration signal contains an external disturbance term. This perturbation can be used to describe factors such as driver behavior uncertainty and sudden braking. Under this assumption, the linearized velocity dynamics of the i-th human-driven vehicle (HDV) can be written as:
[0213] (42)
[0214] in, and These represent the deviations of the vehicle spacing and speed from the equilibrium state, respectively. This describes the driver's sensitivity to changes in distance and speed. Therefore, the linearized dynamic equation of the HDV in (11) becomes:
[0215] (43)
[0216] in , This indicates that the disturbance enters the system through the acceleration channel. Stacking the error states of all vehicles yields the overall state vector of the mixed traffic system.
[0217] (44)
[0218] Combined with the control input of AV vehicles The mixed traffic system after adding external disturbance signals can be written as:
[0219] (45)
[0220] Among them, the system matrix With input matrix Consistent with the previous text, the matrix consists of the disturbance channels of each vehicle. Composition. Formula (45) reflects the propagation mechanism of disturbances in the vehicle queue.
[0221] To comprehensively consider vehicle spacing, speed stability, and control energy consumption, the performance output is defined as:
[0222] (46)
[0223] This performance output takes into account vehicle spacing deviation, speed deviation, and AV control input energy. It has positive weights. Equivalently, the above performance state can also be expressed in a weighted form:
[0224] (47)
[0225] in, and These represent the square roots of the state and control performance weights, respectively.
[0226] Next, assume that the autonomous vehicle uses a linear state feedback control law.
[0227] (48)
[0228] in Let be the feedback gain matrix to be designed. Substituting it into system (45), the closed-loop system can be expressed as:
[0229] (49)
[0230] Compared with the literature Different performance metrics exist; this method focuses on the robust performance of the system under worst-case perturbation conditions. Therefore, it adopts a perturbation-based approach. To performance output The closed-loop transfer function Norm as a performance metric, i.e.:
[0231] (50)
[0232] in, This represents the maximum singular value. This indicator characterizes the maximum amplification of disturbances across all frequency ranges, directly reflecting the sensitivity of the traffic system to disturbances under worst-case conditions.
[0233] Therefore, the robust control design problem for autonomous vehicles can be formulated as the following optimization problem:
[0234] (51)
[0235] because Norm design involves upper bound operations and maximum singular value calculations in the frequency domain. This optimization problem is generally non-convex and non-smooth, making it difficult to solve using traditional LMI or semi-normal programming methods.
[0236] To address the aforementioned issues, a Direct Search Strategy (DPS) method is employed to adjust the feedback gain. Optimization is performed. In each iteration, the singular value vector information of the closed-loop transfer function at the peak frequency can be used to construct... The Clarke subgradient of the performance function is used to update the control gain. Related theoretical results show that for a non-degenerate dynamic controller, all Clarke time invariant points correspond to the global optimum. Controller.
[0237] For autonomous vehicles The DPS control strategy utilizes system-level objectives. Instead of passively responding to traffic disturbances, it considers the global behavior of the entire mixed traffic flow, thereby proactively mitigating adverse disturbances such as... Figure 3 As shown. Compared to existing strategies, The DPS controller performs better in smoothing traffic flow and improving fuel economy.
[0238] Vehicle 2 is subjected to an initial disturbance, and the vehicle parameters are chosen to simulate stop-and-go behavior in a real-world experiment. (a) All vehicles are manually driven, in which case the disturbance is amplified, resulting in stop-and-go waves. (b) Vehicle 1 is equipped with CACC (Continuous Acceleration and Adaptive Cruise Control), and its behavior is passively adjusted based on the vehicle in front. In this case, the disturbance is not amplified, but small traffic waves still persist for a longer period. (c) Vehicle 1 uses parameters that consider the global behavior of the entire mixed traffic flow. The DPS control strategy actively mitigates adverse disturbances. In this scenario, the disturbance is attenuated, and the traffic flow quickly stabilizes. Considering the presence of disturbances, the controller design for the mixed traffic system is transformed into a... Robust performance optimization issues, and through DPS obtains a stable feedback gain value by directly searching in the policy space. This enables AV to actively suppress disturbance propagation. The reachability analysis of the target equilibrium is based on a reasoning framework of system invariant modes and steady-state equilibrium constraints.
[0239] Consider the definition of the state vector and denote the feedback gain. Then the control law It can be expanded as follows:
[0240] (52)
[0241] in, For the equilibrium traffic state of HDV, the equilibrium equation (2) is satisfied, and It is the autonomous vehicle at the target speed The corresponding expected vehicle spacing.
[0242] Because mixed transportation systems have uncontrollable modes, if Even if a stable controller is found, an improperly selected controller may prevent the system from converging precisely to the desired speed. To reveal this reachability constraint, we further analyze that the closed-loop system will eventually converge to a steady-state form, and based on this, we give the conditions that make the system exactly reach the desired equilibrium state. Consider an autonomous vehicle and... A mixed traffic system consisting of manually driven vehicles on a ring road. If a stable state feedback controller exists, the system can converge to the desired traffic speed. If and only if the desired distance of the autonomous vehicle satisfies
[0243] (53)
[0244] in The total length of the road. The equilibrium spacing for HDV is given. This result indicates that in mixed traffic systems, the accessibility of the desired traffic speed depends not only on the stability of the controller but also on the geometric constraints of road length and the number of vehicles.
[0245] In real-world traffic scenarios, the distance between autonomous vehicles must always be positive, that is... Therefore, we can conclude that:
[0246] (54)
[0247] Based on the equilibrium relationship of HDV (2) and the characteristic that speed increases with increasing spacing in actual driving behavior, the above equation further defines the upper bound of achievable traffic speed. Therefore, the following conclusions can be drawn.
[0248] Traffic speeds in a mixed transportation system (11) have an reachable range.
[0249] (55)
[0250] in satisfy
[0251] (56)
[0252] This result reveals that a single autonomous vehicle can precisely guide a mixed traffic system to a target speed. However, this speed must fall within the range Inside. It is worth noting that, Typically, this is higher than the equilibrium speed of a system consisting solely of HDVs. Physically, this means that autonomous vehicles can follow the vehicle in front with a smaller car-following distance, thus freeing up more space for HDVs behind. In most car-following models, the larger available distance will induce HDVs to travel at a higher speed in equilibrium, thereby improving the overall traffic speed level, such as... Figure 4 As shown, (a) when all vehicles are driven by humans, the distance between the two vehicles is equal under uniform driving dynamics. (b) In a mixed traffic system, autonomous vehicles can be controlled to follow the vehicle in front with a shorter distance, while other human-driven vehicles maintain a larger distance under equilibrium conditions.
[0253] When a mixed transportation system contains only one autonomous vehicle, the system is always stable. However, in reality, multiple autonomous vehicles and a large number of manually driven vehicles often coexist. First, a dynamic model is established for the case of multiple autonomous vehicles, and the controllability and stability of the system are analyzed.
[0254] Consider a mixed traffic system on a ring road, which includes a total of Vehicles, including The vehicles are autonomous vehicles, and Let the numbers of the autonomous vehicles be respectively... And define the set of autonomous vehicles as
[0255] (57)
[0256] For the number is For manually driven vehicles, the error state is still defined as ,in To satisfy the equilibrium condition (2) for the HDV equilibrium state, the corresponding linearized car-following model remains unchanged and is still given by equation (4).
[0257] For the number is In autonomous vehicles, the longitudinal acceleration signal is directly used as the control input. Its linearized vehicle dynamics model can be expressed as:
[0258] (58)
[0259] in, , Indicates the target speed The next The desired spacing between vehicles is set for a group of autonomous vehicles. Similar to the case of a single autonomous vehicle, this desired spacing is an adjustable design parameter.
[0260] To obtain a global dynamic description of the system, the error states of all vehicles are summarized into a system state vector. And aggregate all control inputs from autonomous vehicles into
[0261] (59)
[0262] Therefore, a hybrid transportation system containing multiple autonomous vehicles can be represented as the following state-space model:
[0263] (60)
[0264] The coefficient matrix is expressed as follows: (61)
[0265] System Matrix It still has a block-loop structure, with sub-blocks corresponding to different vehicles distinguished based on whether they are autonomous vehicles, and the input matrix... It describes the role and position of each autonomous vehicle in the system.
[0266] This model is structurally a generalization of the single-vehicle autonomous driving scenario. A key question arises when multiple autonomous vehicles are introduced into the system: will the controllability and stability of the mixed traffic system change? The conclusions are summarized below.
[0267] Considering a hybrid transportation system with multiple autonomous vehicles described by equation (60), we have:
[0268] 1. The system (60) is not completely controllable, and there is still an uncontrollable mode;
[0269] 2. System (60) is stable.
[0270] To analyze the controllability of the system, a virtual control input is first introduced. for
[0271] (62)
[0272] in Under this transformation, system (58) can be written as: (63)
[0273] Where the matrix The definition is consistent with the case of a single autonomous vehicle in formula (20). Utilizing As a coordinate transformation matrix, the system can be transformed. Convert to (64)
[0274] in It is a block diagonal matrix, and its structure is consistent with that of formula (21) for a single autonomous vehicle. It is composed of the superimposed input channels corresponding to each autonomous vehicle. In this coordinate system, the system can be decoupled into independent subsystems. It can be directly verified that there always exists one that satisfies… An uncontrollable mode corresponds to a zero eigenvalue. Since the algebraic multiplicity of this zero eigenvalue is 1, the uncontrollable mode is stable.
[0275] In a mixed traffic system with multiple AVs, the expression for the uncontrollable mode is:
[0276] (65)
[0277] It remains unchanged during the evolution of system dynamics. The physical meaning of this uncontrollable mode is consistent with the case when a single CAV exists, that is, due to the aggregate properties of the ring road, the sum of the distances between each vehicle should remain unchanged.
[0278] The preceding text has been based on The DPS framework constructs a control strategy for a single autonomous vehicle. Since the control method is optimized directly in the feedback control strategy space and does not depend on the specific value of the control input dimension, it can be naturally extended to hybrid transportation systems with multiple autonomous vehicles (60).
[0279] In the case of multiple autonomous vehicles, the system state feedback control rate is defined as follows: The feedback gain matrix It is composed of the superposition of the sub-feedback gains corresponding to each autonomous vehicle, i.e.
[0280] (66)
[0281] in Indicates the number is The feedback gain used in autonomous vehicles corresponds to the control input as follows: .exist Under the DPS framework, each feedback gain The design objective, derived through a direct search strategy, is to minimize the impact of disturbance inputs on the performance output of the hybrid traffic system. Norm. The optimal feedback gain varies among different autonomous vehicles due to their different spatial positions and operational channels within the transportation system. They are not exactly the same.
[0282] In obtaining Then, consider the expression described by equation (60) containing A ring road mixed traffic system with autonomous vehicles. Assume that... If the DPS method obtains the state feedback gain (66) and the coefficient matrix of the linear equation system (69) is non-singular, then the traffic system can be stabilized to the target speed. If and only if each autonomous vehicle expects a certain distance satisfy:
[0283] (67)
[0284] in The equilibrium spacing of HDV is given by equilibrium condition (2).
[0285] Assume the system is under steady control at time and . The system converges to an equilibrium state. Similar to the previous analysis, the final state of the system can be expressed as:
[0286] (68)
[0287] in and These are undetermined constants. Since the closed-loop system satisfies the zero acceleration condition at the equilibrium point, i.e. Combining this with the inherent uncontrollable modal constraints in the system, the following system of linear equations can be obtained:
[0288] (69)
[0289] in
[0290] (70)
[0291] The system reaches the desired equilibrium speed Equivalence and If and only if a constant vector When the above system of equations (69) is zero, the reachability condition (71) is obtained.
[0292] (71)
[0293] Because the desired spacing of each autonomous vehicle must meet Therefore, there is
[0294] (72)
[0295] Therefore, it can be concluded that the maximum achievable vehicle spacing of HDV under equilibrium conditions satisfies...
[0296] (73)
[0297] This determined that the mixed transportation system contains Maximum achievable equilibrium speed in the case of an autonomous vehicle .
[0298] Including In a hybrid transportation system (60) with autonomous vehicles, the achievable speed range of traffic flow is: ,in It is determined by equation (74).
[0299] (74)
[0300] This conclusion, when applied to a single autonomous vehicle, clearly demonstrates that as the proportion of autonomous vehicles increases, the upper bound of the achievable equilibrium speed for a hybrid transportation system will significantly improve.
[0301] S3. Using the direct policy search method, solve the controller optimization problem in S2 to obtain the result. Minimize the optimal gain matrix K*.
[0302] The direct policy search reduces the overall state feedback gain matrix K by iteratively adjusting it. This continues until the preset convergence condition is met.
[0303] S4. Configure K* in each AV; during system operation, each AV acquires or estimates x(t) in real time, and calculates and executes acceleration control commands according to the control law u(t) = -K* x(t).
[0304] In practice, when acquiring or estimating x(t) in real time, the vehicle spacing and speed information of each vehicle are obtained through vehicle-mounted sensing devices and / or roadside sensing devices.
[0305] Example 2
[0306] To better illustrate the effectiveness of this method, the following simulation experiment is conducted.
[0307] To evaluate its effectiveness under real-world conditions, simulation experiments were conducted to assess the nonlinearity of the car-following behavior. Multiple sets of simulation experiments were performed to verify the effectiveness based on the nonlinear optimal velocity model (OVM). All experiments were completed using MATLAB.
[0308] The parameters set for numerical simulation in the OVM model (5) are as follows: .
[0309] Verification shows that this set of parameters violates the stability condition, which means that the system is unstable when only manually driven vehicles are present in the circular road, and any disturbance may produce stop-and-go wave phenomena.
[0310] For the weighting coefficients in the performance output function (47), select This parameter achieves a reasonable trade-off between vehicle spacing adjustment, speed stability, and control input energy, and remains consistent across all numerical experiments to ensure fair comparisons between different control strategies. Based on DPS control framework, feedback gain for autonomous vehicles Obtained through a direct search method, i.e., by directly minimizing the closed-loop system. Performance metrics are used to determine control parameters. The controller design process does not rely on any convex relaxation or semidefinite programming solvers, but rather updates the control gain iteratively in the policy space through non-smooth optimization until convergence. To avoid vehicle collisions and ensure the physical feasibility of the system, it is assumed that all vehicles are equipped with a standard automatic emergency braking system, and its control logic is described as follows:
[0311]
[0312] The maximum deceleration is set as follows: .
[0313] Verification through numerical simulation The stability capability and traffic efficiency improvement effect of the DPS control strategy in a nonlinear mixed traffic system are investigated. Consider a circular road scenario with a total of n = 20 vehicles and a road length of L = 400. In the initial trial state, each vehicle is subjected to a small perturbation near its equilibrium position and velocity. The initial position of the i-th vehicle is... The initial test speed was ,in To correspond with and balance the vehicle distance Equilibrium velocity, disturbance term , , express arrive A uniform distribution on the surface.
[0314] When all vehicles are manually driven, such as Figure 5 As shown in (a), the gray curve represents the speed changes of 20 human-driven vehicles, and the black curve represents the average speed. The simulation results clearly show that the traffic disturbance gradually amplifies. The initial small speed propagates along the convoy and continuously increases, with the speeds of individual vehicles exhibiting continuous oscillations, eventually forming a stop-and-go wave. This phenomenon indicates that mixed traffic is in an unstable state when only human-driven vehicles are present. Conversely, as... Figure 5 (b) shows the introduction of a vehicle using... The autonomous vehicle using the DPS control strategy experiences slight disturbances in the initial stage. However, the autonomous vehicle actively applies adjustment mechanisms during the early stages of disturbance propagation, effectively suppressing the amplification of disturbances in the worst-case scenario. This allows the traffic flow to stabilize to its original average speed of 15 m / s within a short period. This verifies that the proposed control strategy can stabilize the entire nonlinear traffic flow and prevent traffic wave propagation using only a single AV vehicle. Furthermore, based on the conclusions of the accessibility analysis, adjusting the equilibrium speed... Adjust the corresponding balance spacing and Autonomous vehicles can not only stabilize traffic flow but also guide it to higher speeds. The increase from 15 m / s to 16 m / s indicates that controlled autonomous vehicles can not only suppress traffic waves and maintain stability but also improve the overall efficiency of traffic flow through control strategies. (See...) Figure 5 (c) Figure 5 In the OVM model, (a) when the parameters are set to a certain value, the traffic system consisting only of manually driven vehicles is unstable; (b) after introducing an autonomous vehicle with an appropriate control strategy, the mixed traffic system becomes stable; and (c) by controlling the autonomous vehicle, traffic flow can be guided to a higher stable speed. In this case, it can be observed that when only 5% of the vehicles in the mixed traffic system are autonomous vehicles (1 AV / 20 vehicles), the overall traffic speed increases by approximately 6%.
[0315] A scenario closer to real-world traffic conditions is presented, where sudden disturbances occur in the traffic flow caused by infrastructure bottlenecks, merging from nearby ramps, or frequent lane changes. A vehicle on a circular road undergoes a rapid deceleration within a short period. Initially, all vehicles are traveling steadily at 15 m / s. At t=20s, the i-th vehicle rapidly decreases its speed from 15 m / s to 5 m / s within 2 seconds, introducing a representative strong disturbance source into the traffic flow. This setting corresponds to common real-world scenarios such as sudden braking by a vehicle ahead, speed limits due to road construction, or sudden driver error.
[0316] Figure 6 (a) illustrates the evolution when all vehicles are manually driven. The initial disturbance gradually amplifies upstream in the traffic flow, eventually forming distinct stop-and-go waves, causing the overall traffic system to enter an unstable operating state. This phenomenon is consistent with accident-free congestion commonly seen on actual highways, where even without continuous external interference, a single forced movement can induce traffic oscillations that propagate over long periods and distances. The curves on the left side of each subplot reflect the change in vehicle speed over time, with red lines corresponding to disturbed vehicles and black lines representing the system's average speed. The spatiotemporal diagram on the right further visually depicts the backward propagation of traffic waves.
[0317] After introducing an autonomous vehicle using a PI with Saturation control strategy into the system, Figure 6(b) The disturbance was effectively weakened after approximately t=60s, indicating that the AV based on feedback control has a certain ability to suppress fluctuations. However, this method typically achieves safety and stability by maintaining a large vehicle spacing, which may introduce new problems in real multi-lane traffic environments, such as inducing frequent lane changes by vehicles in adjacent lanes, thereby introducing new disturbance sources over a larger area and weakening overall traffic efficiency. The results of using the FollowerStopper control strategy are as follows: Figure 6 As shown in (c), compared to PI control, this method improves both the disturbance dissipation speed and the smoothness of average speed recovery, while maintaining a relatively moderate vehicle spacing. This is consistent with observations from existing real-vehicle experiments, namely that strategies based on the "slow in, fast out" principle can, to some extent, balance safety and traffic flow stability. However, this type of method still relies on heuristic rules in essence, and its control effect is quite sensitive to parameter settings and traffic conditions.
[0318] When using the LQR optimal control strategy Figure 6 (d) The system achieves near-energy-optimal disturbance suppression, with the average speed rapidly recovering to equilibrium within approximately 35 seconds, and the vehicle speed curve exhibiting relatively smooth convergence characteristics. This indicates that LQR can provide superior performance under conditions of accurate model and predictable disturbance forms. However, in real-world traffic, differences in driving behavior, perception errors, and unmodeled dynamics are often difficult to fully characterize, which limits its robustness to some extent.
[0319] Figure 6 (e) demonstrates the method's... The performance of the DPS control strategy under strong disturbance scenarios. It can be seen that, compared to LQR, DPS control provides a smoother response throughout the process, more moderate suppression of velocity abrupt changes, and stronger adaptability to model uncertainties. Strong disturbances are completely eliminated within approximately 30-35 seconds, and almost no significant backpropagation region is observed in the spatiotemporal velocity field. This result demonstrates that the strategy not only achieves rapid stabilization but also effectively avoids secondary fluctuations, thus better meeting the comprehensive requirements of comfort and safety in actual traffic operations. From a control theory perspective, this method achieves a balance between local performance optimization and global stability.
[0320] In summary, while the aforementioned control strategies can all suppress traffic oscillations to some extent, their practical applicability varies significantly. FollowerStopper and PI control rely on empirical rules, sacrificing vehicle spacing to some extent. LQR achieves energy optimality under nominal model conditions but is highly sensitive to uncertainties. This method... The DPS control strategy demonstrates superior overall performance in terms of disturbance suppression speed, speed evolution smoothness, and adaptability to complex traffic environments, better meeting the requirements of autonomous vehicle control strategies in real-world traffic conditions. To further evaluate the potential advantages of multi-AV cooperative control in actual traffic systems, the performance of systems with one and two AVs was compared and analyzed. The settling time and control energy consumption characteristics of the traffic system under the DPS control strategy.
[0321] The experimental setup remained consistent with the single-AV scenario, with the length of the circular road being L=20n. Initially, all vehicles experienced minor disturbances near their equilibrium state to simulate unavoidable fluctuations in real-world traffic caused by differences in driving behavior, perception errors, or slight disturbances. This setup more closely resembles real-world road conditions rather than an idealized, undisturbed initial state.
[0322] Simulation results show that, under different system sizes and AV configurations, the hybrid transportation system can achieve [the desired performance]. The system recovered to a stable operating state under DPS control. Relevant performance indicators are as follows: Figure 7 As shown, the control energy is defined as The metric is used to measure the cumulative control cost required by the AV throughout the stabilization process. The results show that as the system size increases, the propagation path of disturbances in the traffic flow becomes longer, and the number of vehicles the AV needs to coordinate and influence increases. Therefore, both the system stabilization time and control energy show an upward trend. This phenomenon aligns with traffic dynamics intuition: in longer traffic flows, disturbance dissipation typically requires longer times and greater adjustment. However, regardless of the system size, the traffic flow eventually recovers to a steady state, indicating that the proposed control strategy has good scalability.
[0323] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those skilled in the art should understand that any modifications or equivalent substitutions to the technical solutions of the present invention without departing from the spirit and scope of the present invention should be covered within the scope of the claims of the present invention.
Claims
1. A hybrid traffic flow cooperative control method based on H∞ direct policy search, characterized in that, A mixed traffic system applied to a ring road, the system comprising at least one AV and at least one HDV; the method comprising the following steps: S1. Establish a linear state-space model of the hybrid transportation system: ; Where x(t) is the state vector, representing the motion deviation of each vehicle relative to the target steady state; u(t) is the control input vector of the AV; w(t) is the disturbance input vector; A, B, and H are system matrices; S2, Set the target equilibrium speed v* for the mixed traffic system; Each HDV is assumed to have the same dynamic characteristics, and its equilibrium spacing s* at v* is determined based on these dynamic characteristics; Set the desired vehicle spacing for each AV at v* And each The accessibility conditions determined by the geometric constraints of the ring road must be met. Set the control law of AV as u(t) = -K x(t), where K is the feedback gain matrix to be optimized; Define the performance output vector z(t): ; Where Q and R are positive definite weighting matrices; To minimize the H∞ norm of the closed-loop transfer function from w(t) to z(t) To optimize the objective, we construct a controller optimization problem with K as the optimization variable; S3. Using the direct policy search method, solve the controller optimization problem in S2 to obtain the result. Minimize the optimal gain matrix K*; S4. Configure K* in each AV; during system operation, each AV acquires or estimates x(t) in real time, and calculates and executes acceleration control commands according to the control law u(t) = -K* x(t).
2. The hybrid traffic flow cooperative control method based on H∞ direct policy search as described in claim 1, characterized in that, In step S1, based on the optimal speed model, a linearized car-following dynamics model of HDV containing external disturbance terms is established; a linearized dynamics model with acceleration as the control input is established for AV; the vehicle spacing deviation state and speed deviation state of all vehicles are aggregated into a system state vector x(t), the control input of all AVs is aggregated into a control input vector u(t), and the external disturbances experienced by all vehicles are aggregated into a disturbance input vector w(t). In step S2, the equilibrium distance s* of HDV under v* is determined according to the linearized car-following dynamics model of HDV.
3. The hybrid traffic flow cooperative control method based on H∞ direct strategy search according to claim 2, characterized in that, In step S1, a linearized car-following dynamics model of HDV is constructed by applying the car-following function of the optimal velocity model. At the equilibrium point The result is obtained by performing a first-order Taylor expansion at that point.
4. The hybrid traffic flow cooperative control method based on H∞ direct strategy search according to claim 3, characterized in that, The linearized car-following dynamics model takes the following form: ; ; in, , Let represent the distance and speed of vehicle i at time t, respectively; Let α1, α2, and α3 represent the distance deviation and speed deviation of vehicle i at time t, respectively; α1, α2, and α3 are constants greater than zero, representing the optimal speed model at the equilibrium point. linearization coefficient at; .
5. The hybrid traffic flow cooperative control method based on H∞ direct policy search as described in claim 4, characterized in that, The optimal velocity model is: ; Among them, s i (t) represents the distance between vehicle i and the preceding vehicle i-1; v represents the relative speed between vehicle i and the vehicle in front i-1; i (t) represents the speed of vehicle i at time t; A parameter indicating the driver's sensitivity to the speed difference between the current vehicle speed and the desired vehicle speed; A parameter indicating the driver's sensitivity to the relative speed between their vehicle and the vehicle in front; This represents the driver's desired speed function; The linearization coefficients α1, α2, α3 are: .
6. The hybrid traffic flow cooperative control method based on H∞ direct policy search as described in claim 5, characterized in that, The mathematical expression for the driver's desired speed function is: ; in, Indicates stationary distance and driving distance. This indicates the maximum distance between vehicles.
7. The hybrid traffic flow cooperative control method based on H∞ direct policy search according to claim 6, characterized in that, In step S2, the target equilibrium velocity v* and the HDV equilibrium spacing s* are related through the desired velocity function, v*=V(s*).
8. The hybrid traffic flow cooperative control method based on H∞ direct policy search as described in claim 1, characterized in that, In step S2, the accessibility condition determined by the geometric constraints of the ring road is: ; Among them, S AV Let L represent the set of autonomous vehicles, L be the total length of the ring road, n be the total number of vehicles, and k be the number of autonomous vehicles (AVs).
9. The hybrid traffic flow cooperative control method based on H∞ direct strategy search according to claim 1, characterized in that, In step S3, the direct policy search reduces the overall state feedback gain matrix K by iteratively adjusting it. This continues until the preset convergence condition is met.
10. The hybrid traffic flow cooperative control method based on H∞ direct policy search according to claim 1, characterized in that, In step S4, when acquiring or estimating x(t) in real time, the vehicle spacing and speed information of each vehicle are acquired through the vehicle-mounted sensing device and / or the roadside sensing device.