Signal intersection vehicle cloud multi-level cooperative control method
Through the multi-level collaborative control method of vehicle-cloud at signal intersections, combined with deep reinforcement learning and model prediction control, the speed and path planning of intelligent vehicles are optimized, and the problems of vehicle queues and high energy consumption at signal intersections are solved, and traffic efficiency is improved and energy saving is achieved.
Patent Information
- Application Number
- CN202510654238.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing traffic control methods are difficult to adapt to complex traffic needs, resulting in vehicle queues, waiting and frequent start and stopping at signal intersections, increasing energy consumption and reducing traffic efficiency. The simulation test fails to fully consider the impact of dynamic interaction between vehicles and traffic signals and the dissipation of queues on subsequent traffic flows.
Build a multi-level collaborative control method for vehicle-cloud at signal intersections, combine deep reinforcement learning and model prediction control, optimize smart vehicle speed and path planning through the vehicle-cloud layered architecture, and use DRL to adjust MPC output to achieve collaborative optimization control.
Reduce waiting time before traffic lights, save energy consumption, improve traffic efficiency, improve the adaptability and robustness of control strategies, and adapt to complex traffic environments.
Smart Images

Figure CN120472690A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent traffic control, and specifically relates to a vehicle-cloud multi-level collaborative control method for a signalized intersection. Background Art
[0002] In recent years, the development of intelligent connected technology has promoted the implementation of related applications such as autonomous driving. However, most of the current research focuses on autonomous driving technology for single vehicles, lacking research on simulation testing and collaborative control strategies for intelligent vehicles in complex traffic scenarios, especially in scenarios with continuous signalized intersections where social vehicles mix. Existing traffic control methods are mostly based on traditional signal timing or simple sensor control, which are difficult to adapt to complex traffic needs. Vehicles queuing, waiting, and frequent starting and stopping are common at signalized intersections, resulting in increased energy consumption and reduced traffic efficiency. At the same time, existing simulation testing methods often fail to fully consider the dynamic interaction between vehicles and traffic signals and the impact of queue dissipation on subsequent traffic flow.
[0003] Based on this, the present invention proposes a multi-level collaborative control method for vehicles and clouds at signalized intersections. This method constructs a vehicle-cloud hierarchical control platform, combines deep reinforcement learning (DRL) and model predictive control (MPC) technology, and optimizes the operation of smart vehicles. The intelligent vehicle information-physical system is divided into multiple operating scales, and hierarchical deconstruction is adopted to achieve separation of concerns. Through the vehicle-cloud hierarchical architecture, the advantages of cloud computing and vehicle-side real-time control are brought into play to optimize the speed and path planning of smart vehicles. Using DRL to adjust the MPC output, the adaptability and robustness of the strategy are improved, waiting time is reduced, energy consumption is reduced, and traffic efficiency is improved, providing a new path for the optimization of intelligent transportation systems. Summary of the Invention
[0004] In view of this, the present invention aims to provide a multi-level vehicle-cloud collaborative control method for signalized intersections, which realizes collaborative optimization control of vehicles in a signalized intersection environment through a multi-level control strategy that combines deep reinforcement learning with model predictive control.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A vehicle-cloud multi-level collaborative control method for a signalized intersection includes the following steps:
[0007] S1. Build an operational framework for the cyber-physical system of intelligent vehicles based on the traffic demand and spatiotemporal characteristics of signalized intersections.
[0008] S2. Build a vehicle-cloud layered control platform;
[0009] The physical layer uses SUMO to simulate road scenarios and load vehicle dynamic models, while vehicle-side computation and control are implemented in Python. The information layer also uses Python to integrate SUMO and platform data. After inputting data into the cloud control application platform, the predictive cruise algorithm is used to solve the optimal speed sequence and output the results to the vehicle.
[0010] S3. Optimize traffic control at signalized intersections using a multi-level strategy based on deep reinforcement learning and model predictive control (DRL-MPC). Use DRL to adjust MPC outputs and collaboratively optimize control outcomes.
[0011] S4. Collect data by interacting with the environment, use the proximal policy optimization algorithm (PPO) to train the DRL module, and update the parameters of the DRL module policy network and value network to continuously improve the control strategy.
[0012] Furthermore, the specific contents of step S1 of constructing the intelligent vehicle cyber-physical system operation framework are as follows:
[0013] Divide the cyber-physical system of intelligent vehicles into multiple levels from the perspective of time and space to match the actual needs of traffic control at signalized intersections;
[0014] Specifically, the hierarchical structure of the system is determined based on the layout of signalized intersections, traffic flow distribution, and the dynamic characteristics of vehicles. Each layer corresponds to a specific functional module, among which the bottom layer is responsible for real-time vehicle data collection and preliminary processing, the middle layer performs local optimization of traffic flow, and the top layer implements an overall collaborative control strategy. Information exchange between each layer is carried out through standardized interfaces to ensure the modularity and scalability of the entire system, so that the system configuration and functions can be flexibly adjusted according to different traffic scenarios and control requirements.
[0015] Furthermore, the specific contents of step S2 of constructing the vehicle-cloud hierarchical control platform are:
[0016] Use SUMO software to create realistic road network scenarios at the physical layer and define vehicle type parameters;
[0017] Integrate the vehicle's dynamics and kinematics models into the SUMO simulation environment. Model parameters include the vehicle's acceleration, deceleration, maximum speed, and turning radius to accurately simulate the vehicle's driving state under different traffic conditions.
[0018] Use Python programming language to implement vehicle-side computing and control logic, including data acquisition, processing vehicle sensor data, and executing control commands received from the cloud control application platform to control vehicle acceleration, deceleration, and lane changes;
[0019] At the information layer, Python is used to write data fusion code to integrate vehicle status data and road condition data obtained from the SUMO simulator, as well as traffic flow data and signal light status data obtained from the support platform. Data cleaning, alignment, and fusion processing are performed to eliminate data noise and inconsistency. The fused data is input into the cloud control application platform, which runs the predictive cruise algorithm solver to calculate the optimal speed sequence as the vehicle control instruction based on the current traffic conditions and vehicle status. The solved optimal speed sequence is converted into specific vehicle control instructions, and the control instructions are sent back to the vehicle-side platform at the physical layer through the communication interface to achieve real-time collaborative control of the vehicle.
[0020] Furthermore, in step S2, the road network scenario includes road layout, number of lanes, signal intersection locations and phase timing.
[0021] Furthermore, in step S2, the vehicle type parameters include vehicle length, maximum speed, acceleration and deceleration limits.
[0022] Furthermore, the specific content of step S3 is:
[0023] First, the state space is defined. The state space of the DRL covers the current traffic conditions of the signalized intersection and the preliminary control signals output by the MPC module. After normalization, they are input into the deep neural network.
[0024] Then the action space is defined. The action space of DRL is consistent with the output dimension of MPC. The initial value of the action element is limited to the range of [-1, 1], and is subsequently scaled to the corresponding amplitude according to the actual control requirements.
[0025] The reward function is then designed, taking into account traffic flow, vehicle waiting time, and energy consumption, while introducing penalty terms to avoid state constraint violations;
[0026] Finally, a collaborative optimization mechanism is designed. The DRL module updates the policy network parameters based on the reward signal and adjusts the action output, so that MPC and DRL work together to improve the adaptability and robustness of the control system.
[0027] Furthermore, the specific content of step S4 is:
[0028] By interacting with the environment to collect data, the Proximal Policy Optimization (PPO) algorithm is used to train the DRL module, and its policy network and value network parameters are updated to continuously improve the control strategy.
[0029] First, the policy network outputs the mean and standard deviation of continuous actions to generate actions that follow a Gaussian distribution. The value network estimates the state value and provides a benchmark for advantage function calculation.
[0030] Then, the experience replay buffer is managed, and the data generated by the interaction with the environment is stored in the experience replay buffer. A first-in-first-out (FIFO) strategy is used to maintain a constant data volume, and the data is updated regularly or in batches.
[0031] Then, the generalized advantage estimation (GAE) calculation is performed. By combining n-step returns with the temporal difference method, the advantage value is calculated using the GAE formula, the advantage estimation is smoothed, and the training stability is improved.
[0032] Furthermore, in step S4, the data generated by the interaction with the environment includes state, action, reward and next state.
[0033] Beneficial effects:
[0034] 1. Through the DRL-MPC multi-level control strategy, the coordinated optimization control of vehicles in signalized intersection environments is achieved, reducing waiting time at traffic lights and saving energy consumption.
[0035] 2. Through the vehicle-cloud layered collaborative control architecture, the advantages of cloud computing power and vehicle-side real-time control are fully utilized to achieve efficient optimization of the transportation system.
[0036] 3. Use deep reinforcement learning (DRL) to adjust the model predictive control (MPC) output to improve the adaptability and robustness of the control strategy, enabling it to better cope with complex traffic environments.
[0037] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a simulation platform framework;
[0039] Figure 2 Provides a layered control framework for the car-cloud;
[0040] Figure 3 This is the DRL-MPC framework diagram;
[0041] Figure 4 Adopting time-scale diagram for DRL-MPC control. DETAILED DESCRIPTION
[0042] To make the technical solutions, advantages, and purposes of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0043] The present invention provides a vehicle-cloud multi-level collaborative control method for a signalized intersection, comprising the following steps:
[0044] S1. Build an operational framework for the cyber-physical system of intelligent vehicles based on the traffic demand and spatiotemporal characteristics of signalized intersections.
[0045] The operational framework of the constructed intelligent vehicle cyber-physical system is specifically divided into multiple layers from the perspective of time and space to match the actual needs of traffic control at signalized intersections. Specifically, the hierarchical structure of the system is determined based on the layout of the signalized intersection, the distribution of traffic flow, and the dynamic characteristics of vehicles. Each layer corresponds to a specific functional module. For example, the bottom layer is responsible for real-time vehicle data collection and preliminary processing, the middle layer performs local optimization of traffic flow, and the top layer implements an overall collaborative control strategy. Through this layered approach, the system can more efficiently process complex traffic information and achieve precise control of traffic at signalized intersections. At the same time, information exchange between each layer is carried out through standardized interfaces, ensuring the modularity and scalability of the entire system, so that the system configuration and functions can be flexibly adjusted according to different traffic scenarios and control requirements.
[0046] S2. Build a vehicle-cloud layered control platform. The physical layer uses SUMO to simulate road scenarios and load vehicle dynamic models. Vehicle-side computation and control are implemented in Python. The information layer also uses Python to integrate SUMO and platform data. After inputting data into the cloud control application platform, a predictive cruise control algorithm is used to determine the optimal speed sequence, and the results are output to the vehicle.
[0047] SUMO software and Pythen software are used to realize the road scene and the cloud control system designed in this paper. The simulation platform framework is as follows Figure 1As shown in the figure, SUMO software is used at the physical layer to create a realistic road network scenario, including detailed information such as road layout, number of lanes, signal intersection locations, and phase timing. Vehicle type parameters such as vehicle length, maximum speed, acceleration, and deceleration limits are also defined. The vehicle's dynamic and kinematic models are integrated into the SUMO simulation environment, including parameters such as acceleration, deceleration, maximum speed, and turning radius, to accurately simulate the vehicle's driving state under different traffic conditions. The Python programming language is used to implement vehicle-side computing and control logic, including data acquisition, processing vehicle sensor data (such as speed, position, acceleration, etc.), and executing control commands received from the cloud control application platform to control vehicle acceleration, deceleration, lane changing, and other behaviors. In the information layer, Python is used to write data fusion code to integrate vehicle status data and road condition data obtained from the SUMO simulator, as well as traffic flow data and signal light status data obtained from the support platform, and perform data cleaning, alignment and fusion processing to eliminate data noise and inconsistency; the fused data is input into the cloud control application platform, which runs the predictive cruise algorithm solver and calculates the optimal speed sequence as the vehicle control instruction based on the current traffic conditions and vehicle status; the optimal speed sequence obtained is converted into specific vehicle control instructions (such as acceleration, speed setting value, etc.), and these instructions are sent back to the vehicle-side platform of the physical layer through the communication interface to realize real-time collaborative control of the vehicle. The vehicle-cloud layered control framework is as follows: Figure 2 .
[0048] S3. Use a multi-level strategy of deep reinforcement learning-model predictive control (DRL-MPC) to optimize traffic control at signalized intersections. Use DRL to adjust the MPC output and coordinately optimize the control effect. The DRL-MPC framework is as follows: Figure 3 As shown;
[0049] The MPC module operates at the lower level and provides the basic control input, which is optimized over the forecast window based on the objective function of the MPC and the traffic demand predicted by the associated nominal model. The objective function is given according to the control purpose, and the state and input constraints are explicitly considered in the optimization process. In order to improve the optimality of the MPC output and avoid serious constraint violations, the DRL module works at the upper level and modifies the MPC output through a learning process that interacts with the real system. The state space of the DRL agent includes the signalized intersection traffic state and the MPC output, while the reward function is designed to be a complement to the MPC objective function so that the two modules can collaboratively optimize the overall goal. In addition, the traffic demand is also input into the DRL agent, and the penalty for constraint violation is added to the reward function. The actions of the DRL have the same dimensionality as the MPC output, but the order of magnitude of its elements is smaller. The proposed MPC-DRL framework has a hierarchical structure, such as Figure 2 As shown:
[0050] Assume that the dynamic model of the signalized intersection is a discrete time model and the simulation sampling time is T s , the DRL module works at a control sampling time of T d Then T s 、T d 、T c The relationship between them is described as:
[0051] T c =m1·T d =m1·m2·T s , m1,m2∈N + ,m1>1 (1)
[0052] Note that for simplicity, we assume that the simulation sampling step, DRL control step, and MPC control step coincide. Therefore, the overall control input of the combined framework is the combination of the MPC output and the DRL output, and every T d Updated once per time unit. c The corresponding control step k d (i.e. k d T d ∈[k c T c ,(k+1)T c The overall control input of )) is:
[0053] u c (k d )=sat(u rl (k d )+u b (k c )) (2)
[0054] Among them, u rl is the output of DRL, u b is the output of the MPC, using a saturation function to ensure additive control input u c (k d ) satisfies the constraints, defined on the element as:
[0055]
[0056] Among them, u min and u max To control the minimum and maximum allowed values of the corresponding elements in the input, Figure 3 The different time scales of MPC and DRL control sampling times are illustrated, as well as u rl How to modify u b , DRL-MPC control uses a time scale such as Figure 4 shown.
[0057] The standard MPC procedure is executed in the MPC module. The system state is x, and each simulation sampling step is k. s Update, with MPC control step size k c The corresponding simulation sampling step is:
[0058] {k C m,k C m+1,…,k C m+m-1}(4)
[0059] Among them, m = m1m2, so the step size is k in the simulation sampling S =k C At each control step k, the actual state of the signalized intersection is measured and input into the MPC module. C Solve the following optimization problem:
[0060]
[0061] in, Indicates that the length N p,c The control variables to be optimized on the forecast window, u b (k c ) is the MPC in control step k C The output of s (k s ) is the simulation sampling step k s The MPC output at and represents the simulation sampling step k s The predicted future state. In addition, d(k^s) contains the simulated sampling step k s Estimated traffic demand. Equation (1) represents the system state evolution driven by the signalized intersection dynamics f, and Equations (2) and (3) represent the constraints of the MPC state set X and the MPC output set U, respectively. Since the frequency of system sampling is higher than the frequency of MPC output generation, Equation (4) maps the MPC output to each simulation sampling step within the prediction window, so that the output can be realized in the system dynamics.
[0062] In the formula, J(k s ) represents the time interval [k s T s ,(k s +1)T s ], Indicates that the length is N P,C The variables that need to be optimized within the prediction window of Represents the predicted future state under the control step size. P,S and N P,Care the predicted horizon lengths calculated according to the simulation sampling step and the MPC time step, respectively, where N P,S =N P,C m.
[0063] The signalized intersection network is considered as a Markov decision process (MDP), which can be represented by a five-tuple,<S,A,P,R,γ> The present invention defines the state space S, the action space A and the reward distribution R. P represents the transition probability between states, which is implicitly defined by the signal intersection network model. γ is a discount factor used to define future rewards. The RL module works at a low level and has a higher frequency than the MPC module. In order to avoid too frequent changes in the control input of the signal intersection network, the control sampling time T of the RL module is set to 0. d Greater than the simulation sampling time T s Therefore, the corresponding RL control step size k d The simulation sampling step is:
[0064] {k d m2,k d m2+1,…,k d m2+m2-1} (6)
[0065] The state, action, and reward of RL are updated at each control step and are defined as follows:
[0066] Status x rl (k d )∈S: The state space of RL should contain all the necessary information of the framework. The introduced RL uses a deep neural network. To facilitate learning, the state of the input layer is normalized to the same order of magnitude. Therefore, the normalized state is:
[0067]
[0068] in, are the normalized state of the signalized intersection network, MPC output, and simulation sampling step k. d The actual requirements of m2 and the framework in the previous control step k d Note that the purpose of adding the fourth element is to provide additional knowledge about the overall control input of the combined framework, which helps to avoid wild fluctuations in the control input.
[0069] Action rl (k d )∈A: action u rl Used to modify the output u of MPC b . Therefore they have the same dimension, namely dimu rl =dimub For simplicity we assume that the action space is continuous. Note that the action u rl is generated from the output layer of the DNN, and its elements have initial values between [-1, 1]. Therefore, these values are scaled down to the actual control input before being added. rl The magnitude ratio of the elements u b The element is small, so in this frame, u b is the dominant control input, providing basic performance, while u rl It is an auxiliary function control input, its function is to improve performance, set to high frequency. rl Satisfy the inequality defining the action space A:
[0070] -w u ΔU≤u rl ≤w u ΔU (8)
[0071] Where ΔU=u max -u min , where u max and u min is the upper and lower bounds of the MPC output, w u ∈[0,1] is the factor that determines u rl To u b Scaling parameter of the influence degree.
[0072] Reward r(x rl (k d ),u rl (k d ))∈R: In order to increase the interactivity of MPC, it is necessary to coordinate MPC and RL to achieve optimal control performance. Therefore, the objective function J of the MPC module should be included in the reward function. MPC :
[0073]
[0074] Where, P s Represents the violation of state constraints, w p >0 is the penalty weight parameter, and the penalty state violates the constraint. Here we assume that r t (k d ) is r(x rl (k d ),u rl (k d )) is an equivalent representation of RL in controlling step size k d In order to evaluate the objective function J of MPC, the real-time state x(k s ), the output u of MPCs (k s ) can be obtained by the formula. In addition, according to the traffic state x(k s ) Directly calculate P s (k s ), where k s =k d m2+1,…,k d m2+m2. Here the reward is a negative value, so R is a set of negative numbers.
[0075] A Deep Actor-Critic training framework was considered. Since lane changing and green wave traffic involve both discrete and continuous actions, the RL agent uses the Proximal Policy Optimization (PPO) algorithm.
[0076] In reinforcement learning, the goal of the policy gradient algorithm is to optimize a parameterized policy π θ (a|s), update the policy parameters θ by maximizing the expected return. This goal can be expressed by the policy gradient theorem:
[0077]
[0078] in, R represents the expectation of sampling the current policy, and γ is the discount factor. t is the cumulative reward starting from time step t. To optimize this objective, we need to calculate the gradient and use it to update the policy:
[0079]
[0080] Where logπ θ (a t |s t ) is the current strategy in state s t Next select action a t The logarithmic probability of A. t Is the advantage function, which indicates the current action a t The advantage relative to the benchmark value is usually estimated by the value function (Critic), the advantage function A t Used to measure the current action relative to the benchmark value (such as the state value function V(s t )) is good or bad, it can be calculated by the following formula:
[0081] A t =R t -V(s t ) (12)
[0082] Among them, V(s t ) is the state s t The estimated value of is usually estimated by the value network (Critic).
[0083] S4. Collect data by interacting with the environment, train the DRL module using the Proximal Policy Optimization (PPO) algorithm, and update the parameters of the DRL module's policy network and value network to continuously improve the control strategy;
[0084] The clipping objective function of the PPO algorithm, PPO adds a clipping mechanism to the traditional policy gradient algorithm to limit the amplitude of each policy update, thereby avoiding instability caused by excessive policy updates. The objective function of PPO is a clipped objective function, which consists of two parts: the original policy gradient objective and a clipping term, which can be expressed as:
[0085]
[0086] Among them, π θ (a t |s t ) is the new policy (current policy) in state s t Select action a t The probability of is the probability of the old strategy (the strategy before the update), ε is a small hyperparameter, usually between [0.1, 0.3], used to limit the amplitude of each strategy update, the role of the clip function is to convert the probability ratio It is limited to the range of [1-ε,1+ε] to prevent the policy update from being too large.
[0087] The ultimate goal of PPO is to maximize the clipped objective function L CLIP (θ), that is:
[0088]
[0089] This optimization process is carried out through multiple updates, and each update calculates the advantage function A based on the sampling data of the current strategy t , and then use the above objective function to optimize the policy. This implementation is suitable for both offline training and can be embedded in an online control system. By continuously updating the network parameters, the agent can gradually learn a better policy.
[0090] The training process of the vehicle-cloud layered PPO-MPC algorithm is shown in Algorithm 1. Both the policy network and the value network adopt a multi-layer feedforward neural network structure. The policy network outputs the mean and standard deviation of continuous actions (for reference trajectory parameters and platoon formation parameters) and the category probability of discrete actions (for lane change decisions). The main hyperparameter settings are as follows:
[0091] Table 1
[0092]
[0093] Algorithm 1
[0094] 1. Initial input: Initialize policy network parameters θ, value network parameters φ, and training hyperparameters
[0095] 2. Initial conditions: Initialize the policy network π θ and value network V φ , initialize the experience playback buffer D
[0096] 3. Loop through e=1, 2, ...M rounds of training;
[0097] a. Initialize the traffic environment, set the initial traffic demand and signal status
[0098] b. Cycle execution of MPC control cycle k c =0,1,…T / T c -1, solve the formula
[0099] c. Loop execution of PPO control step k d =k c m1 to (k c m1+1)-1
[0100] 4. Randomly sample mini-batch data from buffer D
[0101] 5. Calculate the target value and generalized advantage estimate (GAE)
[0102] 6. Perform multiple PPO updates;
[0103] 7. Update target network parameters;
[0104] 8. Output: Optimized strategy parameters and value parameters
[0105] End
[0106] Algorithm 4-1 details the training process for the vehicle-cloud layered PPO-MPC intersection control algorithm. This algorithm employs a multi-timescale control framework, where the MPC module operates at a lower frequency, providing the basic control strategy, while the PPO module operates at a higher frequency, adjusting and optimizing the MPC output. The training process consists of four core steps:
[0107] The first step is environment interaction and data collection. The algorithm uses MPC and PPO to interact with the traffic environment, collecting state, action, reward, and next-state transition tuples and storing them in the experience replay buffer. This step ensures the diversity and representativeness of the training data.
[0108] The second step is advantage function calculation. The algorithm uses the generalized advantage estimation (GAE) method to calculate the advantage value of each state-action pair. This method combines the low bias of n-step returns and the low variance of the temporal difference method to help improve training stability.
[0109] The third step is to update the policy and value networks. The algorithm updates the policy network using the PPO clipping objective, which limits policy updates to avoid excessive policy changes while maximizing policy performance. The value network is updated by minimizing the mean squared error between the predicted and target values.
[0110] The fourth step is target network update. The algorithm regularly updates the target network parameters, which further improves the stability of training.
[0111] Through multiple rounds of iterative training, the algorithm gradually optimizes the parameters of the policy network and the value network, ultimately achieving a control strategy that effectively coordinates intersection traffic. This training process fully utilizes the model prediction capabilities of MPC and the adaptive learning capabilities of PPO to achieve multi-scale control optimization for vehicle-cloud collaboration.
[0112] Simulation experiments have verified that the proposed vehicle-cloud hierarchical DRL-MPC framework demonstrates excellent generalization and optimization capabilities, significantly improving intersection efficiency and reducing vehicle energy consumption while ensuring vehicle safety. This hierarchical framework, which integrates deep reinforcement learning and model predictive control, provides a new technical approach for traffic management in intelligent connected environments and lays a theoretical foundation for subsequent research on multi-vehicle collaborative control.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of protection of the present invention.
Claims
1. A vehicle-cloud multi-level collaborative control method for a signalized intersection, characterized in that: The following steps are involved: S1. Build an operational framework for the cyber-physical system of intelligent vehicles based on the traffic demand and spatiotemporal characteristics of signalized intersections. S2. Build a vehicle-cloud layered control platform; The physical layer uses SUMO to simulate road scenarios and load vehicle dynamic models, while vehicle-side computation and control are implemented in Python. The information layer also uses Python to integrate SUMO and platform data. After inputting data into the cloud control application platform, the predictive cruise algorithm is used to solve the optimal speed sequence and output the results to the vehicle. S3. Optimize traffic control at signalized intersections using a multi-level strategy based on deep reinforcement learning and model predictive control (DRL-MPC). Use DRL to adjust MPC outputs and collaboratively optimize control outcomes. S4. Collect data by interacting with the environment, use the proximal policy optimization algorithm (PPO) to train the DRL module, and update the parameters of the DRL module policy network and value network to continuously improve the control strategy.
2. The vehicle-cloud multi-level collaborative control method for a signalized intersection according to claim 1 is characterized by: The specific contents of step S1 of constructing the intelligent vehicle cyber-physical system operation framework are as follows: Divide the cyber-physical system of intelligent vehicles into multiple levels from the perspective of time and space to match the actual needs of traffic control at signalized intersections; Specifically, the hierarchical structure of the system is determined based on the layout of signalized intersections, traffic flow distribution, and the dynamic characteristics of vehicles. Each layer corresponds to a specific functional module, among which the bottom layer is responsible for real-time vehicle data collection and preliminary processing, the middle layer performs local optimization of traffic flow, and the top layer implements an overall collaborative control strategy. Information exchange between each layer is carried out through standardized interfaces to ensure the modularity and scalability of the entire system, so that the system configuration and functions can be flexibly adjusted according to different traffic scenarios and control requirements.
3. The vehicle-cloud multi-level collaborative control method for a signalized intersection according to claim 2 is characterized by: The specific contents of step S2 of constructing the vehicle-cloud hierarchical control platform are: Use SUMO software to create realistic road network scenarios at the physical layer and define vehicle type parameters; Integrate the vehicle's dynamics and kinematics models into the SUMO simulation environment. Model parameters include the vehicle's acceleration, deceleration, maximum speed, and turning radius to accurately simulate the vehicle's driving state under different traffic conditions. Use Python programming language to implement vehicle-side computing and control logic, including data acquisition, processing vehicle sensor data, and executing control commands received from the cloud control application platform to control vehicle acceleration, deceleration, and lane changes; At the information layer, Python is used to write data fusion code to integrate vehicle status data and road condition data obtained from the SUMO simulator, as well as traffic flow data and signal light status data obtained from the support platform. Data cleaning, alignment, and fusion processing are performed to eliminate data noise and inconsistency. The fused data is input into the cloud control application platform, which runs the predictive cruise algorithm solver to calculate the optimal speed sequence as the vehicle control instruction based on the current traffic conditions and vehicle status. The solved optimal speed sequence is converted into specific vehicle control instructions, and the control instructions are sent back to the vehicle-side platform at the physical layer through the communication interface to achieve real-time collaborative control of the vehicle.
4. The vehicle-cloud multi-level collaborative control method for a signalized intersection according to claim 3 is characterized by: In step S2, the road network scenario includes road layout, number of lanes, signalized intersection locations and phase timing.
5. The vehicle-cloud multi-level collaborative control method for a signalized intersection according to claim 4 is characterized by: In step S2, the vehicle type parameters include vehicle length, maximum speed, acceleration and deceleration limits.
6. The vehicle-cloud multi-level collaborative control method for a signalized intersection according to claim 5 is characterized by: The specific content of step S3 is: First, the state space is defined. The state space of the DRL covers the current traffic conditions of the signalized intersection and the preliminary control signals output by the MPC module. After normalization, they are input into the deep neural network. Then the action space is defined. The action space of DRL is consistent with the output dimension of MPC. The initial value of the action element is limited to the range of [-1, 1], and is subsequently scaled to the corresponding amplitude according to the actual control requirements. The reward function is then designed, taking into account traffic flow, vehicle waiting time, and energy consumption, while introducing penalty terms to avoid state constraint violations; Finally, a collaborative optimization mechanism is designed. The DRL module updates the policy network parameters based on the reward signal and adjusts the action output, so that MPC and DRL work together to improve the adaptability and robustness of the control system.
7. The vehicle-cloud multi-level collaborative control method for a signalized intersection according to claim 6 is characterized by: The specific content of step S4 is: By interacting with the environment to collect data, the Proximal Policy Optimization (PPO) algorithm is used to train the DRL module, and its policy network and value network parameters are updated to continuously improve the control strategy. First, the policy network outputs the mean and standard deviation of continuous actions to generate actions that follow a Gaussian distribution. The value network estimates the state value and provides a benchmark for advantage function calculation. Then, the experience replay buffer is managed, and the data generated by the interaction with the environment is stored in the experience replay buffer. A first-in-first-out (FIFO) strategy is used to maintain a constant data volume, and the data is updated regularly or in batches. Then, the generalized advantage estimation (GAE) calculation is performed. By combining n-step returns with the temporal difference method, the advantage value is calculated using the GAE formula, the advantage estimation is smoothed, and the training stability is improved.
8. The vehicle-cloud multi-level collaborative control method for a signalized intersection according to claim 7 is characterized by: In step S4, the data generated by the interaction with the environment includes state, action, reward and next state.
Citation Information
Patent Citations
Regional boundary main intersection signal control method based on deep reinforcement learning
CN113392577A
Signal lamp control method based on two-stage attention mechanism and deep reinforcement learning
CN114038212A
Traffic signal control method and system combining reinforcement learning and model prediction control
CN119889065A